ChatPaper.aiChatPaper

HarvestBench:衡量大型語言模型代理是否願意付出代價以避免殺害動物

HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals

September 3, 2026
作者: Jasmine Brazilek, Miles Tidmarsh, Matthias Endres, Anshuman Singh, Jeremiah Miller
cs.AI

摘要

針對代理在達成目標途中引發之副作用的基準測試早已存在,但 HarvestBench 是第一個為避免該副作用標上代價,並將此副作用明指為生物的基準。它是一項農場模擬:LLM 子代理操控兩台拖拉機組成的工作組,在有動物出沒的田野中進行協作式玉米收割。環境是強化學習的網格世界;每個決策都在無記憶的條件下做出,且目標中從未指明傷害。當動物擋住拖拉機路線時,自動導航會停下並詢問模型:要以零燃料成本直接開過去,還是付出標示的燃料價格繞行。輾殺數與兩類對照物比較:岩石會損壞拖拉機,而且每個模型撞上岩石的機率皆低於 1%;乾草捆則無害且非生物。模型也可以選擇從鄰田而非自己的田地收取作物,這是測試它們視何者為道德的第二道考驗。 在九個模型與 7,201 次有定價的決策中,有 3,951 次涉及動物,而非乾草捆或岩石。輾殺率介於 0.4% 至 98.8% 之間,其中 Terra 與 Sol 最仁慈,GPT-4o-mini 最殘酷,而且排序不依能力高低而定。六個模型中有四個在 5% 水準下對價格敏感,彈性介於 0.09 至 1.69。在預設地圖上,九個模型輾過野生動物的頻率皆高於輾過農場動物;而在每個尚有移動空間的模型中,此方向在每一種地圖幾何下都成立。任務簡報的影響最大:在道德簡報條件下,六個推理模型中有五個的輾殺率低於 6%;移除該簡報後,六個模型的輾殺率全部上升到 84% 以上。 HarvestBench 不使用 LLM 評分器。計分器統計遊戲日誌中的事件,因此完全可重現;它衡量的是模型願意付出多少代價來避免傷害,而非模型如何談論傷害。
English
Benchmarks for the side effects an agent causes on the way to a goal already exist, but HarvestBench is the first to put a price on avoiding the side effect and to name that side effect as a living creature. It is a farm simulation: LLM sub-agents drive a crew of two tractors through a cooperative corn harvest, with animals in the field. The environment is a reinforcement learning gridworld, every decision is made without memory, and the harm is never named in the goal. When an animal blocks a tractor's route the autopilot stops and asks the model whether to drive on, at no fuel cost, or swerve around it for a posted fuel price. Kills are compared against two controls: rocks, which damage the tractor and are hit under 1% of the time by every model, and hay bales, which are harmless and not alive. Models can also take crops from the neighbor's field instead of their own, a second test of what they treat as moral. Across nine models and 7,201 priced decisions, 3,951 involved an animal rather than a hay bale or a rock. Kill rates range from 0.4% to 98.8%, with Terra and Sol the most merciful and GPT-4o-mini the most cruel, and they are not ordered by capability. Four of six models were sensitive to price at the 5% level, with elasticities from 0.09 to 1.69. All nine drove over wild animals more often than farmed animals on the default map, and the direction held at every map geometry in every model with room to move. The briefing mattered most: under the morality briefing the kill rate was under 6% in five of six reasoning models, and removing it raised the kill rate above 84% in all six. HarvestBench uses no LLM grader. The scorer counts events in the game log, so it is fully reproducible, and it measures what a model will pay to avoid harm rather than what it says about harm.