大型語言模型交易代理在生產環境中實際做了什麼:一份來自兩個機群、為期六個月的群體規模紀錄
What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets
September 4, 2026
作者: T. J. Barton, Chris Constantakis, Patti Hauseman, Annie Mous, Alaska Hoffman, Brian Bergeron, Hunter Goodreau
cs.AI
摘要
我們提出一份連續、群體規模的測量紀錄,記錄自主語言模型交易代理在生產環境中的運作,涵蓋兩個具有同一設計血統的系統:DX Terminal Pro(3,505 個使用者出資金庫,於 2026 年 2 月至 3 月的 21 天期間在 Base 迷因幣市場交易真實 ETH)以及 DXAP 實盤 alpha 艦隊(全歷史 500 至 599 個使用者建立的代理,91 至 117 個同時活躍,交易 Hyperliquid 永續合約,2026 年 6 月至 8 月)。此紀錄橫跨約六個月、750 萬次單一模型呼叫,其中約 30 萬次鏈上操作,另有 231,638 個多工具回合,產生 14,596 筆成交。本文的核心為四項發現。第一,操作層對行為的決定程度高於策略文本中所寫的任何內容:風險滑桿可解釋槓桿(每級 +0.425),代理固定效應吸收 60% 的變異,而排行榜渲染邊界會因果性地引導選擇(前三名切點處的迴歸不連續為 1.75 倍)。第二,部位規模設定對波動率視而不見:在每一個波動率六分位中,槓桿中位數都是 5.0 倍,而單一姿態滑桿格(占部位簿 11%)就承受了 62% 的清算。第三,代理幾乎未捕捉到它們曾觸及的上漲空間:43.2% 的部位在 24 小時內曾出現至少 +300 個基點的有利偏移,然而其中 49.3% 以負交易報酬平倉;機械式括號單可為每個部位挽回 +39.0 個基點。第四,兩支艦隊皆未展現方向性優勢。DXAP 艦隊並不獲利,且落後於配對的 Hyperliquid 散戶基準(來回交易勝率 41% 對 50%)。在 416 個擷取的生產情境上,前沿模型的配對重播競賽發現,在此時間範圍內決策品質在統計上無法區分,而選擇穩定性在不同模型家族之間差異顯著。每一項主要結果都經得起按日分群的推論、置換虛無檢定,以及共同費用重述;本文最後以一套 17 條規則的方法學正典作結,那是以我們自身的撤稿換來的。
English
We present a continuous, population-scale measurement record of autonomous language-model trading agents operating in production across two systems with one design lineage: DX Terminal Pro (3,505 user-funded vaults trading real ETH in Base memecoin markets for 21 days, February to March 2026) and the DXAP live alpha fleet (500 to 599 user-created agents all-history, 91 to 117 concurrently active, trading Hyperliquid perpetuals, June to August 2026). The record spans roughly six months, 7.5M single-model invocations with about 300K onchain actions, and a further 231,638 multi-tool turns producing 14,596 fills. Four findings carry the paper. First, the operating layer determines behavior more than anything written in strategy text: a risk slider explains leverage (+0.425 per level), agent fixed effects absorb 60% of variance, and a leaderboard render boundary causally routes selection (regression discontinuity 1.75x at the top-3 cut). Second, sizing is volatility-blind: median leverage is 5.0x in every volatility sextile, and one posture-slider cell (11% of the book) holds 62% of liquidations. Third, agents capture almost none of the upside they reach: 43.2% of positions saw at least +300 bps of favorable excursion within 24h, yet 49.3% of those closed with a negative trade return; a mechanical bracket recovers +39.0 bps per position. Fourth, neither fleet shows a directional edge. The DXAP fleet is not profitable and trails a matched Hyperliquid retail benchmark (41% vs. 50% roundtrip win rate). A paired-replay league of frontier models on 416 captured production scenarios finds decision quality statistically indistinguishable at this horizon, while choice stability differs sharply across model families. Every headline survives day-clustered inference, permutation nulls, and a common-fee restatement; the paper closes with a 17-rule methodology canon bought with our own retractions.