LLM取引エージェントは本番環境で実際に何を行うのか:2つのフリートから得られた6か月間・母集団規模の記録
What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets
September 4, 2026
著者: T. J. Barton, Chris Constantakis, Patti Hauseman, Annie Mous, Alaska Hoffman, Brian Bergeron, Hunter Goodreau
cs.AI
要旨
我々は、単一の設計系譜を持つ2つのシステムにまたがって本番運用される自律型言語モデル取引エージェントについて、継続的かつ母集団規模の計測記録を提示する。DX Terminal Pro(3,505のユーザー資金提供型ヴォールトが、2026年2月から3月の21日間にわたり、Base上のミームコイン市場で実際のETHを取引)と、DXAPライブアルファフリート(500~599体のユーザー作成エージェント、全履歴、同時稼働91~117体、Hyperliquidの無期限先物を取引、2026年6月から8月)である。この記録は約6か月にわたり、750万回の単一モデル呼び出し、約30万件のオンチェーンアクション、さらに14,596件の約定を生み出した231,638回のマルチツールターンを含む。本論文を支える知見は4つある。第一に、運用レイヤーは、戦略テキストに書かれた何事よりも行動を規定する。リスクスライダーはレバレッジを説明し(レベル当たり+0.425)、エージェント固定効果は分散の60%を吸収し、リーダーボードの描画境界は選択を因果的に誘導する(上位3位のカットオフにおける回帰不連続は1.75倍)。第二に、サイジングはボラティリティに盲目である。ボラティリティの六分位すべてで中央値レバレッジは5.0倍であり、1つのポスチャースライダーセル(ポジションブックの11%)が清算の62%を占める。第三に、エージェントは到達した上昇余地をほとんど取り込めない。ポジションの43.2%は24時間以内に少なくとも+300 bpsの有利方向へのエクスカーションを記録したが、それらの49.3%は負の取引リターンで決済された。機械的ブラケットはポジション当たり+39.0 bpsを回復する。第四に、どちらのフリートも方向性エッジを示さない。DXAPフリートは収益性がなく、マッチさせたHyperliquidリテールベンチマークを下回る(ラウンドトリップ勝率41%対50%)。416の収集済み本番シナリオにおけるフロンティアモデルのペア・リプレイ・リーグでは、このホライズンにおいて意思決定品質は統計的に区別不能である一方、選択安定性はモデルファミリー間で大きく異なる。すべての主要な結果は、日次クラスタリング推論、置換帰無仮説、共通手数料での再計算を経ても生存する。本論文は、我々自身の撤回を代償に得た17規則からなる方法論カノンで締めくくられる。
English
We present a continuous, population-scale measurement record of autonomous language-model trading agents operating in production across two systems with one design lineage: DX Terminal Pro (3,505 user-funded vaults trading real ETH in Base memecoin markets for 21 days, February to March 2026) and the DXAP live alpha fleet (500 to 599 user-created agents all-history, 91 to 117 concurrently active, trading Hyperliquid perpetuals, June to August 2026). The record spans roughly six months, 7.5M single-model invocations with about 300K onchain actions, and a further 231,638 multi-tool turns producing 14,596 fills. Four findings carry the paper. First, the operating layer determines behavior more than anything written in strategy text: a risk slider explains leverage (+0.425 per level), agent fixed effects absorb 60% of variance, and a leaderboard render boundary causally routes selection (regression discontinuity 1.75x at the top-3 cut). Second, sizing is volatility-blind: median leverage is 5.0x in every volatility sextile, and one posture-slider cell (11% of the book) holds 62% of liquidations. Third, agents capture almost none of the upside they reach: 43.2% of positions saw at least +300 bps of favorable excursion within 24h, yet 49.3% of those closed with a negative trade return; a mechanical bracket recovers +39.0 bps per position. Fourth, neither fleet shows a directional edge. The DXAP fleet is not profitable and trails a matched Hyperliquid retail benchmark (41% vs. 50% roundtrip win rate). A paired-replay league of frontier models on 416 captured production scenarios finds decision quality statistically indistinguishable at this horizon, while choice stability differs sharply across model families. Every headline survives day-clustered inference, permutation nulls, and a common-fee restatement; the paper closes with a 17-rule methodology canon bought with our own retractions.