LLM 트레이딩 에이전트는 실제 운영 환경에서 무엇을 하는가: 두 플릿에서 얻은 6개월간의 모집단 규모 기록
What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets
September 4, 2026
저자: T. J. Barton, Chris Constantakis, Patti Hauseman, Annie Mous, Alaska Hoffman, Brian Bergeron, Hunter Goodreau
cs.AI
초록
우리는 하나의 설계 계보를 공유하는 두 시스템에서 실제 운영 환경으로 작동하는 자율 언어 모델 트레이딩 에이전트에 대한 연속적이고 모집단 규모의 측정 기록을 제시한다: DX Terminal Pro(2026년 2월부터 3월까지 21일 동안 Base 밈코인 시장에서 실제 ETH를 거래한 3,505개의 사용자 자금 지원 볼트)와 DXAP 라이브 알파 플릿(전체 이력 기준 500~599개의 사용자 생성 에이전트, 동시 활성 91~117개, Hyperliquid 무기한 선물 거래, 2026년 6월~8월)이다. 이 기록은 약 6개월에 걸쳐 있으며, 약 30만 건의 온체인 액션을 동반한 750만 건의 단일 모델 호출과 14,596건의 체결을 생성한 추가 231,638건의 다중 도구 턴을 포함한다. 네 가지 발견이 이 논문의 핵심을 이룬다. 첫째, 운영 계층은 전략 텍스트에 적힌 그 무엇보다도 행동을 더 크게 결정한다: 리스크 슬라이더는 레버리지를 설명하고(레벨당 +0.425), 에이전트 고정효과는 분산의 60%를 흡수하며, 리더보드 렌더링 경계는 선택을 인과적으로 유도한다(상위 3개 컷오프에서 회귀 불연속 1.75배). 둘째, 포지션 크기 결정은 변동성 맹목적이다: 중위 레버리지는 모든 변동성 육분위에서 5.0배이며, 하나의 운용 태세 슬라이더 셀(장부의 11%)이 청산의 62%를 차지한다. 셋째, 에이전트는 도달한 상방 이익의 거의 아무것도 포착하지 못한다: 포지션의 43.2%가 24시간 내 최소 +300bp의 유리한 익스커전을 보였지만, 그중 49.3%는 음의 거래 수익률로 종료되었으며, 기계적 브래킷은 포지션당 +39.0bp를 회복한다. 넷째, 어느 플릿도 방향성 우위를 보이지 않는다. DXAP 플릿은 수익성이 없으며, 매칭된 Hyperliquid 소매 벤치마크에 뒤처진다(왕복 승률 41% 대 50%). 416개의 포착된 프로덕션 시나리오에 대한 프런티어 모델들의 페어드 리플레이 리그는 이 지평에서 의사결정 품질이 통계적으로 구분되지 않음을 발견하는 한편, 선택 안정성은 모델 계열에 따라 크게 다르다. 모든 핵심 결과는 일별 클러스터 추론, 순열 영가설, 공통 수수료 재진술을 견뎌낸다; 논문은 우리 자신의 철회를 대가로 얻은 17개 규칙의 방법론 정전으로 마무리된다.
English
We present a continuous, population-scale measurement record of autonomous language-model trading agents operating in production across two systems with one design lineage: DX Terminal Pro (3,505 user-funded vaults trading real ETH in Base memecoin markets for 21 days, February to March 2026) and the DXAP live alpha fleet (500 to 599 user-created agents all-history, 91 to 117 concurrently active, trading Hyperliquid perpetuals, June to August 2026). The record spans roughly six months, 7.5M single-model invocations with about 300K onchain actions, and a further 231,638 multi-tool turns producing 14,596 fills. Four findings carry the paper. First, the operating layer determines behavior more than anything written in strategy text: a risk slider explains leverage (+0.425 per level), agent fixed effects absorb 60% of variance, and a leaderboard render boundary causally routes selection (regression discontinuity 1.75x at the top-3 cut). Second, sizing is volatility-blind: median leverage is 5.0x in every volatility sextile, and one posture-slider cell (11% of the book) holds 62% of liquidations. Third, agents capture almost none of the upside they reach: 43.2% of positions saw at least +300 bps of favorable excursion within 24h, yet 49.3% of those closed with a negative trade return; a mechanical bracket recovers +39.0 bps per position. Fourth, neither fleet shows a directional edge. The DXAP fleet is not profitable and trails a matched Hyperliquid retail benchmark (41% vs. 50% roundtrip win rate). A paired-replay league of frontier models on 416 captured production scenarios finds decision quality statistically indistinguishable at this horizon, while choice stability differs sharply across model families. Every headline survives day-clustered inference, permutation nulls, and a common-fee restatement; the paper closes with a 17-rule methodology canon bought with our own retractions.