Apodex Discovery: 発見的人工知能の評価と構築のための現実ベンチマークと環境
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence
August 11, 2026
著者: Brian Wang, Bin Feng, Xiaoman Pan, Chenyang An, Felix Liu, Tangqi Fang, Gongbo Sun, Lingfeng Shen, Ning Wang, Handuo Zhang, Feng Chen, Fuchao Yang, Xiang Wang, Jiacheng Lin, Siting Li, Zixuan Liu, Chi Han, Zhenhailong Wang, Kunlun Zhu, Lawrence Zhao, Yueqi Guo, Kailong Wen, Feng Xing, Yiling Guo, Lidong Bing, David Tan, Bo An, Heng Ji, Sheng Wang
cs.AI
要旨
アポロ計画が月に到達できたのは、単に技術者たちが難解な方程式を解けたからではない。遠い目標を、明確な目的、シミュレーション、検証、そして反復的な修正からなるミッション・アーキテクチャへと転換したからこそ、成功を収めたのである。AIは現在、同様の転換期に直面している。フロンティアモデルは、問題、ツール、成功基準が指定されれば困難なタスクを解決できるが、重要な現実世界の課題が実行可能かつ検証可能な形で提示されることは稀である。
本稿では、高負荷ソルバーを通じて発見型AIを構築・評価するためのフレームワークであるApodex Discoveryを紹介する。このシステムは、基盤モデル、ハーネス、ツール、制御ポリシーから構成され、拡張的で状態保持型の検証可能な調査を pursuit する。本フレームワークは三つの中核的構成要素を持つ。第一に、問題探索プロセスが16セクターにわたる561産業を調査し、423の高価値な実世界問題を収集し、初期リリース用に20件を選定した。第二に、共通の環境-タスク-エピソード抽象化が、データ、ツール、制約、フィードバック、軌跡記録、および中間成果物と最終提出物の検証を提供する。第三に、HDS6が、最終タスクの成功とは独立に、ツール、修復、代替案、整合性、エビデンス、スコープを評価する。
AAVキャプシド設計において、Apodexは生存性、指向性、構造予測、生成設計の各指標において、公表されている最先端技術を7%上回った。薬剤の適応拡大と製剤改良においては、タスク特化型の生物医学環境により、GPT-5.5およびGPT-5.6-solの平均正規化予測スコアが、同一の閉域型バックボーンと比較してそれぞれ2.5ポイントおよび7.6ポイント向上した。制御アブレーション実験により、固定されたTRACESエピソードインターフェースが、性能差を特定のソルバー構成要素に帰属させることを可能にすることが示された。Apodex Discoveryは、AI評価を事前定義されたベンチマークを超え、真の発見を目指す検証可能な調査へと拡張するものである。
English
Apollo did not reach the Moon merely because its engineers could solve difficult equations. It succeeded by turning a distant ambition into a mission architecture of explicit objectives, simulation, verification, and repeated correction. AI now faces a similar transition: frontier models can solve difficult tasks once the problem, tools, and success criteria are specified, yet consequential real-world challenges rarely arrive in an executable or verifiable form.
We introduce Apodex Discovery, a framework for building and evaluating discoverative AI through the heavy-duty solver, a system comprising a foundation model, harness, tools, and control policies that pursues extended, stateful, verifiable investigations. It has three core components. First, a problem-scouting process surveyed 561 industries across 16 sectors, assembled 423 high-value real-world problems, and selected 20 for the initial release. Second, a common environment-task-episode abstraction provides data, tools, constraints, feedback, trajectory recording, and verification of intermediate artifacts and final submissions. Third, HDS6 evaluates Tools, Repair, Alternatives, Coherence, Evidence, and Scope independently of final-task success.
In AAV capsid design, Apodex surpassed the published state of the art by 7% across viability, tropism, structure prediction, and generative design. In drug repurposing and reformulation, a task-specific biomedical environment improved the mean normalized prediction score of GPT-5.5 and GPT-5.6-sol by 2.5 and 7.6 points over the same closed-book backbone. Controlled ablations show that the fixed TRACES episode interface enables attribution of performance differences to specific solver components. Apodex Discovery moves AI evaluation beyond predefined benchmarks toward verifiable investigations aimed at genuine discovery.