ChatPaper.aiChatPaper

Apodex Discovery:評估與建構發現型人工智慧之現實基準與環境

Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence

August 11, 2026
作者: Brian Wang, Bin Feng, Xiaoman Pan, Chenyang An, Felix Liu, Tangqi Fang, Gongbo Sun, Lingfeng Shen, Ning Wang, Handuo Zhang, Feng Chen, Fuchao Yang, Xiang Wang, Jiacheng Lin, Siting Li, Zixuan Liu, Chi Han, Zhenhailong Wang, Kunlun Zhu, Lawrence Zhao, Yueqi Guo, Kailong Wen, Feng Xing, Yiling Guo, Lidong Bing, David Tan, Bo An, Heng Ji, Sheng Wang
cs.AI

摘要

阿波羅之所以能登月,並非只因工程師能解開困難的方程式,而在於它把遙遠的抱負轉化為一套具明確目標、模擬、驗證與反覆修正的任務架構。人工智慧如今也正面臨相似的轉型:只要問題、工具與成功標準被清楚界定,前沿模型就能解決困難任務;然而,具重大影響的真實世界挑戰,很少會以可執行或可驗證的形式出現。 我們提出 Apodex Discovery,這是一個透過「高強度求解器」(heavy-duty solver)來建構與評估發現式人工智慧的框架。高強度求解器由基礎模型、執行框架、工具與控制策略組成,能夠進行長時間、具狀態性且可驗證的調查。它包含三大核心組成。第一,問題勘察流程調查了16個產業部門中的561個行業,彙整出423個高價值的真實世界問題,並從中篩選20個作為初始釋出。第二,一個共同的「環境—任務—回合」抽象層,提供資料、工具、限制、回饋、軌跡記錄,以及對中間產物與最終提交結果的驗證。第三,HDS6 獨立於最終任務是否成功之外,分別評估工具、修復、替代方案、連貫性、證據與範疇。 在 AAV 衣殼設計方面,Apodex 在可行性、嗜性、結構預測與生成設計等層面上,超越了已發表的當前最佳成果達 7%。在藥物再利用與再配方方面,一個任務特定的生物醫學環境,使 GPT-5.5 與 GPT-5.6-sol 相較於同一閉卷式骨幹模型的平均正規化預測分數,分別提升了 2.5 分與 7.6 分。受控消融實驗顯示,固定的 TRACES 回合介面能將效能差異歸因於特定的求解器元件。Apodex Discovery 使 AI 評估超越預先定義的基準,朝向以真實發現為目標的可驗證調查邁進。
English
Apollo did not reach the Moon merely because its engineers could solve difficult equations. It succeeded by turning a distant ambition into a mission architecture of explicit objectives, simulation, verification, and repeated correction. AI now faces a similar transition: frontier models can solve difficult tasks once the problem, tools, and success criteria are specified, yet consequential real-world challenges rarely arrive in an executable or verifiable form. We introduce Apodex Discovery, a framework for building and evaluating discoverative AI through the heavy-duty solver, a system comprising a foundation model, harness, tools, and control policies that pursues extended, stateful, verifiable investigations. It has three core components. First, a problem-scouting process surveyed 561 industries across 16 sectors, assembled 423 high-value real-world problems, and selected 20 for the initial release. Second, a common environment-task-episode abstraction provides data, tools, constraints, feedback, trajectory recording, and verification of intermediate artifacts and final submissions. Third, HDS6 evaluates Tools, Repair, Alternatives, Coherence, Evidence, and Scope independently of final-task success. In AAV capsid design, Apodex surpassed the published state of the art by 7% across viability, tropism, structure prediction, and generative design. In drug repurposing and reformulation, a task-specific biomedical environment improved the mean normalized prediction score of GPT-5.5 and GPT-5.6-sol by 2.5 and 7.6 points over the same closed-book backbone. Controlled ablations show that the fixed TRACES episode interface enables attribution of performance differences to specific solver components. Apodex Discovery moves AI evaluation beyond predefined benchmarks toward verifiable investigations aimed at genuine discovery.