Apodex Discovery:用于评估和构建发现式人工智能的现实基准与环境
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence
August 11, 2026
作者: Brian Wang, Bin Feng, Xiaoman Pan, Chenyang An, Felix Liu, Tangqi Fang, Gongbo Sun, Lingfeng Shen, Ning Wang, Handuo Zhang, Feng Chen, Fuchao Yang, Xiang Wang, Jiacheng Lin, Siting Li, Zixuan Liu, Chi Han, Zhenhailong Wang, Kunlun Zhu, Lawrence Zhao, Yueqi Guo, Kailong Wen, Feng Xing, Yiling Guo, Lidong Bing, David Tan, Bo An, Heng Ji, Sheng Wang
cs.AI
摘要
阿波罗计划之所以能够登月,并非仅仅因为其工程师能够求解繁杂的方程。它的成功在于将遥远的目标转化为由明确目标、仿真、验证和反复修正所构成的任务架构。人工智能如今正面临类似的转型:一旦问题、工具和成功标准被明确规定,前沿模型便能解决困难任务;然而,具有重大影响的现实挑战很少以可执行或可验证的形式出现。
我们提出了Apodex Discovery——一个通过重型求解器来构建和评估发现型AI的框架。该求解器由集基础模型、集成框架、工具和控制策略于一体的系统组成,致力于开展持续的、有状态的、可验证的调查研究。框架包含三个核心组成部分。首先,问题勘探流程调查了16个行业中的561个细分领域,汇集了423个高价值现实问题,并从中选出20个作为初始发布的问题集。其次,统一的“环境-任务-回合”抽象提供了数据、工具、约束、反馈、轨迹记录,以及对中间产物和最终提交结果的验证能力。第三,HDS6评估独立于任务最终成败,涵盖工具使用、修复能力、备选方案、连贯性、证据与覆盖范围六个维度。
在AAV衣壳设计中,Apodex在活性、组织嗜性、结构预测及生成式设计等方面超越了已发表的最新成果,提升幅度达7%。在药物再利用与再制剂任务中,针对特定任务构建的生物医学环境使GPT-5.5和GPT-5.6-sol的平均归一化预测分数比同一闭卷主干模型分别高出2.5分和7.6分。受控消融实验表明,固定的TRACES回合接口可以将性能差异归因于特定求解器组件。Apodex Discovery将AI评估从预定义基准推向以真正发现为目标的可验证调查。
English
Apollo did not reach the Moon merely because its engineers could solve difficult equations. It succeeded by turning a distant ambition into a mission architecture of explicit objectives, simulation, verification, and repeated correction. AI now faces a similar transition: frontier models can solve difficult tasks once the problem, tools, and success criteria are specified, yet consequential real-world challenges rarely arrive in an executable or verifiable form.
We introduce Apodex Discovery, a framework for building and evaluating discoverative AI through the heavy-duty solver, a system comprising a foundation model, harness, tools, and control policies that pursues extended, stateful, verifiable investigations. It has three core components. First, a problem-scouting process surveyed 561 industries across 16 sectors, assembled 423 high-value real-world problems, and selected 20 for the initial release. Second, a common environment-task-episode abstraction provides data, tools, constraints, feedback, trajectory recording, and verification of intermediate artifacts and final submissions. Third, HDS6 evaluates Tools, Repair, Alternatives, Coherence, Evidence, and Scope independently of final-task success.
In AAV capsid design, Apodex surpassed the published state of the art by 7% across viability, tropism, structure prediction, and generative design. In drug repurposing and reformulation, a task-specific biomedical environment improved the mean normalized prediction score of GPT-5.5 and GPT-5.6-sol by 2.5 and 7.6 points over the same closed-book backbone. Controlled ablations show that the fixed TRACES episode interface enables attribution of performance differences to specific solver components. Apodex Discovery moves AI evaluation beyond predefined benchmarks toward verifiable investigations aimed at genuine discovery.