Apodex Discovery: 발견적 인공지능의 평가와 구축을 위한 현실 벤치마크 및 환경
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence
August 11, 2026
저자: Brian Wang, Bin Feng, Xiaoman Pan, Chenyang An, Felix Liu, Tangqi Fang, Gongbo Sun, Lingfeng Shen, Ning Wang, Handuo Zhang, Feng Chen, Fuchao Yang, Xiang Wang, Jiacheng Lin, Siting Li, Zixuan Liu, Chi Han, Zhenhailong Wang, Kunlun Zhu, Lawrence Zhao, Yueqi Guo, Kailong Wen, Feng Xing, Yiling Guo, Lidong Bing, David Tan, Bo An, Heng Ji, Sheng Wang
cs.AI
초록
아폴로가 달에 도달할 수 있었던 것은 단지 엔지니어들이 어려운 방정식을 풀 수 있었기 때문만은 아니었다. 그것은 먼 포부를 명시적 목표, 시뮬레이션, 검증, 그리고 반복적 수정으로 구성된 임무 아키텍처로 전환함으로써 성공했다. AI는 이제 유사한 전환점에 직면해 있다: 프런티어 모델은 문제, 도구, 성공 기준이 명시되면 어려운 작업을 해결할 수 있지만, 결과적으로 중요한 실제 세계의 문제는 실행 가능하거나 검증 가능한 형태로 제공되는 경우가 드물다.
우리는 발견적 AI를 구축하고 평가하기 위한 프레임워크인 Apodex Discovery를 소개한다. 이는 파운데이션 모델, 하네스, 도구, 제어 정책으로 구성된 시스템인 heavy-duty solver를 기반으로 하며, 확장된 상태 유지형 검증 가능한 조사를 추구한다. 이 프레임워크는 세 가지 핵심 구성요소를 가진다. 첫째, 문제 발굴 프로세스는 16개 섹터에 걸친 561개 산업을 조사하고, 423개의 고가치 실제 세계 문제를 수집했으며, 초기 릴리스를 위해 20개를 선정했다. 둘째, 공통 환경-작업-에피소드 추상화는 데이터, 도구, 제약 조건, 피드백, 궤적 기록, 그리고 중간 산출물과 최종 제출물의 검증을 제공한다. 셋째, HDS6은 최종 작업 성공과 독립적으로 도구, 수리, 대안, 일관성, 증거, 범위를 평가한다.
AAV 캡시드 설계에서 Apodex는 생존성, 조직 친화성, 구조 예측, 생성 설계 전반에 걸쳐 발표된 최신 기술 대비 7%의 성능 향상을 달성했다. 약물 재창출 및 재제형화에서, 작업 특화 생체의학 환경은 GPT-5.5와 GPT-5.6-sol의 평균 정규화 예측 점수를 동일한 폐쇄형 백본 대비 각각 2.5점과 7.6점 향상시켰다. 통제된 절제 실험은 고정된 TRACES 에피소드 인터페이스가 성능 차이를 특정 솔버 구성요소에 귀속시킬 수 있게 함을 보여준다. Apodex Discovery는 AI 평가를 사전 정의된 벤치마크를 넘어 진정한 발견을 목표로 하는 검증 가능한 조사로 이동시킨다.
English
Apollo did not reach the Moon merely because its engineers could solve difficult equations. It succeeded by turning a distant ambition into a mission architecture of explicit objectives, simulation, verification, and repeated correction. AI now faces a similar transition: frontier models can solve difficult tasks once the problem, tools, and success criteria are specified, yet consequential real-world challenges rarely arrive in an executable or verifiable form.
We introduce Apodex Discovery, a framework for building and evaluating discoverative AI through the heavy-duty solver, a system comprising a foundation model, harness, tools, and control policies that pursues extended, stateful, verifiable investigations. It has three core components. First, a problem-scouting process surveyed 561 industries across 16 sectors, assembled 423 high-value real-world problems, and selected 20 for the initial release. Second, a common environment-task-episode abstraction provides data, tools, constraints, feedback, trajectory recording, and verification of intermediate artifacts and final submissions. Third, HDS6 evaluates Tools, Repair, Alternatives, Coherence, Evidence, and Scope independently of final-task success.
In AAV capsid design, Apodex surpassed the published state of the art by 7% across viability, tropism, structure prediction, and generative design. In drug repurposing and reformulation, a task-specific biomedical environment improved the mean normalized prediction score of GPT-5.5 and GPT-5.6-sol by 2.5 and 7.6 points over the same closed-book backbone. Controlled ablations show that the fixed TRACES episode interface enables attribution of performance differences to specific solver components. Apodex Discovery moves AI evaluation beyond predefined benchmarks toward verifiable investigations aimed at genuine discovery.