Iris:攀登搜索前沿
Iris: Climbing to the Search Frontier
September 3, 2026
作者: Ziyuan Liu, Hengqi Liu, Zichuan Wang, Yang Qin, Jiachen Liang, Xu Chu, Shaowei Chen, Yuantao Gu, Mu Chuan
cs.AI
摘要
我们提出了 Iris-mini 和 Iris-pro,这是两个分别基于 35B-A3B 与 397B-A17B 规模训练的搜索智能体,并介绍了它们背后的数据管道和训练方案。任务是通过对网页语料库超链接结构进行逆向构建而来:我们在从种子页面及其出链提炼出的实体图上构造多跳链,将每个非答案实体改写为描述性指代,使得任何线索都无法通过字符串匹配来解析;我们只纳入这样的问题:参考模型在闭卷情况下无法回答,但一旦提供支持性证据就能解决。随后,这些问题被转化为轨迹,并在监督微调(SFT)之前分别在轨迹层面和轮次层面进行过滤。之后,策略以实时搜索为交互环境,通过强化学习(RL)进行优化;奖励评判器和观测摘要器均部署在训练集群内部,过长的 rollout 在请求层面被中断,并在下一步从其已提交的前缀恢复执行。我们以交替方式执行这两个阶段,这一流程称为“SFT-RL 爬升”:将每一轮 RL 中难度最高且成功解决的 rollout 与效率最高的 rollout 返回到下一轮监督训练。由于在这些基准上,推理时的上下文管理带来的增益大于大多数已报告的系统间差异,我们在每个基准上都分别评测启用和不启用该机制的情况,并保持工具集、上下文长度限制和评判器不变。所有结果均来自单一 ReAct 智能体,没有子代理,也没有测试时验证。启用上下文管理后,在 BrowseComp、BrowseComp-ZH、DeepSearchQA 和 HLE 上,两个模型分别达到 82.2/84.8/86.9/52.3(Iris-mini)和 88.6/85.1/92.9/56.4(Iris-pro),是各自参数范围内开源搜索智能体中综合表现最强的结果。我们计划发布模型权重,以及数据构建、训练和评估的完整流程。
English
We present Iris-mini and Iris-pro, two search agents trained at the 35B-A3B and 397B-A17B scales, together with the data pipeline and training recipe behind them. Tasks are reverse-constructed from the hyperlink structure of a web corpus: we author multi-hop chains over an entity graph distilled from a seed page and its out-links, rewrite every non-answer entity into a descriptive reference so that no clue can be resolved by string matching, and admit only questions that a reference model fails closed-book yet solves once the supporting evidence is supplied. These questions are then turned into trajectories, which are filtered at both the trajectory and the turn level before SFT. The policy is then optimized by RL against live search, with the reward judge and the observation summarizer served inside the training cluster, and with over-long rollouts interrupted at the request level and resumed from their committed prefix at the next step. We alternate the two stages in a procedure we call SFT-RL climbing, returning the hardest solved and most efficient rollouts of each RL round to the next supervised pass. Because inference-time context management is worth more on these benchmarks than most reported differences between systems, we evaluate every benchmark both with and without it, holding the tool set, the context limit, and the judge fixed. All results come from a single ReAct agent, with no sub-agents and no test-time verification. With management enabled, on BrowseComp, BrowseComp-ZH, DeepSearchQA, and HLE the two models reach 82.2/84.8/86.9/52.3 and 88.6/85.1/92.9/56.4, the strongest overall results among open-source search agents in their respective parameter ranges. We plan to release the model weights together with the complete recipe for data construction, training, and evaluation.