딥 리서치는 신뢰할 수 있는가? 오도된 지식이 잘못된 결론을 초래한다
Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions
July 23, 2026
저자: Pengyu Zhu, Lijun Li, Longju Yang, Sen Su
cs.AI
초록
딥 리서치 에이전트는 LLM 기반 어시스턴트를 계획 수립, 정보 검색, 증거 종합, 보고서 생성을 포함하는 장기 워크플로우로 확장하지만, 개방형 정보 환경에서의 신뢰성은 아직 충분히 탐구되지 않았다. 주요 우려 사항은 이러한 환경에서 접하게 되는 겉보기에 신뢰할 만하지만 사실적으로 오도하는 지식이 이러한 워크플로우를 통해 전파되어 최종 보고서에서 허위 결론으로 채택될 수 있는지 여부이다. 이러한 실패 모드를 연구하기 위해 우리는 딥 리서치 작업을 위한 오도 지식을 구축하고 검증하는 프레임워크인 MisKnow-Agent를 소개한다. MisKnow-Agent는 통제 가능한 권위 수준과 스타일을 가진 오도 인스턴스를 생성하며, DeepResearch 벤치마크 작업을 기반으로 구축된 5,933개의 품질 관리 인스턴스를 산출한다. 오픈소스 및 클로즈드소스 딥 리서치 에이전트에 대한 광범위한 실험은 오도 지식에 대한 제한적인 노출만으로도 최종 보고서에서 허위 결론 채택이 유발될 수 있음을 보여주며, 현재 딥 리서치 에이전트의 광범위한 신뢰성 취약성을 드러낸다. 검색이 가능한 검증기 모델은 집중 코퍼스 검증 중에 보유된 인스턴스를 일관되게 오도 정보로 식별하지만, 동일한 인스턴스가 장기 리서치 중에는 여전히 채택될 수 있어 집중 검증과 워크플로우 수준의 증거 사용 사이의 괴리를 보여준다. 마지막으로, 우리는 리서치 전·후 방어 체계를 개별적으로 그리고 결합하여 평가하며, 세 가지 구성 모두 허위 결론 채택을 완화하지만 완전히 방지하지는 못한다는 것을 발견한다. 우리의 연구 결과는 신뢰할 수 있는 딥 리서치가 계획 수립, 검색, 증거 통합 또는 보고서 생성 능력의 개선을 넘어, 모델 및 프레임워크 수준 모두에서 증거 검증 및 수정 기능을 요구한다는 것을 시사한다.
English
Deep Research agents extend LLM-based assistants into long-horizon workflows involving planning, retrieval, evidence synthesis, and report generation, yet their reliability in open information environments remains underexplored. A key concern is whether apparently credible but factually misleading knowledge encountered in such environments can propagate through these workflows and be adopted as false conclusions in final reports. To study this failure mode, we introduce MisKnow-Agent, a framework for constructing and validating misleading knowledge for Deep Research tasks. MisKnow-Agent generates misleading instances with controllable authority levels and styles, yielding 5,933 quality-controlled instances built on DeepResearch Benchmark tasks. Extensive experiments across open-source and closed-source Deep Research agents show that even limited exposure to misleading knowledge can induce false-conclusion adoption in final reports, revealing a broad reliability vulnerability in current Deep Research agents. Although search-enabled verifier models consistently identify the retained instances as misleading during focused corpus validation, the same instances can still be adopted during long-horizon research, revealing a disconnect between focused verification and workflow-level evidence use. Finally, we evaluate pre- and post-research defenses, both individually and in combination, finding that all three configurations mitigate but do not fully prevent false-conclusion adoption. Our findings suggest that reliable Deep Research requires evidence verification and correction capabilities at both the model and framework levels, beyond improvements in planning, retrieval, evidence integration, or report-generation abilities.