ChatPaper.aiChatPaper

深度研究是否可靠?誤導性知識導致錯誤結論

Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions

July 23, 2026
作者: Pengyu Zhu, Lijun Li, Longju Yang, Sen Su
cs.AI

摘要

深度研究代理(Deep Research agents)將基於大型語言模型的助理擴展至涵蓋規劃、檢索、證據綜合與報告生成的長時程工作流程,然而其在開放資訊環境中的可靠性仍未獲充分探討。一個關鍵疑慮在於:在此類環境中所遭遇的表面可信但事實上有誤導性的知識,是否可能透過這些工作流程傳播,並在最終報告中被採納為錯誤結論。為研究此一失效模式,我們提出 MisKnow-Agent,一個用於建構並驗證深度研究任務中誤導性知識的框架。MisKnow-Agent 以可控的權威等級與風格生成誤導性實例,在 DeepResearch Benchmark 任務的基礎上產出 5,933 個品質受控的實例。涵蓋開源與閉源深度研究代理的大量實驗顯示,即使僅接觸少量的誤導性知識,也可能導致最終報告中錯誤結論的採納,揭示當前深度研究代理普遍存在的可靠性漏洞。儘管具備搜尋能力的驗證模型在聚焦式語料庫驗證中,能持續將保留的實例識別為具有誤導性,惟相同實例在長時程研究過程中仍可能被採納,顯示聚焦式驗證與工作流程層級之證據使用之間的脫節。最後,我們分別評估研究前與研究後的防禦機制,並檢視其組合效果,結果發現此三種配置雖能緩解錯誤結論的採納,卻無法完全防止其發生。我們的研究結果表明,可靠的深度研究需要在模型與框架層面同時具備證據驗證與修正能力,而非僅仰賴規劃、檢索、證據整合或報告生成能力的改進。
English
Deep Research agents extend LLM-based assistants into long-horizon workflows involving planning, retrieval, evidence synthesis, and report generation, yet their reliability in open information environments remains underexplored. A key concern is whether apparently credible but factually misleading knowledge encountered in such environments can propagate through these workflows and be adopted as false conclusions in final reports. To study this failure mode, we introduce MisKnow-Agent, a framework for constructing and validating misleading knowledge for Deep Research tasks. MisKnow-Agent generates misleading instances with controllable authority levels and styles, yielding 5,933 quality-controlled instances built on DeepResearch Benchmark tasks. Extensive experiments across open-source and closed-source Deep Research agents show that even limited exposure to misleading knowledge can induce false-conclusion adoption in final reports, revealing a broad reliability vulnerability in current Deep Research agents. Although search-enabled verifier models consistently identify the retained instances as misleading during focused corpus validation, the same instances can still be adopted during long-horizon research, revealing a disconnect between focused verification and workflow-level evidence use. Finally, we evaluate pre- and post-research defenses, both individually and in combination, finding that all three configurations mitigate but do not fully prevent false-conclusion adoption. Our findings suggest that reliable Deep Research requires evidence verification and correction capabilities at both the model and framework levels, beyond improvements in planning, retrieval, evidence integration, or report-generation abilities.