ChatPaper.aiChatPaper

PaperBanana-Interact:基於多輪人類回饋的科學圖表精化

PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback

August 31, 2026
作者: Xueqing Wu, Ashwin Balasubramanian, Bingxuan Li, Dawei Zhu, Kai-Wei Chang, Yale Song, Yiwen Song, Rui Meng, Tomas Pfister, Nanyun Peng
cs.AI

摘要

近期的研究致力於從論文內容自動生成科學圖表(Lin et al., 2026; Zhu et al., 2026a)。然而,在單一輪次中完全滿足作者對於視覺與溝通的偏好極具挑戰:在我們的前導使用者研究中(N = 14),所有參與者在檢視初稿後均要求進一步修改,其中86%的人認為精煉後的圖表更令人滿意。儘管需求明確,多輪次工作流程仍然在很大程度上未被充分探索。為了填補此一缺口,我們提出MTPaperBananaBench,一個用於多輪次圖表生成的基準,包含292張圖像,並標註了3,518項使用者需求。為降低昂貴的人類研究成本並實現可擴展的基準測試,我們建構了一個使用者模擬器,在每一輪中識別未滿足的需求,並將其中k項轉換為自然語言回饋。同時評估需求滿足程度與整體圖表品質後,揭示了基線多輪系統共有的兩個關鍵失敗模式:(1) 品質漂移,即圖表品質在後續輪次中逐漸下降;(2) 遺忘,即先前已實作的特徵在後續輪次中遺失。為了解決這些問題,我們引入了PaperBanana-Interact,這是一個透過內部評論與修正迴圈來優化圖表的多代理系統。PaperBanana-Interact在多輪次中持續提升而非降低圖表品質,在品質評分上超越基線11.9至18.6分,並將遺忘情況減少3.7至6.2分。
English
Recent efforts have aimed to automate scientific diagram generation from paper content (Lin et al., 2026; Zhu et al., 2026a). However, fully satisfying an author's visual and communicative preferences in a single turn is challenging: in our formative user study (N = 14), all participants requested further revisions after viewing an initial draft, and 86% of them rated the refined diagrams as more satisfactory. Despite the clear demand, the multi-turn workflow remains largely underexplored. To bridge this gap, we present MTPaperBananaBench, a benchmark for multi-turn diagram generation containing 292 images annotated with 3,518 user requirements. To reduce expensive human studies and enable scalable benchmarking, we construct a user simulator that, at each turn, identifies unsatisfied requirements and converts k of them into natural language feedback. Evaluating both requirement satisfaction and overall diagram quality reveals two key failure modes shared across baseline multiturn systems: (1) quality drift, where diagram quality progressively declines over turns, and (2) forgetting, where previously implemented features are lost in subsequent turns. To address these issues, we introduce PaperBanana-Interact, a multi-agent system that refines diagrams via an internal critique-and-refine loop. PaperBanana-Interact consistently improves rather than degrades diagram quality across turns, outperforming baselines by 11.9-18.6 points in quality score and reducing forgetting by 3.7-6.2 points.