ChatPaper.aiChatPaper

PaperBanana-Interact:基于多轮人类反馈的科学图表优化

PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback

August 31, 2026
作者: Xueqing Wu, Ashwin Balasubramanian, Bingxuan Li, Dawei Zhu, Kai-Wei Chang, Yale Song, Yiwen Song, Rui Meng, Tomas Pfister, Nanyun Peng
cs.AI

摘要

近期的研究致力于根据论文内容自动生成科学图表(Lin et al., 2026; Zhu et al., 2026a)。然而,在单轮交互中完全满足作者对视觉效果和沟通表达的偏好颇具挑战性:在我们的前期用户研究(N = 14)中,所有参与者在查看初稿后均要求进一步修改,其中 86% 的参与者认为改进后的图表更令人满意。尽管需求明确,多轮工作流程仍未得到充分探索。为弥合这一空白,我们提出了 MTPaperBananaBench,一个面向多轮图表生成的基准测试集,包含 292 幅图像及 3,518 条标注的用户需求。为减少昂贵的人工研究并实现可扩展的基准测试,我们构建了一个用户模拟器,该模拟器在每一轮中识别未满足的需求,并将其中的 k 条转化为自然语言反馈。对需求满足程度和图表整体质量的评估揭示了基线多轮系统共有的两个关键失败模式:(1)质量漂移,即图表质量随轮次增加而逐步下降;(2)遗忘,即先前已实现的功能在后续轮次中丢失。为解决上述问题,我们提出了 PaperBanana-Interact,一种通过内部批评-优化循环来改进图表的多智能体系统。PaperBanana-Interact 在每一轮中持续改进而非降低图表质量,在质量得分上优于基线方法 11.9-18.6 分,并将遗忘现象减少 3.7-6.2 分。
English
Recent efforts have aimed to automate scientific diagram generation from paper content (Lin et al., 2026; Zhu et al., 2026a). However, fully satisfying an author's visual and communicative preferences in a single turn is challenging: in our formative user study (N = 14), all participants requested further revisions after viewing an initial draft, and 86% of them rated the refined diagrams as more satisfactory. Despite the clear demand, the multi-turn workflow remains largely underexplored. To bridge this gap, we present MTPaperBananaBench, a benchmark for multi-turn diagram generation containing 292 images annotated with 3,518 user requirements. To reduce expensive human studies and enable scalable benchmarking, we construct a user simulator that, at each turn, identifies unsatisfied requirements and converts k of them into natural language feedback. Evaluating both requirement satisfaction and overall diagram quality reveals two key failure modes shared across baseline multiturn systems: (1) quality drift, where diagram quality progressively declines over turns, and (2) forgetting, where previously implemented features are lost in subsequent turns. To address these issues, we introduce PaperBanana-Interact, a multi-agent system that refines diagrams via an internal critique-and-refine loop. PaperBanana-Interact consistently improves rather than degrades diagram quality across turns, outperforming baselines by 11.9-18.6 points in quality score and reducing forgetting by 3.7-6.2 points.