PaperBanana-Interact: 다중 턴 인간 피드백을 통한 과학적 다이어그램 정제
PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback
August 31, 2026
저자: Xueqing Wu, Ashwin Balasubramanian, Bingxuan Li, Dawei Zhu, Kai-Wei Chang, Yale Song, Yiwen Song, Rui Meng, Tomas Pfister, Nanyun Peng
cs.AI
초록
최근 논문 내용으로부터 과학 다이어그램 생성을 자동화하려는 시도가 있어 왔다(Lin et al., 2026; Zhu et al., 2026a). 그러나 한 번의 턴으로 저자의 시각적 및 의사소통 선호도를 완전히 충족시키는 것은 어렵다. 우리의 형성적 사용자 연구(N=14)에서 모든 참가자가 초안을 본 후 추가 수정을 요청했으며, 86%가 개선된 다이어그램이 더 만족스럽다고 평가했다. 명확한 수요에도 불구하고 다중 턴 워크플로우는 크게 탐구되지 않은 상태로 남아 있다. 이러한 격차를 해소하기 위해, 우리는 292개의 이미지와 3,518개의 사용자 요구사항으로 주석이 달린 다중 턴 다이어그램 생성 벤치마크인 MTPaperBananaBench를 제시한다. 값비싼 인간 연구를 줄이고 확장 가능한 벤치마킹을 가능하게 하기 위해, 각 턴에서 충족되지 않은 요구사항을 식별하고 그중 k개를 자연어 피드백으로 변환하는 사용자 시뮬레이터를 구축한다. 요구사항 충족과 전반적인 다이어그램 품질을 모두 평가한 결과, 기준 다중 턴 시스템들에서 공통적으로 나타나는 두 가지 주요 실패 모드가 드러났다: (1) 품질 저하(quality drift), 즉 턴이 진행됨에 따라 다이어그램 품질이 점진적으로 하락하는 현상, (2) 망각(forgetting), 즉 이전 턴에서 구현된 기능이 이후 턴에서 손실되는 현상. 이러한 문제를 해결하기 위해, 우리는 내부 비평 및 개선 루프를 통해 다이어그램을 개선하는 다중 에이전트 시스템인 PaperBanana-Interact를 도입한다. PaperBanana-Interact는 턴이 지남에 따라 다이어그램 품질을 떨어뜨리는 대신 지속적으로 개선하며, 품질 점수에서 베이스라인 대비 11.9–18.6점 우수하고 망각을 3.7–6.2점 감소시킨다.
English
Recent efforts have aimed to automate scientific diagram generation from paper content (Lin et al., 2026; Zhu et al., 2026a). However, fully satisfying an author's visual and communicative preferences in a single turn is challenging: in our formative user study (N = 14), all participants requested further revisions after viewing an initial draft, and 86% of them rated the refined diagrams as more satisfactory. Despite the clear demand, the multi-turn workflow remains largely underexplored. To bridge this gap, we present MTPaperBananaBench, a benchmark for multi-turn diagram generation containing 292 images annotated with 3,518 user requirements. To reduce expensive human studies and enable scalable benchmarking, we construct a user simulator that, at each turn, identifies unsatisfied requirements and converts k of them into natural language feedback. Evaluating both requirement satisfaction and overall diagram quality reveals two key failure modes shared across baseline multiturn systems: (1) quality drift, where diagram quality progressively declines over turns, and (2) forgetting, where previously implemented features are lost in subsequent turns. To address these issues, we introduce PaperBanana-Interact, a multi-agent system that refines diagrams via an internal critique-and-refine loop. PaperBanana-Interact consistently improves rather than degrades diagram quality across turns, outperforming baselines by 11.9-18.6 points in quality score and reducing forgetting by 3.7-6.2 points.