ChatPaper.aiChatPaper

가르칠 수 있는 것을 넘어서는 탐색: 에이전틱 시각 생성에서 지식 경계의 진화

Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation

July 9, 2026
저자: Haozhe Wang, Weijia Feng, Jinpeng Yu, Che Liu, Ping Nie, Fangzhen Lin, Jiaming Liu, Ruihua Huang, Jimmy Lin, Wenhu Chen, Cong Wei
cs.AI

초록

시각 생성 모델은 렌더링에 탁월하지만, 자신이 알지 못하는 내용은 자신 있게 지어낸다. 사용자 요청은 무한하고, 진화하며, 깊이 있는 롱테일(long-tail)을 형성한다: 새로운 캐릭터, 유행하는 개체, 데이터 수집 이후의 사건 등이 포함된다. 이러한 세계 지식 병목 현상은 구조적이다. 생성 모델은 고정된 코퍼스로 학습되지만, 시각 세계는 개방형이기 때문이다. 우리는 12가지 실패 범주와 22개 도메인에 걸친 20,839개의 프롬프트로 구성된 SearchGen-20K와 SearchGen-Bench를 구축하고, 사전 실행된 다중 모달 SearchGen-Corpus-1M을 함께 제공하여 오프라인 및 재현 가능한 연구를 지원한다. SearchGen-Bench에서 최첨단 오픈 생성 모델은 100점 만점에 21~28점에 불과하며, 이는 기존 벤치마크에서는 드러나지 않는 40포인트 하락이다. 자연스러운 해결책은 검색 도구를 활용하여 에이전트 기반 시각 생성을 가능하게 하는 것이다. 그러나 우리는 단순한 검색이 실패한다는 사실을 발견했다. 검색이 무분별하게 결과를 가져와 생성 모델이 이미 처리할 수 있는 프롬프트에 잡음을 주입하기 때문이다. 우리는 이 근본 원인을 생성 모델 특화적이며 진화하는 지식 경계, 즉 생성 모델이 학습을 통해 내재화할 수 있는 것과 외부 맥락에 남아 있어야 하는 것 사이의 경계에서 찾는다. 이 경계는 사전에 명시하기 어렵지만, '가르친 후 탐색하는(teach-then-search)' 공동 학습 프레임워크를 통해 발견 가능함을 보여준다. 이 공동 학습 방법의 최소 버전조차도 단조적 개선을 만들어내며, 세계 지식에 기반한 요청을 충족할 수 있는 시각 생성의 재귀적 자기 개선을 위한 기반을 마련한다. 우리는 전체 데이터셋, 공동 학습 코퍼스, 검색 코퍼스를 도구 기반의 세계 지식 중심 시각 생성을 위한 재현 가능한 프레임워크로 공개한다.
English
Visual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbounded, evolving, and deeply long-tailed: new characters, trending entities, post-cutoff events, and more. This world-knowledge bottleneck is structural: generators are trained on fixed corpora, but the visual world is open-ended. We construct SearchGen-20K and SearchGen-Bench, with 20,839 prompts spanning twelve failure categories and twenty-two domains, paired with a pre-executed multimodal SearchGen-Corpus-1M to support offline, reproducible research. On SearchGen-Bench, frontier open generators score only 21 to 28 out of 100, a 40-point collapse invisible to existing benchmarks. The natural remedy is to employ search tools, enabling agentic visual generation. However, we find that naive search fails: it retrieves indiscriminately, injecting noise into prompts the generator already handles. We trace the root cause to a generator-specific, evolving knowledge boundary: the divide between what a generator can internalize through training and what must remain in external context. Although this boundary is hard to specify in advance, we show that it is discoverable through a teach-then-search co-training framework. Even a minimal version of this co-training recipe produces monotonic improvement, laying the foundation for recursive self-improvement in visual generation that can meet world-knowledge-grounded requests. We release the full dataset, co-training corpus, and search corpus as a replayable harness for tool-augmented, world-knowledge-grounded visual generation.