ChatPaper.aiChatPaper

Omni-Persona: Systematische Benchmarking en Verbetering van Omnimodale Personalisatie

Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization

May 11, 2026
Auteurs: Yeongtak Oh, Dongwook Lee, Sangkwon Park, Heeseung Kim, Sungroh Yoon
cs.AI

Samenvatting

Hoewel multimodale grote taalmodellen vooruitgang hebben geboekt op het gebied van tekst, beeld en audio, is personalisatieonderzoek voornamelijk visueel-talig gebleven, waarbij unified omnimodale benchmarking die gezamenlijk tekst, beeld en audio bestrijkt nog beperkt is en de methodologische strengheid mist om rekening te houden met afwezige-persona-scenario's of systematische grondingstudies. We introduceren Omni-Persona, de eerste uitgebreide benchmark voor omnimodale personalisatie. We formaliseren de taak als cross-modale routering over de Persona Modality Graph, die 4 taakgroepen en 18 fijnmazige taken omvat over circa 750 items. Om het grondgedrag rigoureus te diagnosticeren, stellen we Calibrated Accuracy (mathrm{Cal}) voor, dat gezamenlijk correcte gronding en passende onthouding beloont en afwezige-persona-query's integreert binnen een uniform evaluatiekader. In onze specifieke experimenten komen drie diagnostische bevindingen naar voren: (i) open-source modellen vertonen een consistente audio-versus-visuele grondingskloof die RLVR gedeeltelijk verkleint door middel van dichte regelgebaseerde supervisie; (ii) beantwoordbare recall en parameterschaal zijn onvolledige diagnostische middelen, aangezien sterke recall kan samengaan met afwezige-persona-hallucinatie en grotere modellen niet altijd een hogere Cal bereiken, wat kalibratie blootlegt als een aparte evaluatie-as; en (iii) SFT wordt begrensd door de moeilijkheid om op schaal geannoteerde ground-truth supervisie te construeren, terwijl RLVR consistenter generaliseert door middel van outcome-level verifieerbare feedback, maar onder ons beloningsontwerp afglijdt naar conservatief gedrag en lagere generatiekwaliteit. Omni-Persona dient dus als een diagnostisch raamwerk dat de valkuilen van omnimodale personalisatie blootlegt en toekomstige post-training en beloningsontwerp begeleidt.
English
While multimodal large language models have advanced across text, image, and audio, personalization research has remained primarily vision-language, with unified omnimodal benchmarking that jointly covers text, image, and audio still limited, and lacking the methodological rigor to account for absent-persona scenarios or systematic grounding studies. We introduce Omni-Persona, the first comprehensive benchmark for omnimodal personalization. We formalize the task as cross-modal routing over the Persona Modality Graph, encompassing 4 task groups and 18 fine-grained tasks across {sim}750 items. To rigorously diagnose grounding behavior, we propose Calibrated Accuracy (mathrm{Cal)}, which jointly rewards correct grounding and appropriate abstention, incorporating absent-persona queries within a unified evaluation framework. On our dedicated experiments, three diagnostic findings emerge: (i) open-source models show a consistent audio-vs-visual grounding gap that RLVR partially narrows via dense rule-based supervision; (ii) answerable recall and parameter scale are incomplete diagnostics, since strong recall can coexist with absent-persona hallucination and larger models do not always achieve higher Cal, exposing calibration as a separate evaluation axis; and (iii) SFT is bounded by the difficulty of constructing annotated ground-truth supervision at scale, while RLVR generalizes more consistently through outcome-level verifiable feedback yet drifts toward conservative behavior and lower generation quality under our reward design. Omni-Persona thus serves as a diagnostic framework that surfaces the pitfalls of omnimodal personalization, guiding future post-training and reward design.