基於多模態大型語言模型的計算幽默:方法、資料集、評估與挑戰
Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges
July 21, 2026
作者: Tuo Liang, Zhe Hu, Disheng Liu, Jing Li, Yu Yin
cs.AI
摘要
迷因、卡通與漫畫中的多模態幽默對人工智能系統而言仍具挑戰性,因為預期意義依賴於非字面機制、共享文化知識及溝通意圖,而非單純的場景字面描述。本調查聚焦於單圖及多格作品中的視覺幽默理解,同時將幽默生成視為新興的前沿應用領域。我們有別於先前的幽默、諷刺及通用多模態大語言模型(MLLM)綜述,並以能力為核心的分層架構來組織文獻,涵蓋辨識、解讀與推理,以及生成等層面。在此視角下,我們綜合分析基準設計、評估協議與建模範式,追蹤該領域從任務專用融合模型,朝基於多模態對齊、證據基礎推理及受控生成的大型模型方法的轉變。最後,我們指出阻礙進展的主要障礙:易受捷徑影響的評估、有限的文化與敘事覆蓋範圍、薄弱的證據基礎,以及未解決的安全與版權疑慮。
English
Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communicative intent rather than literal scene description. This survey focuses on visual humor understanding in single-image and multi-panel artifacts, while treating humor generation as an emerging downstream frontier. We position the literature against prior humor, sarcasm, and general MLLM surveys and organize it using a capability-centric hierarchy spanning recognition, interpretation and reasoning, and generation. Under this lens, we synthesize benchmark design, evaluation protocols, and modeling paradigms, tracing the field's shift from task-specific fusion models to large-model approaches based on multimodal alignment, evidence-grounded reasoning, and controlled generation. We conclude by highlighting the main barriers to progress: shortcut-prone evaluation, limited cultural and narrative coverage, weak evidence grounding, and unresolved safety and ownership concerns.