ChatPaper.aiChatPaper

マルチモーダルLLMによる計算ユーモア:手法、データセット、評価、課題

Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges

July 21, 2026
著者: Tuo Liang, Zhe Hu, Disheng Liu, Jing Li, Yu Yin
cs.AI

要旨

マルチモーダルなユーモア(ミーム、風刺漫画、コミック等)は、意図された意味が文字通りの情景記述ではなく、非字義的なメカニズム、共有された文化的知識、および伝達意図に依存するため、AIシステムにとって依然として困難である。本サーベイでは、単一画像および複数パネルからなる作品における視覚的ユーモア理解に焦点を当てつつ、ユーモア生成を新たな下流フロンティアとして位置づける。我々は、既存のユーモア、皮肉、および一般的なMLLMサーベイに対して本文献を位置づけ、認識・解釈と推論・生成を網羅する能力中心の階層構造を用いて整理する。この観点に基づき、ベンチマーク設計、評価プロトコル、モデリングパラダイムを統合し、タスク特化型融合モデルからマルチモーダルアライメント、エビデンスに基づく推論、制御可能生成を活用した大規模モデルアプローチへの分野の移行を追跡する。最後に、進歩を阻む主要な障壁として、ショートカットに依存しやすい評価、限定的な文化的・物語的カバレッジ、弱いエビデンスの根拠付け、および未解決の安全性・所有権に関する懸念を強調する。
English
Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communicative intent rather than literal scene description. This survey focuses on visual humor understanding in single-image and multi-panel artifacts, while treating humor generation as an emerging downstream frontier. We position the literature against prior humor, sarcasm, and general MLLM surveys and organize it using a capability-centric hierarchy spanning recognition, interpretation and reasoning, and generation. Under this lens, we synthesize benchmark design, evaluation protocols, and modeling paradigms, tracing the field's shift from task-specific fusion models to large-model approaches based on multimodal alignment, evidence-grounded reasoning, and controlled generation. We conclude by highlighting the main barriers to progress: shortcut-prone evaluation, limited cultural and narrative coverage, weak evidence grounding, and unresolved safety and ownership concerns.