멀티모달 LLM을 이용한 계산 유머: 방법론, 데이터셋, 평가, 및 도전 과제
Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges
July 21, 2026
저자: Tuo Liang, Zhe Hu, Disheng Liu, Jing Li, Yu Yin
cs.AI
초록
밈, 만화, 카툰에 나타난 멀티모달 유머는 의도된 의미가 문자 그대로의 장면 묘사가 아닌 비문자적 메커니즘, 공유된 문화적 지식, 그리고 의사소통 의도에 의존하기 때문에 AI 시스템이 이해하기 어려운 과제로 남아 있다. 본 설문조사는 단일 이미지 및 다중 패널 아티팩트에서의 시각적 유머 이해에 초점을 맞추되, 유머 생성을 새로운 하위 영역으로 간주한다. 우리는 기존의 유머, 풍자, 일반 MLLM 관련 조사 논문들과의 차별점을 제시하고, 인식, 해석 및 추론, 생성으로 이어지는 능력 중심의 계층 구조를 활용하여 문헌을 체계화한다. 이러한 관점에서 벤치마크 설계, 평가 프로토콜, 모델링 패러다임을 종합하며, 해당 분야가 과제 특화 융합 모델에서 멀티모달 정렬, 증거 기반 추론, 제어된 생성에 기반한 대규모 모델 접근법으로 전환되어 온 과정을 추적한다. 결론적으로, 진전을 가로막는 주요 장벽으로 지름길 평가, 제한된 문화적 및 내러티브 범위, 취약한 증거 기반, 그리고 해결되지 않은 안전 및 소유권 문제를 강조한다.
English
Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communicative intent rather than literal scene description. This survey focuses on visual humor understanding in single-image and multi-panel artifacts, while treating humor generation as an emerging downstream frontier. We position the literature against prior humor, sarcasm, and general MLLM surveys and organize it using a capability-centric hierarchy spanning recognition, interpretation and reasoning, and generation. Under this lens, we synthesize benchmark design, evaluation protocols, and modeling paradigms, tracing the field's shift from task-specific fusion models to large-model approaches based on multimodal alignment, evidence-grounded reasoning, and controlled generation. We conclude by highlighting the main barriers to progress: shortcut-prone evaluation, limited cultural and narrative coverage, weak evidence grounding, and unresolved safety and ownership concerns.