Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges

Tuo Liang, Zhe Hu, Disheng Liu, Jing Li, Yu Yin

arXiv:2607.19011 · 2026-07-26 공개 · arXiv · PDF

multimodal-llms benchmark-design controlled-generation evaluation-protocols multimodal-alignment sarcasm-detection visual-humor humor-generation

Abstract

Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communicative intent rather than literal scene description. This survey focuses on visual humor understanding in single-image and multi-panel artifacts, while treating humor generation as an emerging downstream frontier. We position the literature against prior humor, sarcasm, and general MLLM surveys and organize it using a capability-centric hierarchy spanning recognition, interpretation and reasoning, and generation. Under this lens, we synthesize benchmark design, evaluation protocols, and modeling paradigms, tracing the field's shift from task-specific fusion models to large-model approaches based on multimodal alignment, evidence-grounded reasoning, and controlled generation. We conclude by highlighting the main barriers to progress: shortcut-prone evaluation, limited cultural and narrative coverage, weak evidence grounding, and unresolved safety and ownership concerns.

한국어 요약

한 줄 요약

다중모달 풍자 이해는 AI가 비직관적 의미와 문화적 맥락을 해석하는 능력을 요구하며, 이를 위한 모델 평가와 데이터셋 설계가 주요 과제이다.

핵심 기여도

핵심 아이디어

기존 AI 모델은 텍스트와 이미지를 결합해 처리하는 능력이 향상되었으나, **비직관적 의미**(non-literal meaning)를 해석하는 데에는 여전히 한계가 있다. 예를 들어, **meme**나 **cartoon**에서 나타나는 **humor**, **satire**, **irony**는 단순한 시각 인식이나 텍스트 분석으로는 해석할 수 없으며, **문화적 지식**, **의도**(communicative intent), **상황적 맥락**을 종합적으로 고려해야 한다. 이에 따라, 본 논문은 **multimodal humor understanding**을 **인식**, **해석 및 추론**, **생성**의 세 단계로 구분하고, 각 단계별로 필요한 모델링 패러다임과 평가 프로토콜을 정리한다. 특히, **GPT-4o**는 **NYCC**와 **MangaUB**에서 시각 인식과 캡션 매칭 능력이 뛰어난 반면, **Qwen3.5-27B**는 **YesBut-v2**와 **HumorDB**에서 도덕적 추론과 일반 풍자 감지에서 우수한 성능을 보였다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 **multimodal humor understanding**을 체계적으로 분류하고, **인식**, **해석**, **생성**의 세 단계로 구분함으로써 AI 모델의 능력 평가를 구조화했다는 점에서 학술적 의의가 있다. 또한, **NYCC**, **MemeQA**, **DarkHumor** 등 다양한 데이터셋을 종합적으로 분석하여 풍자 이해의 복잡성을 드러냈다. 그러나 **shortcut-prone evaluation**, **weak evidence grounding**, **cultural and narrative coverage** 등의 문제는 여전히 해결되지 않았으며, **safety and ownership** 문제는 향후 연구에서 반드시 고려해야 할 사항이다.

실용적 활용

본 연구는 **콘텐츠 모니터링**, **소셜 미디어 분석**, **크리에이티브 도구 개발** 등에서 AI 모델의 풍자 이해 능력을 평가하고 개선하는 데 활용될 수 있다. 특히, **MemeQA**나 **HumorDB**와 같은 데이터셋은 모델이 비직관적 의미를 해석하는 능력을 테스트하는 데 유용하며, **GPT-4o**나 **Qwen3.5-27B** 같은 모델은 실제 애플리케이션에서 즉시 활용 가능한 성능을 보인다.