Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges
Abstract
Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communicative intent rather than literal scene description. This survey focuses on visual humor understanding in single-image and multi-panel artifacts, while treating humor generation as an emerging downstream frontier. We position the literature against prior humor, sarcasm, and general MLLM surveys and organize it using a capability-centric hierarchy spanning recognition, interpretation and reasoning, and generation. Under this lens, we synthesize benchmark design, evaluation protocols, and modeling paradigms, tracing the field's shift from task-specific fusion models to large-model approaches based on multimodal alignment, evidence-grounded reasoning, and controlled generation. We conclude by highlighting the main barriers to progress: shortcut-prone evaluation, limited cultural and narrative coverage, weak evidence grounding, and unresolved safety and ownership concerns.
Community
This survey examines visual humor understanding across memes, cartoons, and comics through a capability-centric hierarchy comprising recognition, interpretation/reasoning, and generation. This taxonomy reorganizes a fragmented literature by focusing on the capabilities that benchmarks and models actually evaluate. Beyond synthesizing prior work, the authors conduct a cross-benchmark evaluation of recent MLLMs, showing that while current models perform reasonably well on visual recognition, they still lag far behind humans on interpretation-intensive tasks. The results also reveal that model rankings vary substantially across humor capabilities and that explicit reasoning variants do not consistently improve performance.
Get this paper in your agent:
hf papers read 2607.19011 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper