Computational Humor with Multimodal LLMs: Methods & Challenges

A survey on computational humor using multimodal LLMs: methods, datasets, evaluation, and challenges in understanding memes, cartoons, and comics.

jueves, 23 de julio de 2026 • 4 min read • Q2BSTUDIO Team

Cómo la IA entiende el humor visual en memes y cómics

Computational humor represents one of the most complex and fascinating frontiers of modern artificial intelligence. Understanding jokes, memes, and cartoons is not limited to processing pixels or words; it requires capturing non-literal mechanisms, shared cultural knowledge, and the communicator's intent. Multimodal large language models (MLLMs) have opened new possibilities for tackling this task, combining vision, language, and reasoning. In this article we explore current methods for visual humor recognition, interpretation, and generation, the main challenges faced by the research community, and how businesses can leverage these technologies through specialized custom software development.

The field has evolved from early task-specific fusion models to architectures based on large pre-trained models. In the recognition phase, systems must identify whether an image or sequence contains humor, distinguishing mechanisms such as incongruity, exaggeration, or satire. Aligned multimodal embeddings —for instance, CLIP— are used to compare visual and textual representations. Interpretation goes a step further: it requires explaining why something is funny, identifying the conflict between what is expected and what is shown, as well as implicit cultural references. Here, evidence-grounded reasoning techniques come into play, where the model builds logical chains linking visual elements with external knowledge. Finally, humor generation —still an emerging frontier— aims to create original comedic content, either by modifying images or generating dialogues with double meanings. Current approaches rely on controlled generation through prompt engineering and multimodal alignment supervised by human feedback.

Despite progress, significant barriers persist. Evaluating humorous systems is especially prone to shortcuts: models can learn spurious correlations (e.g., associating certain colors or keywords with humor) without truly understanding the joke. This leads to inflated metrics that do not reflect real generalization ability. Furthermore, cultural and narrative coverage is limited; most datasets come from Anglophone and Western contexts, ignoring humorous traditions from other regions. The reliance on tacit knowledge makes it difficult for a model trained on American memes to grasp a Japanese cartoon or an Andalusian double entendre. On the other hand, the lack of solid grounding in verifiable evidence causes systems to generate offensive or inappropriate humor when output is not controlled. Issues such as safety and intellectual property are critical: humor generated by AI can infringe copyrights or perpetuate harmful stereotypes if not properly managed.

For companies that wish to integrate computational humor capabilities into their products —from virtual assistants to marketing platforms— the key lies in developing customized and robust solutions. At Q2BSTUDIO, as experts in AI and software, we offer services that address these challenges practically. For example, building humorous content classification systems requires finely tuned multimodal language models with local data and safety mechanisms. Our experience in custom software allows us to design data pipelines, train models, and deploy them on scalable cloud infrastructures, whether on AWS or Azure. The cloud facilitates storage and processing of large volumes of images and texts, as well as orchestration of AI agents that combine humorous reasoning with other capabilities. Additionally, cybersecurity is essential to protect training data and generated interactions; we perform security audits and pentesting to ensure the system cannot be manipulated to produce harmful content. We also implement Business Intelligence solutions with Power BI to analyze the performance of these models in real time, measuring acceptance metrics and cultural bias.

One of the most promising developments is the incorporation of specialized AI agents for humor. These agents not only recognize jokes but can adapt their tone according to the user's profile, generating contextual jokes in customer service assistants or educational applications. For example, a virtual language tutor could use humor to make learning more enjoyable, always within safe boundaries. Implementing such agents requires a modular architecture integrating language models, cultural rule engines, and feedback systems. At Q2BSTUDIO we design these architectures using microservices, vector databases, and multimodal model APIs, all orchestrated in hybrid cloud environments. Process automation also plays a relevant role: through automated workflows we can periodically update models with new humor sources, ensuring the system evolves with cultural trends.

The future of computational humor lies in overcoming current limitations with more robust approaches. Researchers are exploring contrastive multimodal learning methods to reduce shortcuts, as well as multilingual and multicultural datasets. Controlled generation through decoding techniques with safety constraints will improve output reliability. In the business sphere, companies that invest in this type of technology will be able to differentiate themselves by offering more human and engaging user experiences. At Q2BSTUDIO we help organizations navigate this path, combining our expertise in AI, cloud, cybersecurity, and business intelligence to build ethical, effective, and scalable computational humor solutions. If your company seeks to integrate these capabilities or explore new applications of multimodal AI, feel free to contact us. Humor is, after all, one of the most human expressions; getting a machine to understand and generate it responsibly is a challenge worth tackling.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.