We are pleased to invite participation in MMCultureQA, a SemEval-2027 Shared Task on Multilingual and Multimodal Culturally Grounded Question Answering.
What is MMCultureQA?
The goal of this shared task is to evaluate whether multimodal systems can answer questions that require not only understanding an image, but also linguistic and cultural knowledge.
Given an image and a spoken or written question, systems must generate a short, open-ended textual answer. Questions cover culturally grounded topics such as food, places, customs, everyday objects, traditions, and social practices, where visual information alone may not be sufficient.
The shared task includes two main tasks:
- Task 1: Spoken Visual Question Answering
Input: Image + spoken question
Output: Open-ended textual answer - Task 2: Textual Visual Question Answering
Input: Image + written question
Output: Open-ended textual answer
Each task includes multiple language tracks, and participants may take part in one or multiple tracks.
Languages include: English, Modern Standard Arabic, Dialectal Arabic (Egyptian, Leventine), Bangla, Hindi, Amharic, Assamese, Gujarati, Italian, Marathi, Oromo, Somali, Sinhala, Nepali, Tigrinya, Urdu.
Dataset
The MMCultureQA data builds on the EverydayMMQA framework, which enables the development of culturally grounded multimodal QA datasets across image, text, and speech modalities.
Dataset and training data:
https://huggingface.co/datasets/QCRI/MMCQA-SemEval27
Important Dates
- Training data release: 23 September 2026
- Evaluation starts: 10 January 2027
- Evaluation ends: 31 January 2027
- System description papers: February 2027 (tentative)
- Notification to authors: March 2027 (tentative)
- Camera-ready papers: April 2027 (tentative)
- SemEval-2027 Workshop: Summer 2027
How to Participate
Register your team:
https://docs.google.com/forms/d/15-Vbs82rBF6TCwxT4aAiJDWiiDqPI6uaggz95qStSkQ/viewform
Join our Slack community for announcements, questions, and discussions:
https://join.slack.com/t/mm-eval/shared_invite/zt-4a869o7xc-nP8fMfETqSczgpNDr08fJQ
Task website:
https://mmcultureqa-semeval27.github.io/
Dataset:
https://huggingface.co/datasets/QCRI/MMCQA-SemEval27
We welcome participation from researchers and teams working on multimodal LLMs, vision-language models, AudioLLMs, multilingual NLP, speech processing, visual question answering, under-resourced languages, and culturally grounded AI.
Please feel free to share this call with colleagues, students, and research groups who may be interested.
We look forward to your participation in MMCultureQA at SemEval-2027!
- The MMCultureQA Organizing Team