Par-PARSEME - a multilingual shared task on paraphrasing of idioms

Event Notification Type: 
Call for Participation
Abbreviated Title: 
Par-PARSEME
Location: 
SemEval 2027
State: 
Country: 
City: 
Contact: 
Carlos Ramisch
Agata Savary
Takuya Nakamura
Submission Deadline: 
Saturday, 5 September 2026

Overview
---------------
The PARSEME and UniDive communities are happy to introduce the Par-PARSEME, task 4 in SemEval-2027. It is a follow-up of previous shared tasks organized by PARSEME.

Par-PARSEME is a shared task whose goal is to assess systems’ ability to **paraphrase idioms** in written texts. Idioms, such as a hot dog or to pull one’s leg, are subsets of multiword expressions (MWEs), and are combinations of words exhibiting non-compositional semantics. This makes them challenging for semantic tasks like paraphrasing. For instance, a hot dog can be paraphrased as a bread roll with a hot sausage but not as #a warm dog or #dog which is hot.

The input for this task is a raw sentence containing exactly one idiom not explicitly marked in text. Systems must paraphrase the sentence so that the original idiom no longer occurs (which requires idiom identification), but the meaning is kept. For all details, see the Par-PARSEME website.

Evaluation is done along 2 axes: **performance** (automatic and manual measures) and **diversity** (automatic measures).

Quick links
---------------
Par-PARSEME website: https://unidive.lisn.upsaclay.fr/doku.php?id=other-events:par-parseme
SemEval 2027 webpage: https://semeval.github.io/SemEval2027/tasks
Par-PARSEME data repository: https://gitlab.com/parseme/sharedtask-data/-/tree/master/par-parseme?ref...
Mailing list for the shared task participants: https://groups.google.com/g/par-parseme-semeval-2025
Mailing list to contact the shared task organizers: parseme dash core at lisn DOT upsaclay DOT fr
Expression of interest form: https://forms.gle/yUbc2FWhPUU41Wma7
Baseline system: https://github.com/racai-ai/mwe_baseline

Languages
---------------
Thirteen languages are covered in the training data: French, Georgian, Modern Greek, Hebrew, Japanese, Latvian, Persian, Polish, Brazilian Portuguese, Romanian, Slovene, Swedish, and Ukrainian. The test data will cover the same languages, and possibly also English and Serbian (to be confirmed).

Data
---------------
The Par-PARSEME Gitlab repository contains:
training data with 60 to 150 data points per language
tiny trial data in English and French to illustrate the expected outcomes and file formats
tools for system evaluation and file format validation
baseline results obtained by the baseline system which we make available to the participants
Test data will contain about 100 sentences per language.

Important dates
---------------
* 8 September 2026: Publication of the training data
* early January 2027: Evaluation starts
* January 31, 2027: Evaluation ends
* February 2027 (tentative): Submission of system description papers
* March 2027 (tentative): Notification to authors
* April 2027 (tentative): Camera-ready due
* Summer 2027: SemEval-2027 workshop, co-located with a major NLP conference

Organizers
---------------
Carlos Ramisch, Aix-Marseille Université, LIS, France
Agata Savary, Université Paris-Saclay, LISN, France
Takuya Nakamura, Université Paris-Saclay, LISN, France
Éric Bilinski, Université Paris-Saclay, LISN, France
Manon Scholivet, Université Paris-Saclay, LISN, France

Advisory board
---------------
20 PARSEME native experts of the 13 participating languages