Findings of the WMT 2018 Shared Task on Parallel Corpus Filtering

Philipp Koehn, Huda Khayrallah, Kenneth Heafield, Mikel L. Forcada


Abstract
We posed the shared task of assigning sentence-level quality scores for a very noisy corpus of sentence pairs crawled from the web, with the goal of sub-selecting 1% and 10% of high-quality data to be used to train machine translation systems. Seventeen participants from companies, national research labs, and universities participated in this task.
Anthology ID:
W18-6453
Volume:
Proceedings of the Third Conference on Machine Translation: Shared Task Papers
Month:
October
Year:
2018
Address:
Belgium, Brussels
Editors:
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana Neves, Matt Post, Lucia Specia, Marco Turchi, Karin Verspoor
Venue:
WMT
SIG:
SIGMT
Publisher:
Association for Computational Linguistics
Note:
Pages:
726–739
Language:
URL:
https://aclanthology.org/W18-6453
DOI:
10.18653/v1/W18-6453
Bibkey:
Cite (ACL):
Philipp Koehn, Huda Khayrallah, Kenneth Heafield, and Mikel L. Forcada. 2018. Findings of the WMT 2018 Shared Task on Parallel Corpus Filtering. In Proceedings of the Third Conference on Machine Translation: Shared Task Papers, pages 726–739, Belgium, Brussels. Association for Computational Linguistics.
Cite (Informal):
Findings of the WMT 2018 Shared Task on Parallel Corpus Filtering (Koehn et al., WMT 2018)
Copy Citation:
PDF:
https://aclanthology.org/W18-6453.pdf