Boosting Neural Machine Translation with Similar Translations

Jitao Xu, Josep Crego, Jean Senellart


Abstract
This paper explores data augmentation methods for training Neural Machine Translation to make use of similar translations, in a comparable way a human translator employs fuzzy matches. In particular, we show how we can simply present the neural model with information of both source and target sides of the fuzzy matches, we also extend the similarity to include semantically related translations retrieved using sentence distributed representations. We show that translations based on fuzzy matching provide the model with “copy” information while translations based on embedding similarities tend to extend the translation “context”. Results indicate that the effect from both similar sentences are adding up to further boost accuracy, combine naturally with model fine-tuning and are providing dynamic adaptation for unseen translation pairs. Tests on multiple data sets and domains show consistent accuracy improvements. To foster research around these techniques, we also release an Open-Source toolkit with efficient and flexible fuzzy-match implementation.
Anthology ID:
2020.acl-main.144
Volume:
Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics
Month:
July
Year:
2020
Address:
Online
Editors:
Dan Jurafsky, Joyce Chai, Natalie Schluter, Joel Tetreault
Venue:
ACL
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
1580–1590
Language:
URL:
https://aclanthology.org/2020.acl-main.144
DOI:
10.18653/v1/2020.acl-main.144
Bibkey:
Cite (ACL):
Jitao Xu, Josep Crego, and Jean Senellart. 2020. Boosting Neural Machine Translation with Similar Translations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1580–1590, Online. Association for Computational Linguistics.
Cite (Informal):
Boosting Neural Machine Translation with Similar Translations (Xu et al., ACL 2020)
Copy Citation:
PDF:
https://aclanthology.org/2020.acl-main.144.pdf
Video:
 http://slideslive.com/38929113