Segmentation for Domain Adaptation in Arabic

Mohammed Attia, Ali Elkahky


Abstract
Segmentation serves as an integral part in many NLP applications including Machine Translation, Parsing, and Information Retrieval. When a model trained on the standard language is applied to dialects, the accuracy drops dramatically. However, there are more lexical items shared by the standard language and dialects than can be found by mere surface word matching. This shared lexicon is obscured by a lot of cliticization, gemination, and character repetition. In this paper, we prove that segmentation and base normalization of dialects can help in domain adaptation by reducing data sparseness. Segmentation will improve a system performance by reducing the number of OOVs, help isolate the differences and allow better utilization of the commonalities. We show that adding a small amount of dialectal segmentation training data reduced OOVs by 5% and remarkably improves POS tagging for dialects by 7.37% f-score, even though no dialect-specific POS training data is included.
Anthology ID:
W19-4613
Volume:
Proceedings of the Fourth Arabic Natural Language Processing Workshop
Month:
August
Year:
2019
Address:
Florence, Italy
Editors:
Wassim El-Hajj, Lamia Hadrich Belguith, Fethi Bougares, Walid Magdy, Imed Zitouni, Nadi Tomeh, Mahmoud El-Haj, Wajdi Zaghouani
Venue:
WANLP
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
119–129
Language:
URL:
https://aclanthology.org/W19-4613
DOI:
10.18653/v1/W19-4613
Bibkey:
Cite (ACL):
Mohammed Attia and Ali Elkahky. 2019. Segmentation for Domain Adaptation in Arabic. In Proceedings of the Fourth Arabic Natural Language Processing Workshop, pages 119–129, Florence, Italy. Association for Computational Linguistics.
Cite (Informal):
Segmentation for Domain Adaptation in Arabic (Attia & Elkahky, WANLP 2019)
Copy Citation:
PDF:
https://aclanthology.org/W19-4613.pdf