Towards Lexical Encoding of Multi-Word Expressions in Spanish Dialects

Diana Bogantes, Eric Rodríguez, Alejandro Arauco, Alejandro Rodríguez, Agata Savary


Abstract
This paper describes a pilot study in lexical encoding of multi-word expressions (MWEs) in 4 Latin American dialects of Spanish: Costa Rican, Colombian, Mexican and Peruvian. We describe the variability of MWE usage across dialects. We adapt an existing data model to a dialect-aware encoding, so as to represent dialect-related specificities, while avoiding redundancy of the data common for all dialects. A dozen of linguistic properties of MWEs can be expressed in this model, both on the level of a whole MWE and of its individual components. We describe the resulting lexical resource containing several dozens of MWEs in four dialects and we propose a method for constructing a web corpus as a support for crowdsourcing examples of MWE occurrences. The resource is available under an open license and paves the way towards a large-scale dialect-aware language resource construction, which should prove useful in both traditional and novel NLP applications.
Anthology ID:
L16-1358
Volume:
Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16)
Month:
May
Year:
2016
Address:
Portorož, Slovenia
Editors:
Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Sara Goggi, Marko Grobelnik, Bente Maegaard, Joseph Mariani, Helene Mazo, Asuncion Moreno, Jan Odijk, Stelios Piperidis
Venue:
LREC
SIG:
Publisher:
European Language Resources Association (ELRA)
Note:
Pages:
2255–2261
Language:
URL:
https://aclanthology.org/L16-1358
DOI:
Bibkey:
Cite (ACL):
Diana Bogantes, Eric Rodríguez, Alejandro Arauco, Alejandro Rodríguez, and Agata Savary. 2016. Towards Lexical Encoding of Multi-Word Expressions in Spanish Dialects. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16), pages 2255–2261, Portorož, Slovenia. European Language Resources Association (ELRA).
Cite (Informal):
Towards Lexical Encoding of Multi-Word Expressions in Spanish Dialects (Bogantes et al., LREC 2016)
Copy Citation:
PDF:
https://aclanthology.org/L16-1358.pdf