Adaptive Dictionary for Bilingual Lexicon Extraction from Comparable Corpora

Amir Hazem, Emmanuel Morin


Abstract
One of the main resources used for the task of bilingual lexicon extraction from comparable corpora is : the bilingual dictionary, which is considered as a bridge between two languages. However, no particular attention has been given to this lexicon, except its coverage, and the fact that it can be issued from the general language, the specialised one, or a mix of both. In this paper, we want to highlight the idea that a better consideration of the bilingual dictionary by studying its entries and filtering the non-useful ones, leads to a better lexicon extraction and thus, reach a higher precision. The experiments are conducted on a medical domain corpora. The French-English specialised corpus 'breast cancer' of 1 million words. We show that the empirical results obtained with our filtering process improve the standard approach traditionally dedicated to this task and are promising for future work.
Anthology ID:
L12-1462
Volume:
Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12)
Month:
May
Year:
2012
Address:
Istanbul, Turkey
Editors:
Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Mehmet Uğur Doğan, Bente Maegaard, Joseph Mariani, Asuncion Moreno, Jan Odijk, Stelios Piperidis
Venue:
LREC
SIG:
Publisher:
European Language Resources Association (ELRA)
Note:
Pages:
288–292
Language:
URL:
http://www.lrec-conf.org/proceedings/lrec2012/pdf/784_Paper.pdf
DOI:
Bibkey:
Cite (ACL):
Amir Hazem and Emmanuel Morin. 2012. Adaptive Dictionary for Bilingual Lexicon Extraction from Comparable Corpora. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12), pages 288–292, Istanbul, Turkey. European Language Resources Association (ELRA).
Cite (Informal):
Adaptive Dictionary for Bilingual Lexicon Extraction from Comparable Corpora (Hazem & Morin, LREC 2012)
Copy Citation:
PDF:
http://www.lrec-conf.org/proceedings/lrec2012/pdf/784_Paper.pdf