The Romanian Neuter Examined Through A Two-Gender N-Gram Classification System

Liviu P. Dinu, Vlad Niculae, Octavia-Maria Şulea


Abstract
Romanian has been traditionally seen as bearing three lexical genders: masculine, feminine and neuter, although it has always been known to have only two agreement patterns (for masculine and feminine). A recent analysis of the Romanian gender system described in (Bateman and Polinsky, 2010), based on older observations, argues that there are two lexically unspecified noun classes in the singular and two different ones in the plural and that what is generally called neuter in Romanian shares the class in the singular with masculines, and the class in the plural with feminines based not only on agreement features but also on form. Previous machine learning classifiers that have attempted to discriminate Romanian nouns according to gender have so far taken as input only the singular form, presupposing the traditional tripartite analysis. We propose a classifier based on two parallel support vector machines using n-gram features from the singular and from the plural which outperforms previous classifiers in its high ability to distinguish the neuter. The performance of our system suggests that the two-gender analysis of Romanian, on which it is based, is on the right track.
Anthology ID:
L12-1379
Volume:
Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12)
Month:
May
Year:
2012
Address:
Istanbul, Turkey
Editors:
Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Mehmet Uğur Doğan, Bente Maegaard, Joseph Mariani, Asuncion Moreno, Jan Odijk, Stelios Piperidis
Venue:
LREC
SIG:
Publisher:
European Language Resources Association (ELRA)
Note:
Pages:
907–910
Language:
URL:
http://www.lrec-conf.org/proceedings/lrec2012/pdf/651_Paper.pdf
DOI:
Bibkey:
Cite (ACL):
Liviu P. Dinu, Vlad Niculae, and Octavia-Maria Şulea. 2012. The Romanian Neuter Examined Through A Two-Gender N-Gram Classification System. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12), pages 907–910, Istanbul, Turkey. European Language Resources Association (ELRA).
Cite (Informal):
The Romanian Neuter Examined Through A Two-Gender N-Gram Classification System (Dinu et al., LREC 2012)
Copy Citation:
PDF:
http://www.lrec-conf.org/proceedings/lrec2012/pdf/651_Paper.pdf