Identifying Incorrect Labels in the CoNLL-2003 Corpus

Frederick Reiss, Hong Xu, Bryan Cutler, Karthik Muthuraman, Zachary Eichenberger


Abstract
The CoNLL-2003 corpus for English-language named entity recognition (NER) is one of the most influential corpora for NER model research. A large number of publications, including many landmark works, have used this corpus as a source of ground truth for NER tasks. In this paper, we examine this corpus and identify over 1300 incorrect labels (out of 35089 in the corpus). In particular, the number of incorrect labels in the test fold is comparable to the number of errors that state-of-the-art models make when running inference over this corpus. We describe the process by which we identified these incorrect labels, using novel variants of techniques from semi-supervised learning. We also summarize the types of errors that we found, and we revisit several recent results in NER in light of the corrected data. Finally, we show experimentally that our corrections to the corpus have a positive impact on three state-of-the-art models.
Anthology ID:
2020.conll-1.16
Volume:
Proceedings of the 24th Conference on Computational Natural Language Learning
Month:
November
Year:
2020
Address:
Online
Editors:
Raquel Fernández, Tal Linzen
Venue:
CoNLL
SIG:
SIGNLL
Publisher:
Association for Computational Linguistics
Note:
Pages:
215–226
Language:
URL:
https://aclanthology.org/2020.conll-1.16
DOI:
10.18653/v1/2020.conll-1.16
Bibkey:
Cite (ACL):
Frederick Reiss, Hong Xu, Bryan Cutler, Karthik Muthuraman, and Zachary Eichenberger. 2020. Identifying Incorrect Labels in the CoNLL-2003 Corpus. In Proceedings of the 24th Conference on Computational Natural Language Learning, pages 215–226, Online. Association for Computational Linguistics.
Cite (Informal):
Identifying Incorrect Labels in the CoNLL-2003 Corpus (Reiss et al., CoNLL 2020)
Copy Citation:
PDF:
https://aclanthology.org/2020.conll-1.16.pdf
Code
 codait/text-extensions-for-pandas +  additional community code
Data
CoNLL 2003CoNLL++RCV1