Development of a Web-Scale Chinese Word N-gram Corpus with Parts of Speech Information

Chi-Hsin Yu, Yi-jie Tang, Hsin-Hsi Chen


Abstract
Web provides a large-scale corpus for researchers to study the language usages in real world. Developing a web-scale corpus needs not only a lot of computation resources, but also great efforts to handle the large variations in the web texts, such as character encoding in processing Chinese web texts. In this paper, we aim to develop a web-scale Chinese word N-gram corpus with parts of speech information called NTU PN-Gram corpus using the ClueWeb09 dataset. We focus on the character encoding and some Chinese-specific issues. The statistics about the dataset is reported. We will make the resulting corpus a public available resource to boost the Chinese language processing.
Anthology ID:
L12-1340
Volume:
Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12)
Month:
May
Year:
2012
Address:
Istanbul, Turkey
Editors:
Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Mehmet Uğur Doğan, Bente Maegaard, Joseph Mariani, Asuncion Moreno, Jan Odijk, Stelios Piperidis
Venue:
LREC
SIG:
Publisher:
European Language Resources Association (ELRA)
Note:
Pages:
320–324
Language:
URL:
http://www.lrec-conf.org/proceedings/lrec2012/pdf/596_Paper.pdf
DOI:
Bibkey:
Cite (ACL):
Chi-Hsin Yu, Yi-jie Tang, and Hsin-Hsi Chen. 2012. Development of a Web-Scale Chinese Word N-gram Corpus with Parts of Speech Information. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12), pages 320–324, Istanbul, Turkey. European Language Resources Association (ELRA).
Cite (Informal):
Development of a Web-Scale Chinese Word N-gram Corpus with Parts of Speech Information (Yu et al., LREC 2012)
Copy Citation:
PDF:
http://www.lrec-conf.org/proceedings/lrec2012/pdf/596_Paper.pdf