Annotating the Contemporary Chinese Corpus

Zhou, Qiang; Yu, Shiwen

doi:10.1075/ijcl.2.2.05qia

Article published In:

International Journal of Corpus Linguistics
Vol. 2:2 (1997) ► pp.239–258

Annotating the Contemporary Chinese Corpus

Qiang Zhou | Institute of Computational Linguistics, Peking University

Shiwen Yu | Institute of Computational Linguistics, Peking University

In recent years, great progress has been made in Chinese corpus processing. A fifty-million-word Chinese National Corpus project has been put into effect, and many automatic corpus processing programs have also been developed. In this paper, we will briefly introduce our work on constructing a large scale annotated corpus for Chinese grammatical research and developing a Chinese Corpus Multilevel Processing system—CCMP. First, we present our annotation scheme. Second, we discuss some basic methodologies for Chinese corpus analysis and propose a man-machine mutually dependent corpus processing model. Finally, we introduce the survey of our CCMP. We hope our work will give impetus to further research in Chinese corpus linguistics.

Keywords: Bracketing, Part-of-Speech Tagging, Segmentation, Chinese Corpus Annotation

Published online: 1 January 1997

https://doi.org/10.1075/ijcl.2.2.05qia

Cited by (1)

Cited by 1 other publications

Xiao-Shen Zheng, Pi-Lian He, Mei Tian, Zhong Wang & Yue-Heng Sun

2002. Proceedings. International Conference on Machine Learning and Cybernetics, ► pp. 1747 ff.

This list is based on CrossRef data as of 4 july 2024. Please note that it may not be complete. Sources presented here have been supplied by the respective publishers. Any errors therein should be reported to them.