This paper discusses the approach of developing a sample of printed corpus in Bangla, one of the national languages of India and the only national language of Bangladesh. It is designed from the data collected from various published documents. The paper highlights different issues related to corpus generation, data-file preparation, language analysis, and processing as well as application potentials to different areas of pure and applied linguistics. It also includes statistical studies on the corpus along with some interpretation of the results. The difficulties that one may face during corpus generation are also pointed out.
2021. In search of a suitable method for disambiguation of word senses in Bengali. International Journal of Speech Technology 24:2 ► pp. 439 ff.
Parameswarappa, S., V. N. Narayana & G. N. Bharathi
2012. 2012 International Conference on Computer Communication and Informatics, ► pp. 1 ff.
Dash, N.S. & B.B. Chaudhuri
2003. Language Engineering Conference, 2002. Proceedings, ► pp. 99 ff.
This list is based on CrossRef data as of 11 september 2024. Please note that it may not be complete. Sources presented here have been supplied by the respective publishers.
Any errors therein should be reported to them.