Named entity recognition in Bengali and Hindi using support vector machine

Ekbal, Asif; Bandyopadhyay, Sivaji

doi:10.1075/li.34.1.02ekb

Article published In:

Lingvisticæ Investigationes
Vol. 34:1 (2011) ► pp.35–67

Named entity recognition in Bengali and Hindi using support vector machine

Asif Ekbal | Indian Institute of Technology Patna, Patna, India

Sivaji Bandyopadhyay

Named Entity Recognition (NER) aims to classify each word of a document into predefined target named entity (NE) classes and is nowadays considered to be fundamental for many Natural Language Processing (NLP) tasks such as information retrieval, machine translation, information extraction, question answering systems and others. This paper reports about the development of a NER system for Bengali and Hindi using Support Vector Machine (SVM). We have used the annotated corpora of 122,467 tokens of Bengali and 502,974 tokens of Hindi tagged with the twelve different NE classes, defined as part of the IJCNLP-08 NER Shared Task for South and South East Asian Languages (SSEAL). An appropriate tag conversion routine has been developed in order to convert the data into the forms tagged with the four NE tags, namely Person name, Location name, Organization name and Miscellaneous name. The system makes use of the different contextual information of the words along with the variety of orthographic word-level features that are helpful in predicting the different NE classes. The system has been tested with the gold standard test sets of 35K, and 38K tokens for Bengali, and Hindi, respectively. Evaluation results have demonstrated the overall recall, precision, and f-score values of 85.11%, 81.74%, and 83.39%, respectively, for Bengali and 82.76%, 77.81%, and 80.21%, respectively, for Hindi. Statistical analysis, ANOVA is performed to show that the improvement in the performance with the use of language dependent features is statistically significant over the language independent features for Bengali and Hindi both.

Article language: French

Published online: 7 July 2011

https://doi.org/10.1075/li.34.1.02ekb

Cited by (11)

Cited by 11 other publications

Order by:

Singh, Navdeep, Munish Kumar, Bavalpreet Singh & Jaskaran Singh

2023. DeepSpacy-NER: an efficient deep learning model for named entity recognition for Punjabi language. Evolving Systems 14:4 ► pp. 673 ff.

Avishka, Wap, Banujan Kuhaneswaran & Hn Gunasinghe

2022. 2022 2nd International Conference on Advanced Research in Computing (ICARC), ► pp. 254 ff.

Jain, Arti, Devendra K. Tayal, Divakar Yadav & Anuja Arora

2020. Research Trends for Named Entity Recognition in Hindi Language. In Data Visualization and Knowledge Engineering [Lecture Notes on Data Engineering and Communications Technologies, 32], ► pp. 223 ff.

Anandika, Amrita & Smita Prava Mishra

2019. 2019 International Conference on Applied Machine Learning (ICAML), ► pp. 153 ff.

Jain, Arti & Anuja Arora

2018. Named Entity System for Tweets in Hindi Language. International Journal of Intelligent Information Technologies 14:4 ► pp. 55 ff.

Remmiya Devi, G., M. Anand Kumar, K.P. Soman, Sabu M. Thampi, El-Sayed M. El-Alfy, Sushmita Mitra & Ljiljana Trajkovic

2018. Co-occurrence based word representation for extracting named entities in Tamil tweets. Journal of Intelligent & Fuzzy Systems 34:3 ► pp. 1435 ff.

Sathyanarayanan, Dinesh, Ashwin Ashok, Debanik Mishra, Santwana Chimalamarri & Dinkar Sitaram

2018. 2018 International Conference on Electrical, Electronics, Communication, Computer, and Optimization Techniques (ICEECCOT), ► pp. 65 ff.

Rahman, Yeasin Ar, Mahtabul Alam Sohan, Khalid Ibn Zinnah & Mohammed Moshiul Hoque

2017. 2017 International Conference on Electrical, Computer and Communication Engineering (ECCE), ► pp. 935 ff.

Talukdar, Gitimoni, Pranjal Protim Borah & Arup Baruah

2014. 2014 International Conference on Contemporary Computing and Informatics (IC3I), ► pp. 187 ff.

Fu, Chunyuan & Guohong Fu

2011. 2011 Eighth International Conference on Fuzzy Systems and Knowledge Discovery (FSKD), ► pp. 1221 ff.

Fu, Chunyuan & Guohong Fu

2012. 2012 9th International Conference on Fuzzy Systems and Knowledge Discovery, ► pp. 2546 ff.

This list is based on CrossRef data as of 5 july 2024. Please note that it may not be complete. Sources presented here have been supplied by the respective publishers. Any errors therein should be reported to them.