Article In: International Journal of Learner Corpus Research: Online-First Articles
L2K-ARG
Introducing and evaluating a learner corpus of Korean argumentative writing
This content is being prepared for publication; it may be subject to changes.
Abstract
The present study introduces an open-access corpus of Korean second language argumentative writing (L2K-ARG)
produced by university-level learners with different first-language backgrounds. The corpus contains 484 essays written in
response to three standardized prompts under timed, prompt-controlled settings. Its design was informed by the framework proposed
by Egbert, J., Biber, D., & Gray, B. (2022). Designing
and evaluating language corpora: A practical framework for corpus
representativeness. Cambridge University Press. , which defines representativeness based on how well a corpus
reflects the target domain and the distribution of linguistic features within that domain. To characterize the learners presented
in the corpus, the dataset includes general Korean proficiency as well as writing proficiency scores based on an analytic rubric.
This resource aims to complement existing Korean learner corpora by providing a publicly available, prompt-controlled
argumentative writing samples with learner proficiency information, supporting reproducible research in L2 Korean.
Article outline
- 1.Introduction
- 2.Backgrounds
- 2.1Available open-access corpora for L2 Korean
- 2.2Design principles for language corpora: Insights from Egbert et al. (2022)
- 3.Corpus design and data collection
- 3.1Domain analysis and design considerations
- 3.1.1Describing the domain
- 3.1.2Operationalizing the domain
- 3.1.3Planning the sample
- 3.2Data collection procedures
- 3.1Domain analysis and design considerations
- 4.Distribution analysis
- 4.1Linguistic feature annotation
- 4.2Evaluation of distributional representation
- 5.Corpus overview
- 5.1Descriptive statistics
- 5.2General Korean proficiency profile
- 5.3Writing proficiency profile
- 5.4Release format
- 6.Conclusion and future directions
- Data Availability Statement
- AI Use Disclosure Statement
- Notes
References
References (38)
Biber, D. (1993). Representativeness
in corpus design. Literary and Linguistic
Computing, 8(4), 243–257.
Centre for English Corpus
Linguistics. (n.d.). Learner corpora around the
world. UCLouvain. [URL]
CLARIN ERIC. (n.d.). L2 learner
corpora. [URL]
Crossley, S. A., Tian, Y., Baffour, P., Franklin, A., Kim, Y., Morris, W., Benner, B., Picou, A., & Boser, U. (2023). The
English language learner insight, proficiency and skills evaluation (ELLIPSE)
corpus. International Journal of Learner Corpus
Research, 9(2), 248–269.
de Marneffe, M. C., Manning, C. D., Nivre, J., & Zeman, D. (2021). Universal
dependencies. Computational
Linguistics, 47(2), 255–308.
Dickinson, M., Israel, R., & Lee, S. H. (2010). Building
a Korean web corpus for analyzing learner language. In Proceedings of
the Sixth Web as Corpus Workshop (NAACL HLT
2010) (pp. 8–16). Association for Computational Linguistics.
Duolingo. (2025). 2025 Duolingo language
report. [URL] (accessed
on 11-March-2026).
Eckes, T., & Grotjahn, R. (2006). A
closer look at the construct validity of C-tests. Language
Testing, 23(3), 290–325.
Egbert, J., Biber, D., & Gray, B. (2022). Designing
and evaluating language corpora: A practical framework for corpus
representativeness. Cambridge University Press.
Fouser, R. J. (2001). Sociolinguistic
transfer from Japanese into Korean as an L≥ 3. Cross-Linguistic Influence in Third Language
Acquisition: Psycholinguistic
Perspectives, 311, 149.
Gilquin, G., & Granger, S. (2015). Learner
language. In D. Biber & R. Reppen (Eds.), The
cambridge handbook of English corpus
linguistics (pp. 418–435). Cambridge University Press.
Granger, S. (2008). The
contribution of learner corpora to second language acquisition and foreign language teaching: A critical
evaluation. In A. Frank, B. Kettemann, & H. Mehlmauer-Larcher (Eds.), Corpora
and language
teaching (pp. 13–32). John Benjamins.
Hirvela, A. (2017). Argumentation
& second language writing: Are we missing the boat?. Journal of Second Language
Writing, 361, 69–74.
Jarvis, S., & Paquot, M. (2015). Native
language identification. In S. Granger, G. Gilquin, & F. Meunier (Eds.), The
cambridge handbook of learner corpus
research (pp. 605–627). Cambridge University Press.
Joshi, P., Santy, S., Budhiraja, A., Bali, K., & Choudhury, M. (2020). The
state and fate of linguistic diversity and inclusion in the NLP
world. In Proceedings of the 58th Annual Meeting of the Association
for Computational
Linguistics (pp. 6282–6293). Association for Computational Linguistics.
Koo, T. K., & Li, M. Y. (2016). A
guideline of selecting and reporting intraclass correlation coefficients for reliability
research. Journal of Chiropractic
Medicine, 15(2), 155–163.
Kyle, K., & Eguchi, M. (2024). Evaluating
NLP models with written and spoken L2 samples. Research Methods in Applied
Linguistics, 3(2), 100120.
Lee, S. H., Jang, S. B., & Seo, S. K. (2009). Annotation
of Korean learner corpora for particle error detection. Calico
Journal, 26(3), 529–544.
Lee, S. H., Dickinson, M., & Israel, R. (2012). Developing
learner corpus annotation for Korean particle errors. In Proceedings
of the Sixth Linguistic Annotation
Workshop (pp. 129–133). Association for Computational Linguistics.
Lee-Ellis, S. (2009). The
development and validation of a Korean C-Test using Rasch Analysis. Language
Testing, 26(2), 245–274.
Lim, K., Song, J., & Park, J. (2023). Neural
automated writing evaluation for Korean L2 writing. Natural Language
Engineering, 29(5), 1341–1363.
Manning, C. D. (2015). Computational
linguistics and deep learning. Computational
Linguistics, 41(4), 701–707.
Masciolini, A., Berdičevskis, A., Szawerna, M. I., & Volodina, E. (2025a). Annotating
second language in Universal Dependencies: A review of current practices and directions for harmonized
guidelines. In Proceedings of the Eighth Workshop on Universal
Dependencies (UDW, SyntaxFest
2025) (pp. 153–163). Association for Computational Linguistics.
Masciolini, A., Caines, A., De Clercq, O., Kruijsbergen, J., Kurfalı, M., Muñoz Sánchez, R., Volodina, E., Östling, R., Allkivi, K., Arhar Holdt, S., Auzina, I., Darģis, R., Drakonaki, E., Frey, J. C., Glišić, I., Kikilintza, P., Nicolas, L., Romanyshyn, M., Rosen, A., ... & Zesch, T. (2025b). Towards
better language representation in Natural Language Processing: A multilingual dataset for text-level Grammatical Error
Correction. International Journal of Learner Corpus
Research, 11(2), 309–335.
Meurers, D. (2015). Learner
corpora and natural language processing. In S. Granger, G. Gilquin, & F. Meunier (Eds.), The
cambridge handbook of learner corpus
research (pp. 537–566). Cambridge University Press.
National Institute of Korean
Language. (n.d.). Korean Learners’ Corpus Search. [URL]
Norris, J. M., & Ortega, L. (2009). Towards
an organic approach to investigating CAF in instructed SLA: The case of complexity. Applied
Linguistics, 30(4), 555–578.
Park, J., & Lee, J. H. (2016). A
Korean learner corpus and its features. En-e-hak
[Linguistics], (75), 69–85.
Song, S., Yuk, J., Choi, C., Yoo, H., Lim, H., Lim, K., & Park, J. (2025). Unified
automated essay scoring and grammatical error correction. In Findings
of the Association for Computational Linguistics: NAACL
2025 (pp. 4412–4426). Association for Computational Linguistics.
Song, J., Lim, K., & Park, J. (2026). Enriching
the Korean learner corpus for grammatical error correction and writing assessment. Language
Resources and
Evaluation, 60(1), 15.
Sung, H., & Shin, G-H. (2023). Towards
L2-friendly pipelines for learner corpora: A case of written production by L2-Korean
learners. In Proceedings of the 18th Workshop on Innovative Use of
NLP for Building Educational Applications (BEA
2023) (pp. 72–82). Association for Computational Linguistics.
(2024). Constructing
a dependency treebank for second language learners of
Korean. In Proceedings of the 2024 Joint International Conference on
Computational Linguistics, Language Resources and Evaluation (LREC-COLING
2024) (pp. 3747–3758). ELRA and ICCL.
(2025a). Second
language Korean Universal Dependency treebank v1. 2: Focus on data augmentation and annotation scheme
refinement. In Proceedings of the Third Workshop on Resources and
Representations for Under-Resourced Languages and Domains
(RESOURCEFUL-2025) (pp. 13–19). University of Tartu Library, Estonia.
(2025b). Towards
robust morphosyntactic analysis of L2 Korean: Evaluating and fine-tuning a Korean language
model. ACM Transactions on Asian and Low-Resource Language Information
Processing, 24(11), 1–21.
(2026). Parser
agreement and disagreement in L2 Korean UD: Implications for human-in-the-loop
annotation. In Proceedings of the 20th Linguistic Annotation Workshop
(LAW
XX) (pp. 12–21). Association for Computational Linguistics.