Article In: International Journal of Learner Corpus Research: Online-First Articles
The swiss learner corpus SWIKO
Young learner texts across languages, modes, and tasks
This content is being prepared for publication; it may be subject to changes.
Abstract
This article introduces the Swiss Learner Corpus SWIKO, a multilingual and multi-modal corpus developed at the
Research Centre on Multilingualism in Fribourg (Switzerland). SWIKO comprises over 2,600 annotated oral and written productions by
340 lower-secondary school students. Between 2017–2022, data was elicited through eight communicative tasks systematically varied
by text type, topic, and task constraint and produced under different conditions with respect to mode, modality, and language
(German, French, and English; as both the L1 and L2/L3). All productions are part of speech tagged. In addition, 1,549 written
L2/L3 productions were rated along the CEFR scales, and 890 written German L1 and L2 texts were semi-automatically error
annotated. Due to its controlled task design, rich annotation, and high comparability of learner groups, SWIKO supports a wide
range of empirical and pedagogical applications, including contrastive analyses and investigations of task effects. Open access to
the corpus is provided via https://ifm-swiko.unifr.ch/.
Keywords: learner corpus, tasks, multilingual, multi-modal, CEFR rating
Article outline
- 1.Introduction
- 2.Corpus design
- 2.1Context and aims
- 2.2Scope
- 3.Data collection
- 3.1Participants and procedure
- 3.2Tasks
- 3.3Production conditions
- 4.Data processing
- 4.1Transcription and initial manual markup
- 4.1.1Oral data
- 4.1.2Written data
- 4.2Annotation
- 4.2.1Tokenization, lemmatisation, and POS tagging
- 4.2.2Target hypothesis and semi-automated error annotation in written German productions
- 4.3Rating
- 4.1Transcription and initial manual markup
- 5.Conclusion
- Open data badges and data availability statement
- AI use disclosure statement
- Notes
References
References (56)
Abdi Tabari, M., & Wang, Y. (2022). Assessing
linguistic complexity features in L2 writing: Understanding effects of topic familiarity and strategic planning within the
realm of task readiness. Assessing
Writing, 52, 100605.
Alexopoulou, T., Michel, M., Murakami, A., & Meurers, D. (2017). Task
effects on linguistic complexity and accuracy: A large-scale learner corpus analysis employing natural language processing
techniques. Language
Learning, 67(S1), 180–208.
Arnet-Clark, I., Frank Schmid, S., Ritter, G., & Rüdiger-Harper, J. (2013). New
World — English as a Second Foreign Language. Klett & Balmer Verlag.
Barras, M., Karges, K., & Lenz, P. (2016). Leseverstehen
überprüfen: Welche Sprache für die Fragen und Antworten in den
Testitems?. Babylonia, 16(2), 13–18.
Beers, S. F., & Nagy, W. E. (2011). Writing
development in four genres from grades three to seven: Syntactic complexity and genre
differentiation. Reading and
Writing, 24(2), 183–202.
Bertschy, I., Cuenat, M. E., & Stotz, D. (2015). Lehrplan
Französisch und Englisch. Passepartout — Fremdsprachen an der Volksschule. [URL]
Boyd, A., Hana, J., Nicolas, L., Meurers, D., Wisniewski, K., Abel, A., Schöne, K., Štindlová, B., & Vettori, C. (2014). The
MERLIN corpus: Learner language and the CEFR. In N. Calzolari, K. Choukri, T. Declerck, H. Loftsson, B. Maegaard, J. Mariani, A. Moreno, J. Odijk, and S. Piperidis (Eds.), Proceedings
of the Ninth International Conference on Language Resources and Evaluation
(LREC’14), (pp. 1281–1288). European Language Resources Association (ELRA).
Callies, M. (2023). Current
perspectives on learner corpus research. Arbeiten Aus Anglistik Und
Amerikanistik, 48(1), 37–51.
(2024). Challenges
in the compilation, annotation, and analysis of learner corpus
data. In M. Kaunisto & M. Schilk (Eds), Challenges
in corpus linguistics: Rethinking corpus compilation and
analysis (pp. 55–67). John Benjamins Publishing Company.
Centre for English Corpus
Linguistics. (2026). Learner Corpora around the
world. Université catholique de Louvain. [URL]
Conférence intercantonale de l’instruction publique de la Suisse romande et du Tessin
(CIIP). (2023). Plan d’études romand. Portail CIIP. [URL]
Council of Europe. (2001). Common European
framework of reference for languages: Learning, teaching, assessment. Cambridge University Press. [URL]
. (2020). Common European
framework of reference for languages: Learning, teaching, assessment: Companion
volume. Council of Europe Publishing.
Deutschschweizer Erziehungsdirektoren-Konferenz, Deutsches Institut für internationale
pädagogische Forschung (DIPF), & Nagarro. (2018a). Computer Based
Assessment (CBA) Execution Environment [Computer
software]. Nagarro. [URL]
Deutsches Institut für internationale pädagogische Forschung (DIPF), &
Nagarro. (2018b). Computer Based Assessment (CBA)
ItemBuilder [Computer
software]. Nagarro.
Dirdal, H., Hasund, I. K., Drange, E. M. D., Vold, E. T., & Berg, E. M. (2022). Design
and construction of the Tracking Written Learner Language (TRAWL) Corpus: A longitudinal and multilingual young learner
corpus. Nordic Journal of Language Teaching and
Learning, 10(2), 115–135.
Eckes, T. (2015). Introduction
to many-facet Rasch measurement (2nd ed., Vol. 22). Peter Lang. [URL]
Ellis, R., Skehan, P., Li, S., Shintani, N., & Lambert, C. (2020). Task-based
language teaching: Theory and practice. Cambridge University Press.
European Commission: European Education and Culture Executive
Agency. (2023). Key data on teaching languages at school in Europe: 2023
edition (P. Birch, Ed.). Publications Office of the European Union.
Fernández-Mira, P., Morgan, E., Davidson, S., Yamada, A., Carando, A., Sagae, K., & Sánchez-Gutiérrez, C. H. (2021). Lexical
diversity in an L2 Spanish learner corpus: The effect of topic-related variables. International
Journal of Learner Corpus
Research, 7(2), 230–258.
Gardner, S., Nesi, H., & Biber, D. (2019). Discipline,
level, genre: Integrating situational perspectives in a new MD analysis of university student
writing. Applied
Linguistics, 40(4), 646–674.
Gilquin, G., De Cock, S., & Granger, S. (2010). The
Louvain international database of spoken English interlanguage [Handbook and
CD-ROM]. Presses universitaires de Louvain.
Glaznieks, A., Frey, J. C., Stopfner, M., Zanasi, L., & Nicolas, L. (2022). Leonide:
A longitudinal trilingual corpus of young learners of Italian, German and
English. International Journal of Learner Corpus
Research, 8(1), 97–120.
Granger, S. (2015). Contrastive
interlanguage analysis: A reappraisal. International Journal of Learner Corpus
Research, 1(1), 7–24.
Granger, S., Dupont, M., Meunier, F., Naets, H., & Paquot, M. (2020). The
International Corpus of Learner English. Version 3. [URL]
Granger, S., Gilquin, G., & Meunier, F. (Eds). (2015). The
Cambridge handbook of learner corpus research. Cambridge University Press. [URL].
Hicks, N., & Studer, T. (2024). Learner
corpora in foreign language education: Examples from the multilingual SWIKO corpus. Babylonia
Journal of Language
Education, 2, 26–35.
Hirschmann, H., Lüdeling, A., Shadrova, A., Bobeck, D., Klotz, M., Akbari, R., Schneider, S., & Wan, S. (2022). FALKO.
A collection of richly annotated learner corpora of German as a foreign language. Korpora
Deutsch als Fremdsprache, 2(2).
Housen, A., De Clercq, B., Kuiken, F., & Vedder, I. (2019). Multiple
approaches to complexity in second language research. Second Language
Research, 35(1), 3–21.
Housen, A., Kuiken, F., & Vedder, I. (Eds.). (2012). Dimensions
of L2 performance and proficiency: Complexity, accuracy and fluency in SLA. John Benjamins.
Huang, Y., Geertzen, J., Baker, R., Korhonen, A., & Alexopoulou, T. (2017). The
EF Cambridge open language database (EFCAMDAT): Information for
users (pp. 1–18). [URL]
Karges, K., Barras, M., & Lenz, P. (2021). Assessing
young language learners’ receptive skills: Should we ask the questions in the language of
schooling? In S. Frisch & J. Rymarczyk (Eds.), Current
research into the assessment of young learners’ language
skills (Vol. 30). Peter Lang.
Karges, K., Studer, T., & Hicks, N. S. (2022). Lernersprache,
Aufgabe und Modalität: Beobachtungen zu Texten aus dem Schweizer Lernerkorpus
SWIKO. Zeitschrift für germanistische
Linguistik, 50(1), 104–130.
Konsortium Überprüfung des Erreichens der Grundkompetenzen
(ÜGK). (2019). Überprüfung der Grundkompetenzen. Nationaler Bericht der ÜGK 2017:
Sprachen 8. Schuljahr. Erziehungsdirektorenkonferenz (EDK) & Service de la recherche en éducation (SRED).
. (2025). Nationaler Bericht zu der Überprüfung des Erreichens der
Grundkompetenzen (ÜGK) 2023, Sprachen 11. Schuljahr: ein Beitrag zum Schweizer
Bildungsmonitoring. Interfaculty Centre for Educational Research (ICER).
Lee, H. K., & Anderson, C. (2007). Validity
and topic generality of a writing performance test. Language
Testing, 24(3), 307–330.
Lenz, P., & Studer, T. (2008). Lingualevel:
Instrumente zur Evaluation von Fremdsprachenkompetenzen : 5.-9.
Schuljahr. Schulverlag.
(2022). Facets
computer program for many-facet Rasch measurement (Version 3.84.0) [Computer
software]. [URL]
Lozano, C., Díaz-Negrillo, A., & Callies, M. (2020). Designing
and compiling a learner corpus of written and spoken narratives:
COREFL. In C. Bongartz & J. Torregrossa (Eds.), What’s
in a narrative? Variation in story-telling at the interface between language and
literacy (pp. 21–46). Peter Lang.
Lüdeling, A., & Hirschmann, H. (2015). Error
annotation systems. In F. Meunier, G. Gilquin, & S. Granger (Eds.), The
Cambridge handbook of learner corpus
research (pp. 135–158). Cambridge University Press.
McEnery, T., Brezina, V., Gablasova, D., & Banerjee, J. (2019). Corpus
linguistics, learner corpora, and SLA: Employing technology to analyze language use. Annual
Review of Applied
Linguistics, 39, 74–92.
Michalke, M. (2019, May 13). Package
‘koRpus’. [URL]
Paquot, M., & Plonsky, L. (2017). Quantitative
research methods and study quality in learner corpus research. International Journal of Learner
Corpus
Research, 3(1), 61–94.
Robinson, P. (2011). Second
language task complexity, the cognition hypothesis, language learning, and
performance. In P. Robinson (Ed.), Second
language task complexity: Researching the cognition hypothesis of language learning and
performance (Vol. 2, pp. 3–38). John Benjamins Publishing.
Schmid, H. (2013). TreeTagger
— A language independent part-of-speech tagger (Version 3.2) [Computer
software]. [URL]
Schmidt, T., & Wörner, K. (2009). EXMARaLDA
— Creating, analysing and sharing spoken language corpora for pragmatic
research. Pragmatics, 19(4), 565–582.
Selting, M., Auer, P., Barth-Weingarten, D., Bergmann, J., Bergmann, P., Birkner, K., Couper-Kuhlen, E., Deppermann, A., Gilles, P., Günthner, S., Hartung, M., Kern, F., Mertzlufft, C., Meyer, C., Morek, M., Oberzaucher, F., Peters, J., Quasthoff, U., Schütte, W., … Uhmann, S. (2009). Gesprächsanalytisches
transkriptionssystem 2 (GAT 2). Gesprächsforschung — Online-Zeitschrift zur verbalen
Interaktion, 10, 353–402.
(2009). Modelling
second language performance: Integrating complexity, accuracy, fluency, and lexis. Applied
Linguistics, 30(4), 510–532.
The Conference of Cantonal Directors of Education of German-speaking Switzerland
(D-EDK). (2023). Lehrplan 21. Lehrplan
21. [URL]
Tracy-Ventura, N., & Paquot, M. (Eds). (2020). The
Routledge handbook of second language acquisition and
corpora. Routledge.
van Rooy, B. (2015). Annotating
learner corpora. In F. Meunier, G. Gilquin, & S. Granger (Eds.), The
Cambridge handbook of learner corpus
research (pp. 79–106). Cambridge University Press.
Weiss, Z., Hicks, N. S., Meurers, D., & Studer, T. (2022). Using
linguistic complexity to probe into genre differences? Insights from the multilingual SWIKO learner
corpus [Conference presentation]. Sixth Learner Corpus Research
Conference, 22–24 September 2022, Padua,
Italy.