Article In: International Journal of Corpus Linguistics: Online-First Articles
Benchmark of linguistic variation in LLM‑generated texts
This content is being prepared for publication; it may be subject to changes.
Abstract
This study investigates the register variation in texts written by humans and comparable texts produced by large language
models (LLMs). Multidimensional analysis (MDA) is applied to a sample of human-written texts and AI-generated counterpart texts to find the
dimensions of variation in which LLMs differ most significantly and most systematically from humans. We introduce the LLM-generated corpus
AI-Brown, which is comparable to BE-21: a Brown family corpus representing contemporary British English. Since all languages except English
are underrepresented in the training data of frontier LLMs, similar analysis is replicated on Czech language data using the AI-Koditex
corpus and a Czech multidimensional model (Cvrček et al., 2018). Sixteen frontier models were examined in various settings and prompts, with
emphasis placed on the difference between base models and instruction-tuned models. We offer a benchmark through which models can be
compared with each other and ranked in interpretable dimensions.
Article outline
- 1.Introduction
- 2.Stylometry and register variation in LLMs
- 3.Data
- 3.1Corpora
- 3.1.1English
- 3.1.2Czech
- 3.2Models
- 3.3Sampling temperature
- 3.4Corpus generation procedure and prompts
- 3.1Corpora
- 4.Analysis methods
- 4.1Multidimensional analysis
- 4.2Statistical processing
- 4.3Clustering of models
- 5.Results and discussion
- 5.1Overall results
- 5.1.1English
- 5.1.2Czech
- 5.2Base versus instruction tuned
- 5.2.1English
- 5.2.2Czech
- 5.3Prompt impact
- 5.3.1English
- 5.3.2Czech
- 5.4Temperature impact
- 5.1Overall results
- 6.Conclusion
- Acknowledgements
- Data availability
- Declaration on using AI
- Notes
- Author queries
References
References (46)
Altman, S. (2023, February 24). Planning
for AGI and beyond. OpenAI. [URL]
Amodei, D. (2024, October 24). Machines
of loving grace: How AI could transform the world for the better. darioamodei.com. [URL]
Anthropic. (2024, June 8). Claude’s
character. Anthropic. [URL]
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Chen, C., Olsson, C., Olah, C., Hernandez, D., Drain, D., Ganguli, D., Li, D., Tran-Johnson, E., Perez, E., Kerr, J., Mueller, J., Ladish, J., Landau, J., Ndousse, K., Lukosuite, K., Lovitt, L., Sellitto, M., Elhage, N., Schiefer, N., Mercado, N., DasSarma, N., Lasenby, R., Larson, R., Ringer, S., Johnston, S., Kravec, S., El Showk, S., Fort, S., Lanham, T., Telleen-Lawton, T., Conerly, T., Henighan, T., Hume, T., Bowman, S. R., Hatfield-Dodds, Z., Mann, B., Amodei, D., Joseph, N., McCandlish, S., Brown, T., & Kaplan, J. (2022). Constitutional
AI: Harmlessness from AI feedback. arXiv
preprint arXiv:2212.08073. [URL]
Baker, P. (2023). A
year to remember? Introducing the BE21 corpus and exploring recent part of speech tag change in British
English. International Journal of Corpus
Linguistics, 28(3), 407–429.
Balloccu, S., Schmidtová, P., Lango, M., & Dušek, O. (2024). Leak,
cheat, repeat: Data contamination and evaluation malpractices in closed-source
LLMs. In Y. Graham & M. Purver (Eds.), Proceedings
of the 18th conference of the European chapter of the Association for Computational
Linguistics (pp. 67–93). Association for Computational Linguistics.
Berber Sardinha, T. (2024). AI-generated
vs human-authored texts: A multidimensional comparison. Applied Corpus
Linguistics, 4(1), 100083.
(2019). Multi-dimensional
analysis: A historical synopsis. In T. Berber Sardinha & M. Veirano Pinto (Eds.), Multi-dimensional
analysis: Research methods and current
issues (pp. 11–26). Bloomsbury.
Bitton, Y., Bitton, E., & Nisan, S. (2025). Detecting
stylistic fingerprints of large language models. arXiv:2503.01659v1.
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., MCandlish, S., Radford, A., Sutskever, I., & Amodei, D. (2020). Language
models are few-shot learners. arXiv:2005.14165v4. [URL]
Common Crawl Foundation. (2025). Common
Crawl. [URL]
Cvrček, V., Komrsková, Z., Lukeš, D., Poukarová, P., Řehořková, A., & Zasina, A. J. (2021). From
extra- to intratextual characteristics: Charting the space of variation in Czech through MDA. Corpus
Linguistics and Linguistic
Theory, 17(2), 351–382.
Cvrček, V., Laubeová, Z., Lukeš, D., Poukarová, P., Řehořková, A., & Zasina, A. J. (2020). Registry
v češtině. NLN.
da Silva, T. H., Furtado, V., Furtado, E., Mendes, M., Almeida, V., & Sales, L. (2024). How
do illiterate people interact with an intelligent voice assistant? International Journal of Human–Computer
Interaction, 40(3), 584–602.
Garg, A. (2025, June 10). Google
claims AI helping engineers do 10 percent more productive tasks, says AI agents are happening. India
Today. [URL]
Goulart, L., Matte, M. L., Mendoza, A., Alvarado, L., & Veloso, I. (2024). AI
or student writing? Analyzing the situational and linguistic characteristics of undergraduate student writing and AI-generated
assignments. Journal of Second Language
Writing, 661, 101160.
Johnson, R. L., Pistilli, G., Menédez-González, N., Duran, L. D. D., Panai, E., Kalpokiene, J., & Bertulfo, D. J. (2022). The
ghost in the machine has an American accent: Value conflict in
GPT-3. arXiv:2203.07785v1.
Kubát, M. (2014). Moving
window type-token ratio and text length. In G. Altmann, R. Čech, J. Mačutek, & L. Uhlířová (Eds.), Empirical
approaches to text and language
analysis (pp. 105–113). RAM-Verlag.
Kumarage, T., Garland, J., Bhattacharjee, A., Trapeznikov, K., Ruston, S., & Liu, H. (2023). Stylometric
detection of AI-generated text in Twitter timelines. arXiv:2303.03697v1. [URL]
Malik, M., Jiang, J., & Chai, K. M. A. (2024). An
empirical analysis of the writing styles of persona-assigned LLMs. In Y. Al-Onaizan, M. Bansal, & Y.-N. Chen (Eds.), Proceedings
of the 2024 conference on empirical methods in natural language
processing (pp. 19369–19388). Association for Computational Linguistics.
Mikros, G. (2025). Beyond
the surface: Stylometric analysis of GPT-4o’s capacity for literary style imitation. Digital Scholarship in
the
Humanities, 40(2), 587–600.
Mikros, G. K., Koursaris, A., Bilianos, D., & Markopoulos, G. (2023). AI-writing
detection using an ensemble of transformers and stylometric features. In M. Montes-y-Gómez, F. Rangel, S. M. Jiménez-Zafra, M. Casavantes, B. Altuna, M. Á. Álvarez-Carmona, G. Bel-Enguix, L. Chiruzzo, I. de la Iglesia, H. J. Escalante, M. Á. García-Cumbreras, J. A. García-Díaz, J. Á. González Barba, R. L. Tamayo, S. Lima, P. Moral, F. Miriam, P. del Arco, & R. Valencia-García (Eds.), Proceedings
of the Iberian Languages Evaluation Forum (IberLEF 2023) co-located with the conference of the Spanish Society for Natural Language
Processing (SEPLN 2023). CEUR Workshop Proceedings. [URL]
Milička, J., Marklová, A., VanSlambrouck, K., Pospíšilová, E., Šimsová, J., Harvan, S., & Drobil, O. (2024). Large
language models are able to downplay their cognitive abilities to fit the persona they simulate. PLOS
One, 19(3), e0298522.
Milička, J., Marklová, A., & Cvrček, V. (2025a). AI
Brown v1. LINDAT/CLARIAH-CZ Digital library at the Institute of Formal and Applied Linguistics (ÚFAL). [URL]
(2025b). AI
Koditex v1. LINDAT/CLARIAH-CZ Digital library at the Institute of Formal and Applied Linguistics (ÚFAL), [URL]
(in
press). AI Brown and AI Koditex: LLM-generated corpora comparable to human-written corpora of English and
Czech texts. Language Resources and Evaluation.
Milička, J., Marklová, A., Drobil, O., & Pospíšilová, E. (2025). Learning
to detect AI texts and learning the limits. PloS
One, 20(10), e0333007.
Ni, S., Kong, X., Li, C., Hu, X., Xu, R., Zhu, J., & Yang, M. (2025). Training
on the benchmark is not all you need. Proceedings of the AAAI Conference on Artificial
Intelligence, 39(23), 24948–24956.
Nini, A. (2019). The
multi-dimensional analysis tagger. In T. Berber Sardinha & M. Veirano Pinto (Eds.), Multi-dimensional
analysis: Research methods and current
issues (pp. 67–94). Bloomsbury.
Peeperkorn, M., Kouwenhoven, T., Brown, D., & Jordanous, A. (2024). Is
temperature the creativity parameter of large language models? arXiv:2405.00492v1.
Przystalski, K., Argasiński, J. K., Grabska-Gradzińska, I., & Ochab, J. K. (2026). Stylometry
recognizes human and LLM-generated texts in short samples. Expert Systems with
Applications, 2961, 129001.
Rao, Z., Mohamed, Y., Liu, S., & Liu, Z. (2025). Two
birds with one stone: Multi-task detection and attribution of LLM-generated
text. arXiv:2508.14190v1.
Reinhart, A., Markey, B., Laudenbach, M., Pantusen, K., Yurko, R., Weinberg, G., & Brown, D. W. (2025). Do
LLMs write like humans? Variation in grammatical and rhetorical styles. Proceedings of the National Academy
of
Sciences, 122(8), e2422455122.
Rudnicka, K. (2025a, July 9). Each
AI chatbot has its own, distinctive writing style just as humans do. Scientific
American. [URL]
Schut, L., Gal, Y., & Farquhar, S. (2025). Do
multilingual LLMs think in English? arXiv:2502.15603v1.
Shanahan, M., McDonell, K., & Reynolds, L. (2023). Role
play with large language
models. Nature, 6231, 493–498.
Tully, T., Redfern, J., Das, D., & Xiao, D. (2025, July 13). 2025
Mid-year LLM market update: Foundation model landscape + economics. Menlo Ventures. [URL]
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention
is all you need. In U. von Luxburg, I. Guyon, S. Bengio, H. Wallach, & R. Fergus (Eds.), NIPS’17:
Proceedings of the 31st international conference on Neural Information Processing
System, (pp. 6000–6010). Association for Computing Machinery.
Xu, H., Shi, Z. J., & Shi, M. (2025). Bonding
with AI: Investigating the love relationships between humans and AI companions [Master’s
thesis]. The Hong Kong University of Science and Technology.
Zasina, J., Lukeš, D., Komrsková, Z., Poukarová, P., & Řehořková, A. (2018). Koditex:
A corpus of diversified texts. Institute of the Czech National Corpus, Faculty of Arts, Charles University. [URL]