Article In: International Journal of Corpus Linguistics: Online-First Articles
What a nano-GPT can (not) tell us about spoken language
This content is being prepared for publication; it may be subject to changes.
Abstract
After the launch of ChatGPT in autumn 2022, a lot of research has focussed on the quality and near-naturalness Large Language Model-based tools present in the texts they produce. While one area of research has focussed on the similarities and differences between machine-produced and human-produced output (e.g. Berber Sardinha, T. (2024). AI-generated vs human-authored texts: A multidimensional comparison. Applied Corpus Linguistics, 4(1), 100083. ), others have explored how far such tools could process more complex tasks (e.g. Curry, N., Baker, P., & Brookes, G. (2024). Generative AI for corpus approaches to discourse studies: a critical evaluation of ChatGPT. Applied Corpus Linguistics, 4(1), 100082. ; Valmeekam, K., Marquez, M., Olmo, A., Sreedharan, S., & Kambhampati, S. (2023). PlanBench: An extensible benchmark for evaluating Large Language Models on planning and reasoning about change. arXiv:2206.10498v4. ). While it can be assumed that ChatGPT makes use of written-to-be-spoken training material, there has been no investigation, as yet, into how far a Generative Pre-trained Transformer (GPT) algorithm is able to process (transcribed) natural, colloquial language. This research will investigate whether spoken language transcripts lead to processing difficulties; whether such generated language can be seen as a suitable reflection of natural speech; and whether machine produced texts offer new insights into the workings of language.
Article outline
- 1.Introduction
- 2.The spoken-written divide: Corpus linguistics meets large language models
- 3.Methodology
- 4.Comparing words, word-clusters and keywords
- 4.1Wording
- 4.2Clusters and key clusters
- 4.3Discussion of findings
- 5.Conclusions
- Notes
- Author queries
References
References (51)
Bhatia, A. (2023, April 27). Watch an A. I. learn to write by reading nothing but Jane Austen. The New York Times. [URL]
Berber Sardinha, T. (2024). AI-generated vs human-authored texts: A multidimensional comparison. Applied Corpus Linguistics, 4(1), 100083.
Biber, D., Johansson, S., Leech, G., Conrad, S., & Finegan, E. (2000). Longman grammar of spoken and written English. Longman.
Bisong, E. (2019). Google Colaboratory. In E. Bisong (ed.), Building machine learning and deep learning models on Google Cloud platform: A comprehensive guide for beginners (pp. 59–64). Apress Berkely, CA.
BNC Consortium. (2007). The British National Corpus, XML edition. Oxford Text Archive. [URL]
Brookes, G. (2024, November 28). Concordance analysis is still something that humans do best. Reading Concordances in the 21st Century (RC21). [URL]
Carter, R. (2004). Grammar and spoken English. In C. Coffin, A. Hewings, & K. O’Halloran (Eds.), Applying English grammar (pp. 25–39). Arnold.
Curry, N., Baker, P., & Brookes, G. (2024). Generative AI for corpus approaches to discourse studies: a critical evaluation of ChatGPT. Applied Corpus Linguistics, 4(1), 100082.
Davies, M. (2025). CORPORA AND AI / LLMs: Overview. English-Corpora.org. [URL]
Dohmatob, E., Feng, Y., Yang, P., Charton, F., & Kempe, J. (2024). ICML’24: Proceedings of the 41st International Conference on Machine Learning, 4451, 11165–11197.
Ebrahimi, A., Mager, M., Oncevay, A., Chaudhary, V., Chiruzzo, L., Fan, A., Ortega, J., Ramos, R., Rios, A., Meza Ruiz, I. V., Giménez-Lugo, G., Mager, E., Neubig, G., Palmer, A., Coto-Solano, R., Vu, T., & Kann, K. (2022). AmericasNLI: Evaluating zero-shot natural language understanding of pretrained multilingual models in truly low-resource languages. In S. Muresan, P. Nakov, & A. Villavicencio (Eds.), Proceedings of the 60th annual meeting of the Association for Computational Linguistics (Volume 1: Long papers) (pp. 6279–6299). Association for Computational Linguistics.
The Economist (2026, February 25). AI tools are being prepared for the physical world: The race to build world models is on. The Economist. [URL]
Goldman Sachs (2024, May 14). AI is poised to drive 160% increase in data center power demand. Goldman Sachs. [URL]
Google Labs. (2026, January 29). Project Genie: Experimenting with infinite, interactive worlds. Google. [URL]
Guo, Y., Shang, G., Vazirgiannis, M., & Clavel, C. (2024). The curious decline of linguistic diversity: Training language models on synthetic text. In K. Duh, H. Gomez, & S. Bethard (Eds.), Findings of the Association for Computational Linguistics: NAACL 2023 (pp. 3589–3604). Association for Computational Linguistics.
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de la Casa, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Vinyals, O., Rae, J. W., & Sifre, L. (2022). Training compute-optimal large language models. NIPS’22: Proceedings of the 36th international conference on Neural Information Processing Systems, 21761, 30016–30030.
Karpathy, A. (2024). minGPT. GitHub. [URL]
Knowles, G. O. (1973). Scouse: The urban dialect of Liverpool [Doctoral dissertation, University of Leeds]. White Rose eTheses Online. [URL]
Landgrebe, J., & Smith, B. (2023). Why machines will never rule the world: Artificial intelligence without fear. Routledge.
Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C., & Carlini, N. (2022). Deduplicating training data makes language models better. In S. Muresan, P. Nakov, & A. Villavicencio (Eds.), Proceedings of the 60th annual meeting of the Association for Computational Linguistics (Volume 1: Long papers) (pp. 8424–8445). Association for Computational Linguistics.
Longpre, S., Mahari, R., Lee, A., Lund, C., Oderinwale, H., Brannon, W., Saxena, N., Obeng-Marnu, N., South, T., Hunter, C., Klyman, K., Klamm, C., Scholkopf, H., Singh, N., Cherep, M., Anis, A. M., Dinh, A., Chitongo, C., Yin, D., Sileo, D., Mataciunas, D., Misra, D., Alghamdi, E., Shippole, E., Zhang, J., Materzynska, J., Qian, K., Tiwary, K., Miranda, L., Dey, M., Liang, M., Hamdy, M., Muennighoff, N., Ye, S., Kim, S., Mohanty, S., Gupta, V., Sharma, V., Chien, V. M., Zhou, X., Li, Y., Xiong, C., Villa, L., Biderman, S., Li, H., Ippolito, D., Hooker, S., Kabbara, J., Pentland, & Data Provenance Initiative. (2024). Consent in crisis: The rapid decline of the AI data commons. NIPS ’24: Proceedings of the 38th international conference on Neural Information Processing Systems, 34311, 108042–108087.
Love, R., Dembry, C., Hardie, A., Brezina, V., & McEnery, T. (2017). The Spoken BNC2014: Designing and building a spoken corpus of everyday conversations. International Journal of Corpus Linguistics, 22(3), 319–344.
Lewis, S., & de Leeuw, E. (2025). An acoustic and articulatory investigation into Liverpool and Wirral lateral production in female and male adolescent speech: Covert articulatory variation in ‘Scouse’ and Wirral speech. English Today, 41(1), 3–19.
Monserrate, S. G. (2022). The Cloud is material: On the environmental impacts of computation and data storage. MIT Case Studies in Social and Ethical Responsibilities of Computing, Winter 2022.
O’Keefe, A., McCarthy, M., & Carter, R. (2007). From corpus to classroom: Language use and language teaching. Cambridge University Press.
OpenAI. ([2022] 2024). ChatGPT [Computer Software]. [URL]
. (2023). Whisper ASR [Computer Software]. [URL]
(2018). Spreading activation, Lexical Priming and the semantic web: Early psycholinguistic theories, corpus linguistics and AI applications. Palgrave Macmillan.
(2025). Large-Language-Model tools and the theory of Lexical Priming: Where technology and human cognition meet and diverge. Journal of Corpora and Discourse Studies, 91, 1–22.
Pace-Sigge, M., & Sumakul, T. (2022). What teaching an algorithm teaches when teaching students how to write academic texts. In J. H. Jantunen, J. Kalja-Voima, M. Laukkarinen, A. Puupponen, M. Salonen, T. Saresma, J. Tarvainen, & S. Ylönen (Eds.), Diversity of methods and materials in digital human sciences: Proceedings of the Digital Research Data and Human Sciences (DRDHum) Conference (pp. 230–243). University of Jyväskylä. [URL]
Rayson, P. (2016). Log-likelihood calculator [Computer Software]. UCREL — University Centre for Computer Corpus Research. [URL]
Raza, S., Bamgbose, O., Ghuge, S., Tavakol, F., Reji, D. J., & Bashir, S. R. (2025). Developing safe and responsible large language model: Can we balance bias reduction and language understanding?. Machine Learning, 1141, 140.
Reisner, A. (2023, September 25). These 183,000 books are fueling the biggest fight in publishing and tech. The Atlantic. [URL]
Sanyal, S., Sanghavi, S., & Dimakis, A. G. (2024). Pre-training small base LMs with fewer tokens. arXiv:2404.08634v1. [URL]
Scott, M. (2023). WordSmith Tools (Version 8) [Computer Software]. Lexical Analysis Software. [URL]
Sinclair, J., Jones, S., & Daley, R. (1970) [2004]. English collocation studies: The OSTI report. Continuum.
Shumailov, I., Shumaylov, Z., Zhao, Y., Gal, Y., Papernot, N., & Anderson, R. (2023). The curse of recursion: Training on generated data makes models forget. arxiv:2305.17493.
Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). AI models collapse when trained on recursively generated data. Nature, 631(8022), 755–759.
Strubell, E., Ganesh, A., & McCallum, A. (2020). Energy and policy considerations for modern deep learning research. In A. Korhonen, D. Traum, & L. Màrquez (Eds.), Proceedings of the 57th annual meeting of the Association for Computational Linguistics (pp. 3645–3650). Association for Computational Linguistics.
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., & Lample, G. (2023). LLaMA: Open and efficient foundation language models. arXiv:2302.13971v1.
Valmeekam, K., Marquez, M., Olmo, A., Sreedharan, S., & Kambhampati, S. (2023). PlanBench: An extensible benchmark for evaluating Large Language Models on planning and reasoning about change. arXiv:2206.10498v4.
Wang, Y., Xu, C., Sun, Q., Hu, H., Tao, C., Geng, X., & Jiang, D. (2022). PromDA: Prompt-based data augmentation for low-resource NLU tasks. In S. Muresan, P. Nakov, & A. Villavicencio (Eds.), Proceedings of the 60th annual meeting of the Association for Computational Linguistics (Volume 1: Long Papers), (pp. 4242–4255). Association for Computational Linguistics.