References (38)
References
Atkins, B. T. S., & Rundell, M. (2008). The Oxford guide to practical lexicography. Oxford University Press. Google Scholar logo with link to Google Scholar
Baker, M. (1993). Corpus linguistics and translation studies: Implications and applications. In M. Baker, G. Francis, & E. Tognini-Bonelli (Eds.), Text and technology: In honour of John Sinclair (pp.233–250). John Benjamins. Google Scholar logo with link to Google Scholar
Bañón, M. et al. (2020). ParaCrawl: Web-scale acquisition of parallel corpora. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp.4555–4567). Association for Computational Linguistics. Google Scholar logo with link to Google Scholar
Bernardini, S., & Castagnoli, S. (2008). Corpora for translator education and professional practice. In Proceedings of the 8th International Conference on Terminology and Artificial Intelligence (TIA).Google Scholar logo with link to Google Scholar
Bowker, L., & Pearson, J. (2002). Working with specialized language: A practical guide to using corpora. Routledge. Google Scholar logo with link to Google Scholar
Braune, F., & Fraser, A. (2010). Improved sentence alignment using maximum entropy classifiers. In Proceedings of the 2010 Conference on Empirical Methods in Natural Language Processing (pp.170–178). Association for Computational Linguistics.Google Scholar logo with link to Google Scholar
Brown, P. F. et al. (1991). Aligning sentences in parallel corpora. In Proceedings of the 29th Annual Meeting of the Association for Computational Linguistics (pp.169–176). Association for Computational Linguistics. Google Scholar logo with link to Google Scholar
Cettolo, M., Girardi, C., & Federico, M. (2012). WIT³: Web inventory of transcribed and translated talks. In Proceedings of the 16th Conference of the European Association for Machine Translation (EAMT) (pp.261–268). European Association for Machine Translation.Google Scholar logo with link to Google Scholar
Doval, I., & Jiménez, T. (2020). Multifuncionalidad de los corpus paralelos, ejemplificado en el corpus alemán/español PaGeS. In M. Blanco, H. Olbertz, & V. Vázquez Rozas (Eds.), Corpus y construcciones. Perspectivas hispánicas (pp.305–322). Universidade de Santiago de Compostela. (Anexo 79 de Verba) [URL]
Doval, I., & Sánchez Nieto, M. T. (2026). Parallel corpora Spanish (PaCorES): A collection of multifunctional parallel corpora. Spanish Journal of Applied Linguistics, 39(2), 1–33.Google Scholar logo with link to Google Scholar
(2026). Parallel Corpora Spanish (PaCorES): A collection of multifunctional parallel corpora. RESLA. Revista Española de Lingüística Aplicada / Spanish Journal of Applied Linguistics., 39(2). 〈[URL]>
El-Kishky, A. et al. (2020). CCAligned: A massive collection of cross-lingual web-document pairs. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp.5960–5969). Association for Computational Linguistics. [URL].
Feng, F., et al. (2022). Language-agnostic BERT sentence embedding. Transactions of the Association for Computational Linguistics, 10, 111–125. [URL].
Gale, W. A., & Church, K. W. (1993). A program for aligning sentences in bilingual corpora. Computational Linguistics, 19(1), 75–102. [URL]
Han, L., Wong, D. F., Chao, L. S., He, L., Zhu, L., & Li, S. (2013). A study of Chinese word segmentation based on the characteristics of Chinese. In I. Gurevych, C. Biemann, & T. Zesch (Eds.), Language Processing and Knowledge in the Web (pp.111–118). Springer. Google Scholar logo with link to Google Scholar
Huang, C. R., & Hsieh, S. K. (2015). Chinese lexical semantics: From radicals to event structure. In W. S. -Y. Wang & C. Sun (Eds.), The Oxford Handbook of Chinese Linguistics (pp.290–305). Oxford University Press. 〈>Google Scholar logo with link to Google Scholar
Kay, M., & Röscheisen, M. (1993). Text-translation alignment. Computational Linguistics, 19(1), 121–142. 〈 [URL]>
Koehn, P. (2020). Neural machine translation. Cambridge University Press. Google Scholar logo with link to Google Scholar
Laviosa, S. (1998). Core patterns of lexical use in a comparable corpus of English narrative prose. Meta, 43(4), 557–570. [URL].
Liu, L., & Zhu, M. (2023). Bertalign: Improved word embedding-based sentence alignment for Chinese—English parallel corpora of literary texts. Digital Scholarship in the Humanities, 38(2), 621–634. Google Scholar logo with link to Google Scholar
McCarthy, P. M., & Jarvis, S. (2010). MTLD, vocd-D, and HD-D: A valid and reliable measure of lexical diversity. Behavior Research Methods, 42(2), 381–392. Google Scholar logo with link to Google Scholar
McEnery, T., & Hardie, A. (2012). Corpus linguistics: Method, theory and practice. Cambridge University Press.Google Scholar logo with link to Google Scholar
Nivre, J. et al. (2020). Universal Dependencies v2: An evergrowing multilingual treebank collection. In Proceedings of the Twelfth Language Resources and Evaluation Conference (pp.4034–4043). European Language Resources Association. [URL]
Prokopidis, P. et al. (2016). Parallel Global Voices: A collection of multilingual corpora with citizen media stories. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC 2016) (pp.900–905). European Language Resources Association. [URL].
Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing (pp.3982–3992). Association for Computational Linguistics. [URL].
Schwenk, H. (2018). Filtering and mining parallel data in a joint multilingual space. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics. Volume 2: Short papers (pp.228–234). Association for Computational Linguistics. [URL].
Schwenk, H. et al. (2021a). CCMatrix: Mining billions of high-quality parallel sentences on the web. Transactions of the Association for Computational Linguistics, 9, 49–64. [URL].
(2021b). WikiMatrix: Mining 135M parallel sentences in 1620 language pairs from Wikipedia. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics (pp.1351–1361). Association for Computational Linguistics. [URL].
Thompson, B., & Koehn, P. (2019). Vecalign: Improved sentence alignment in linear time and space. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (pp.1342–1348). Association for Computational Linguistics. [URL].
Tiedemann, J. (2011). Bitext alignment. Morgan & Claypool. Google Scholar logo with link to Google Scholar
(2012). Parallel data, tools and interfaces in OPUS. In N. Calzolari, K. Choukri, T. Declerck, M. U. Doğan, B. Maegaard, J. Mariani, A. Moreno, J. Odijk, & S. Piperidis (Eds.), In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC’12) (pp.2214–2218). European Language Resources Association (ELRA). [URL].
Tsang, Y. K. et al. (2025). A corpus of Chinese word segmentation agreement. Behavior Research Methods, 57, 25. Google Scholar logo with link to Google Scholar
Varga, D. et al. (2005). Parallel corpora for medium-density languages. In Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2005) (pp.590–596). INCOMA. [URL]
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. In Advances in neural information processing systems 30 (NIPS 2017) (pp.5998–6008). Curran Associates. [URL]
Wu, A. (2003). Chinese word segmentation in MSR-NLP. In Proceedings of the Second SIGHAN Workshop on Chinese Language Processing (pp.172–175). Association for Computational Linguistics. [URL].
Xiong, Y. et al. (2019). A fine-grained Chinese word segmentation and part-of-speech tagging corpus for clinical text. BMC Medical Informatics and Decision Making, 19, 66. Google Scholar logo with link to Google Scholar
Zanettin, F. (2012). Translation-driven corpora: Corpus resources for descriptive and applied translation studies. Routledge.Google Scholar logo with link to Google Scholar
Ziemski, M. et al. (2016). The United Nations Parallel Corpus v1.0. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC 2016) (pp.3530–3534). European Language Resources Association. [URL].
Mobile Menu Logo with link to supplementary files background Layer 1 prag Twitter_Logo_Blue