In:Exploring and Exploiting Parallel Corpora of Spanish: The PaCorES corpus collection in language and translation education and research
Edited by Irene Doval and M. Teresa Sánchez Nieto
[Studies in Corpus Linguistics 129] 2026
► pp. 16–40
The Chinese–Spanish parallel corpus (PaCheS)
A multifunctional resource for typologically distant languages
This content is being prepared for publication; it may be subject to changes.
Abstract
Parallel corpora are a key resource for contrastive
linguistics, translation studies, and natural language processing. However,
for the Chinese–Spanish language pair, existing resources are often
scattered and poorly integrated, domain-specific, or rely on English as a
pivot language, introducing structural and stylistic biases. This paper
presents the Chinese–Spanish Parallel Corpus (PaCheS), a multifunctional bilingual resource designed to
address these limitations. The core of PaCheS consists of approximately 15
million words from 49 contemporary literary works, ensuring high-quality,
direct human translations with balanced bidirectionality. To ensure high
precision in the alignment despite the structural distance between the two
languages, the corpus employs a two-stage alignment workflow: automatic
alignment using Bertalign followed by semantic validation via LaBSE
embeddings and manual revision. Additionally, the resource is supplemented
by over 2 million bisegments from institutional and web-mined sources which
have undergone rigorous filtering and normalization. All materials are
accessible via a web interface which supports metadata filtering and
advanced queries, making PaCheS a valuable resource for research,
translation, and language learning.
Article outline
- 1.Introduction
- 2.Related work
- 3.Corpus design and composition
- 4.Textual preprocessing and segmentation
- 5.Alignment and manual revision
- 6.System architecture and interface
- 7.Conclusions and outlook
- Author queries
Notes References
References (38)
Atkins, B. T. S., & Rundell, M. (2008). The
Oxford guide to practical
lexicography. Oxford University Press.
Baker, M. (1993). Corpus
linguistics and translation studies: Implications and
applications. In M. Baker, G. Francis, & E. Tognini-Bonelli (Eds.), Text
and technology: In honour of John
Sinclair (pp.233–250). John Benjamins.
Bañón, M. et al. (2020). ParaCrawl:
Web-scale acquisition of parallel
corpora. In Proceedings
of the 58th Annual Meeting of the Association for Computational
Linguistics (pp.4555–4567). Association for Computational Linguistics.
Bernardini, S., & Castagnoli, S. (2008). Corpora
for translator education and professional
practice. In Proceedings
of the 8th International Conference on Terminology and Artificial
Intelligence (TIA).
Bowker, L., & Pearson, J. (2002). Working
with specialized language: A practical guide to using
corpora. Routledge.
Braune, F., & Fraser, A. (2010). Improved
sentence alignment using maximum entropy
classifiers. In Proceedings
of the 2010 Conference on Empirical Methods in Natural Language
Processing (pp.170–178). Association for Computational Linguistics.
Brown, P. F. et al. (1991). Aligning
sentences in parallel
corpora. In Proceedings
of the 29th Annual Meeting of the Association for Computational
Linguistics (pp.169–176). Association for Computational Linguistics.
Cettolo, M., Girardi, C., & Federico, M. (2012). WIT³:
Web inventory of transcribed and translated
talks. In Proceedings
of the 16th Conference of the European Association for Machine
Translation
(EAMT) (pp.261–268). European Association for Machine Translation.
Doval, I., & Jiménez, T. (2020). Multifuncionalidad
de los corpus paralelos, ejemplificado en el corpus alemán/español
PaGeS. In M. Blanco, H. Olbertz, & V. Vázquez Rozas (Eds.), Corpus
y construcciones. Perspectivas
hispánicas (pp.305–322). Universidade de Santiago de Compostela. (Anexo
79 de Verba) [URL]
Doval, I., & Sánchez Nieto, M. T. (2026). Parallel
corpora Spanish (PaCorES): A collection of multifunctional parallel
corpora. Spanish Journal of Applied
Linguistics, 39(2), 1–33.
(2026). Parallel
Corpora Spanish (PaCorES): A collection of multifunctional parallel
corpora. RESLA. Revista Española de
Lingüística Aplicada / Spanish Journal of Applied
Linguistics., 39(2). 〈[URL]>
El-Kishky, A. et al. (2020). CCAligned:
A massive collection of cross-lingual web-document
pairs. In Proceedings
of the 2020 Conference on Empirical Methods in Natural Language
Processing
(EMNLP) (pp.5960–5969). Association for Computational Linguistics. [URL].
Feng, F., et al. (2022). Language-agnostic
BERT sentence embedding. Transactions
of the Association for Computational
Linguistics, 10, 111–125. [URL].
Gale, W. A., & Church, K. W. (1993). A
program for aligning sentences in bilingual
corpora. Computational
Linguistics, 19(1), 75–102. [URL]
Han, L., Wong, D. F., Chao, L. S., He, L., Zhu, L., & Li, S. (2013). A
study of Chinese word segmentation based on the characteristics of
Chinese. In I. Gurevych, C. Biemann, & T. Zesch (Eds.), Language
Processing and Knowledge in the
Web (pp.111–118). Springer.
Huang, C. R., & Hsieh, S. K. (2015). Chinese
lexical semantics: From radicals to event
structure. In W. S. -Y. Wang & C. Sun (Eds.), The
Oxford Handbook of Chinese
Linguistics (pp.290–305). Oxford University Press. 〈>
Kay, M., & Röscheisen, M. (1993). Text-translation
alignment. Computational
Linguistics, 19(1), 121–142. 〈 [URL]>
Laviosa, S. (1998). Core
patterns of lexical use in a comparable corpus of English narrative
prose. Meta, 43(4), 557–570. [URL].
Liu, L., & Zhu, M. (2023). Bertalign:
Improved word embedding-based sentence alignment for Chinese—English
parallel corpora of literary
texts. Digital Scholarship in the
Humanities, 38(2), 621–634.
McCarthy, P. M., & Jarvis, S. (2010). MTLD,
vocd-D, and HD-D: A valid and reliable measure of lexical
diversity. Behavior Research
Methods, 42(2), 381–392.
McEnery, T., & Hardie, A. (2012). Corpus
linguistics: Method, theory and
practice. Cambridge University Press.
Nivre, J. et al. (2020). Universal
Dependencies v2: An evergrowing multilingual treebank
collection. In Proceedings
of the Twelfth Language Resources and Evaluation
Conference (pp.4034–4043). European Language Resources Association. [URL]
Prokopidis, P. et al. (2016). Parallel
Global Voices: A collection of multilingual corpora with citizen
media
stories. In Proceedings
of the Tenth International Conference on Language Resources and
Evaluation (LREC
2016) (pp.900–905). European Language Resources Association. [URL].
Reimers, N., & Gurevych, I. (2019). Sentence-BERT:
Sentence embeddings using Siamese
BERT-networks. In Proceedings
of the 2019 Conference on Empirical Methods in Natural Language
Processing (pp.3982–3992). Association for Computational Linguistics. [URL].
Schwenk, H. (2018). Filtering
and mining parallel data in a joint multilingual
space. In Proceedings
of the 56th Annual Meeting of the Association for Computational
Linguistics. Volume 2: Short
papers (pp.228–234). Association for Computational Linguistics. [URL].
Schwenk, H. et al. (2021a). CCMatrix:
Mining billions of high-quality parallel sentences on the
web. Transactions of the Association
for Computational
Linguistics, 9, 49–64. [URL].
(2021b). WikiMatrix:
Mining 135M parallel sentences in 1620 language pairs from
Wikipedia. In Proceedings
of the 16th Conference of the European Chapter of the Association
for Computational
Linguistics (pp.1351–1361). Association for Computational Linguistics. [URL].
Thompson, B., & Koehn, P. (2019). Vecalign:
Improved sentence alignment in linear time and
space. In Proceedings
of the 2019 Conference on Empirical Methods in Natural Language
Processing and the 9th International Joint Conference on Natural
Language Processing
(EMNLP-IJCNLP) (pp.1342–1348). Association for Computational Linguistics. [URL].
(2012). Parallel
data, tools and interfaces in
OPUS. In N. Calzolari, K. Choukri, T. Declerck, M. U. Doğan, B. Maegaard, J. Mariani, A. Moreno, J. Odijk, & S. Piperidis (Eds.), In Proceedings
of the Eighth International Conference on Language Resources and
Evaluation
(LREC’12) (pp.2214–2218). European Language Resources Association (ELRA). [URL].
Tsang, Y. K. et al. (2025). A
corpus of Chinese word segmentation
agreement. Behavior Research
Methods, 57, 25.
Varga, D. et al. (2005). Parallel
corpora for medium-density
languages. In Proceedings
of the International Conference on Recent Advances in Natural
Language Processing (RANLP
2005) (pp.590–596). INCOMA. [URL]
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention
is all you
need. In Advances
in neural information processing systems 30 (NIPS
2017) (pp.5998–6008). Curran Associates. [URL]
Wu, A. (2003). Chinese
word segmentation in
MSR-NLP. In Proceedings
of the Second SIGHAN Workshop on Chinese Language
Processing (pp.172–175). Association for Computational Linguistics. [URL].
Xiong, Y. et al. (2019). A
fine-grained Chinese word segmentation and part-of-speech tagging
corpus for clinical text. BMC Medical
Informatics and Decision
Making, 19, 66.
Zanettin, F. (2012). Translation-driven
corpora: Corpus resources for descriptive and applied translation
studies. Routledge.
Ziemski, M. et al. (2016). The
United Nations Parallel Corpus
v1.0. In Proceedings
of the Tenth International Conference on Language Resources and
Evaluation (LREC
2016) (pp.3530–3534). European Language Resources Association. [URL].
