Home | Publications | ROF26a

Evaluating Latin and Ancient Greek Sentence Alignment Through Parallel Sentence Mining

MCML Authors

Abstract

Cross-lingual detection of intertextuality and translation in Latin and Ancient Greek through computational approaches is of great interest for classical studies. While several systems exist for parallel sentence detection, based on general multilingual or specific models for Latin–Ancient Greek, they have not been compared against each other. Therefore, we present a synthetic benchmark to evaluate the performance of language models regarding crosslingual Ancient Greek and Latin parallel sentence mining. We first compare six language models to encode sentences and then further improve the cross-lingual alignment through post-processing, fine-tuning, and knowledge distillation. We find that the whitening transformation in combination with knowledge distillation provides excellent results. Specifically, SPhilBERTa, a trilingual language model for Ancient Greek and Latin, benefits the most from the improvements and achieves a notable mining score of 97.6 on our benchmark.

inproceedings ROF26a


NLP4DH 2026

6th International Conference on Natural Language Processing for Digital Humanities. San Diago, CA, USA, Jul 04, 2025.

Authors

S. Reichbauer • S. OkabeA. Fraser

Links

DOI GitHub

Research Area

 B2 | Natural Language Processing

BibTeXKey: ROF26a

Back to Top