Home | Publications | WPL26

Bi-Text Mining Across German Dialects: On the Role of Synthetic Training Data for Dialect Adaptation

MCML Authors

Abstract

Cross-dialect bi-text mining relies on robust multilingual sentence representations to identify semantically equivalent sentence pairs across languages. While recent multilingual bi-encoder models achieve strong performance on standardized written languages, their behavior on dialectal varieties is largely unknown. In this study, we use Tatoeba to evaluate the performance of four widely-used bi-encoders on dialect-to-standard German translation retrieval, covering German documents and queries written in three dialects: Low German, Bavarian, and Alemannic. Motivated by the lack of resources, we examine the extent to which synthetic translations (from dictionaries and large language models; LLMs) can serve as weak supervision for dialect adaptation. Our results reveal that bi-encoders, when applied in a zero-shot setting, exhibit deficiencies in capturing semantic similarity between German and dialects, while fine-tuning on synthetic data substantially improves their retrieval effectiveness, with larger gains obtained from LLM-translated training data. We further analyze retrieval performance on Bavarian across varying dialect word proportions and observe a drop when dialect words make up more than 60% of the text.

inproceedings WPL26


BUCC @LREC 2026

19th Workshop on Building and Using Comparable Corpora at the 15th International Conference on Language Resources and Evaluation. Palma de Mallorca, Spain, May 11-16, 2026.

Authors

J. Wang • B. PlankR. Litschko

Links

PDF GitHub

Research Area

 B2 | Natural Language Processing

BibTeXKey: WPL26

Back to Top