Home | Publications | JLP+26

MaLA: A Corpus and Data Mix for Massive Language Adaptation of Large Language Models

MCML Authors

Link to Profile Hinrich Schütze

Hinrich Schütze

Prof. Dr.

Core PI

Abstract

In this work, we present the MaLA (Massively multilingual Language Adaptation) suite, a comprehensive collection of open-access resources designed to advance language technology across 546 languages. The primary contribution is the MaLA corpus, a 74-billion-token dataset specifically engineered for the continual pre-training of large language models (LLMs). To ensure high utility and robustness for various tasks, we introduce a diverse data mix that balances underrepresented languages with curated high-resource data. This mix includes: (1) structured scientific literature; (2) literary archives; (3) multilingual instruction sets; and (4) programming data. We document the corpus provenance and sampling strategies—upsampling low-resource and downsampling high-resource languages— to reduce forgetting (particularly in code generation) while expanding language capacity. Leveraging this resource, we release EMMA-500, a Llama2-based model optimized for cross-lingual transfer and language adaptability. Beyond the model weights, we provide the full suite of scripts, processing pipelines, and model generations to support reproducibility and further experimentation in multilingual indexing and retrieval. We release the MaLA suite, including the MaLA corpus, EMMA-500 model weights, scripts, and model generations, under open licenses to provide the community with the testbeds and tools necessary to evaluate and improve global information access.

inproceedings JLP+26


COLM 2026

Conference on Language Modeling. San Francisco, CA, USA, Oct 06-09, 2026.

Authors

S. Ji • Z. Li • J. Paavola • P. Lin • P. Chen • D. O'Brien • H. Luo • H. Schütze • J. Tiedemann • B. Haddow

Links

URL

Research Area

 B2 | Natural Language Processing

BibTeXKey: JLP+26

Back to Top