A Target-Centric Survey of Quantization-Aware Training
MCML Authors
Abstract
Abstract
The rapid development of LLMs incurs prohibitive memory footprints and intensive computational demands. Quantization-Aware Training (QAT) techniques have emerged as a promising solution to address these challenges by explicitly simulating quantization effects during model training, yielding low-bit models that achieve accuracy comparable to their full-precision counterparts. In this work, we provide a target-centric survey of QAT, aimed at clarifying both its theoretical foundations and its evolving implementation landscape. We systematically review existing QAT methods through a target-centric taxonomy and synthesize cross-target differences in error characteristics, numerical formats, and strategy transferability. We further summarize QAT evaluation paradigms and discuss challenges in optimization and deployment, outlining potential directions for future research.
inproceedings SZW+26
EMNLP 2026
Conference on Empirical Methods in Natural Language Processing. Budapest, Hungary, Oct 24-29, 2026. To be published. Preprint available.Authors
J. Song • M. Zhao • Z. Wang • Y. Liu • Q. Li • S. Feng • F. Ren • D. Wang • H. SchützeLinks
arXivResearch Area
BibTeXKey: SZW+26