Do LLMs Give Consistent Opinions? Evaluating Response Reliability Under Varying Likert-Scale Formulations in Survey-Style MCQA
MCML Authors
Abstract
Abstract
Survey methodology requires that equivalent question formulations elicit stable responses from the same respondent, regardless of how the Likert-type answer scale is presented. As LLMs are increasingly deployed as synthetic survey respondents, a critical but underexplored question arises: do they satisfy this basic parallel-forms reliability standard when answer scales vary in wording, cardinality, polarity, or option order? We extend the OpinionQA into two stress-test benchmarks, 1,235 subjective questions each paired with five reworded scale variants and six positional permutations, evaluated across four prompt-label formats. Using a four-metric consistency framework, we benchmark eight open-weight models from four families in both base and instruction-tuned configurations. Three findings emerge. First, base models fail basic reliability requirements: they exhibit severe mechanical instability driven by primacy bias, label sensitivity, and formatting failures, including full polarity inversions for the same question. Second, instruction tuning acts as a reliability filter, reducing response variance by approximately 50% and enabling models to treat semantically equivalent scales as equivalent. Third, for instruction-tuned models under randomized option ordering, removing explicit labels paradoxically increases consistency by eliminating symbol-binding conflicts. These findings establish that response reliability is an engineered property of LLMs requiring both alignment training and deliberate prompt design, with concrete implications for survey-like MCQA evaluations.
inproceedings KMH+26
W-NUT @EMNLP 2026
11th Workshop on Natural User-generated Text at the Conference on Empirical Methods in Natural Language Processing. Budapest, Hungary, Oct 24-29, 2026. To be published.Authors
M. Kandlinger • B. Ma • A.-C. Haensch • M. AßenmacherResearch Areas
BibTeXKey: KMH+26