Beyond Benchmarks: A User-Centric Framework for Evaluating Large Language Models
MCML Authors
Abstract
Abstract
Large language models (LLMs) are increasingly embedded in everyday life, yet most evaluation eforts still emphasize technical performance rather than how users experience these systems. We present UCLLM, a user-centric framework comprising nine dimensions that shift evaluation toward users’ perceptions, needs, and interaction goals. Rather than validating the full framework, we demonstrate how one dimension of UCLLM can be operationalized by conducting a pilot blueprint(N = 24) comparing four widely used LLMs under high- and low-certainty epistemic stances. Lowcertainty responses that ofered alternatives were often perceived as more confdent, reliable, and actionable than high-certainty ones, highlighting a “humility paradox” in user trust. Treating this study as a blueprint illustrates how UCLLM’s conceptual dimensions can be translated into measurable user-centered constructs and experimental procedures. Our fndings highlight the importance of user-centric evaluation and motivate broader empirical work across all UCLLM dimensions to guide the design of more trustworthy and responsible LLMs.
inproceedings AAB+26
CUI 2026
International Conference on Conversational User Interfaces. Bremen, Germany, Jul 21-24, 2026.Authors
Y. Abdrabou • Y. Abdelrahman • E. Bozkir • Y. Mazen • F. Alt • E. KasneciLinks
DOIResearch Area
BibTeXKey: AAB+26