Human-Like Without Engagement: The Measurement Trap in LLM Psychometrics
TL;DR for operators Suppose an AI team administers a standardized values questionnaire to several models, compares their scores with a target population, and selects the model with the smallest statistical distance. That procedure has a failure mode: a model can look close to the human average precisely because it is producing flat or middle-of-the-scale answers rather than responding coherently to the questions. ...