TL;DR for operators

Suppose an AI team administers a standardized values questionnaire to several models, compares their scores with a target population, and selects the model with the smallest statistical distance. That procedure has a failure mode: a model can look close to the human average precisely because it is producing flat or middle-of-the-scale answers rather than responding coherently to the questions.

Andersen and Dichas demonstrate this problem with six open-weight instruction-tuned models evaluated on the Norwegian Moral Foundations Questionnaire, using 1,282 Norwegian respondents as the human reference.1 Only three models pass the study’s baseline attention criterion. Some models that fail it nevertheless appear superficially close to human foundation averages.

For evaluation teams, that changes the order of operations. Establish questionnaire engagement before interpreting profile similarity. Then check whether apparent alignment survives changes in presentation. The study also shows why steering results need the same discipline: a Nordic respondent persona moves all three attention-passing models substantially closer to the human reference, while the tested one-pair activation intervention fails to selectively control individual moral foundations.

A close score can be a measurement failure

A psychometric comparison normally assumes that both sides of the comparison are responding to the instrument in a meaningful way. With LLMs, that assumption cannot be taken for granted.

The study makes the problem visible with two attention-check items embedded in the questionnaire. Downstream moral-profile interpretation is gated on whether a model reaches the 92.2% joint pass-rate threshold observed in the human reference sample.

At forward-scale baseline, Gemma 4-E4B-it, Qwen3-14B, and Qwen3-8B pass. Qwen2.5-1.5B-Instruct and both NorMistral models fail. Under the reversed response scale, only Gemma 4 and Qwen3-14B still pass at baseline.

The significant result is not merely that some models fail a test. It is that failure can coexist with apparently plausible averages. Flat or central-tendency responses pull foundation scores toward the middle of the scale, which can place an inattentive model near human means without evidence that it tracked the distinctions among questionnaire items.

That makes numerical proximity conditional evidence. A dashboard that reports only “distance from target population” can reward a model for producing an uninformative response pattern.

For audit teams, model-selection teams, or developers validating an alignment persona, an engagement check therefore belongs upstream of the similarity metric. Otherwise the evaluation can confuse absence of differentiation with human-like values.

Similarity should account for the shape of human variation

For models that do pass the engagement screen, the paper compares their five moral-foundation scores with the Norwegian reference population.

The human sample is not equally variable in every direction. Care and fairness are relatively strongly correlated, while loyalty, authority, and purity also form a correlated group. The authors therefore use squared Mahalanobis distance rather than treating each foundation as an independent axis:

$$ d^{2}=(\mathbf{x}-\boldsymbol{\mu})^{\top}\Sigma^{-1}(\mathbf{x}-\boldsymbol{\mu}) $$

Here, $\mathbf{x}$ is the model’s five-foundation profile, $\boldsymbol{\mu}$ is the Norwegian human centroid, and $\Sigma$ is the human covariance matrix. Deviations in directions where human respondents naturally vary little receive more weight than deviations along common patterns of human variation.

This is a better target metric than five unweighted score gaps, but it does not solve the engagement problem. Nor does a low $d^2$ establish that the model reproduces human psychometric structure. Each questionnaire item is queried independently, so the experiment cannot test whether model responses reproduce the human cross-item covariance pattern.

The metric answers a narrower question: how close is this five-dimensional mean profile to the human centroid, given the covariance geometry observed in the reference sample?

A persona changes both the profile and whether a profile appears

Once the analysis is restricted to models that demonstrate engagement, prompt-level steering produces the paper’s clearest intervention result.

The authors first use individualizing and binding personas as a sanity probe. The three attention-passing models move strongly in the theoretically expected directions, with persona contrast scores of 32.6 for Gemma 4-E4B-it, 21.9 for Qwen3-14B, and 20.0 for Qwen3-8B. Two hard failers remain nearly flat.

The main treatment is nordic_b, a respondent persona combining a Norwegian demographic anchor with a welfare-state framing without specifying the desired questionnaire answers. Relative to baseline, squared Mahalanobis distance falls:

Model Baseline $d^2$ nordic_b $d^2$ Reduction
Gemma 4-E4B-it 3.65 2.05 44%
Qwen3-14B 6.26 2.37 62%
Qwen3-8B 11.59 2.69 77%

The Qwen3-8B result is especially revealing. Its reversed-scale attention is only about 30% at bare baseline but rises to about 98% under nordic_b. The persona is therefore doing more than shifting an already stable numerical profile. It changes whether the model engages coherently with the altered questionnaire format.

For production evaluation, system prompts and demographic framing should be treated as part of the measured system. A psychometric profile obtained under one persona is not automatically a context-free property of the underlying model.

Moving scores is not the same as controlling a construct

The paper then tests a more internal intervention: one-pair Activation Addition, or ActAdd.

A steering direction is formed from the difference between hidden activations elicited by one positive and one negative contrastive prompt:

$$ \mathbf{v}=\bar{\mathbf{h}}_{\ell}^{+}-\bar{\mathbf{h}}_{\ell}^{-} $$

During questionnaire inference, that direction is injected into the residual stream:

$$ \mathbf{h}_{\ell}\leftarrow\mathbf{h}_{\ell}+\alpha\mathbf{v} $$

The main experiments use layer 15 and mostly $\alpha=5$.

This configuration does move scores, but it does not selectively control the intended moral foundation. For Gemma 4, a positive loyalty intervention raises loyalty by 5.7 points, yet authority rises by 8.4 and purity by 5.9. Reversing the coefficient does not produce a symmetric loyalty decrease. For Qwen3-8B and Qwen3-14B, the intervention often reduces differentiation among foundations.

That is evidence against this specific steering configuration, not activation steering as a category. The paper uses one contrastive pair, a fixed layer across models, a coefficient chosen from limited tuning, and author-written contrasts that may contain additional semantic or sentiment signals.

The operational standard should therefore be selectivity, not movement. If an internal steering intervention changes the target score but perturbs adjacent constructs as much or more, the measured change does not demonstrate control over the intended property.

What an evaluation pipeline should change

The study supports a more disciplined sequence for organizations using questionnaires to compare models or validate behavioral controls:

  1. Test engagement before scoring alignment. Reject or separately classify response modes that fail content-sensitive checks.
  2. Measure profile distance only among interpretable runs. Covariance-aware metrics can improve comparison, but they do not rescue invalid measurements.
  3. Perturb the presentation. Reverse scales, vary answer formatting, or otherwise test whether the measured construct survives superficial changes.
  4. Re-run engagement checks after steering. A persona can alter not only scores but the model’s ability to participate coherently in the evaluation.
  5. Test steering selectivity. Track intended and neighboring dimensions rather than celebrating movement on one headline metric.

These steps matter most when questionnaire results feed an operational decision: approving a system prompt, selecting a model for a regional market, reporting an alignment KPI, or deploying an activation-level control.

The evidence supports a validation lesson, not a stable moral portrait

Several boundaries prevent a broader conclusion.

Only six open-weight instruction-tuned models are tested, with one run per model-condition cell. The main Nordic persona is evaluated against the same Norwegian reference population that informed its demographic design, so the reported 44–77% reductions are not held-out evidence of generalization to Norwegians more broadly.

The activation experiments are also non-exhaustive: one-pair ActAdd is tested with a fixed layer and limited coefficient search. Poor selectivity in that configuration does not establish that richer activation-steering methods cannot separate moral foundations.

Finally, the experiment measures context-conditioned questionnaire behavior. Because items are independently queried, even a model whose five foundation means align closely with the Norwegian centroid has not been shown to reproduce the joint psychometric structure observed among human respondents.

The durable result is therefore methodological. Human-like scores become interpretable only after the evaluation establishes that the model is actually responding to the instrument. Once that condition is enforced, the study shows that respondent framing can materially alter both the resulting profile and the existence of coherent engagement itself. For teams using psychometric evaluation in AI governance, that is a reason to treat measurement validity as part of alignment testing rather than as a preprocessing detail.

Cognaptus: Automate the Present, Incubate the Future.


  1. Hans Andersen and David Dichas (2026). Steering LLMs Responses Towards Moral Foundations on the Norwegian MFQ-30. arXiv:2609.21636. https://arxiv.org/abs/2609.21636 ↩︎