TL;DR for operators
Choosing between a and an depends on how the next word sounds, even when spelling misleads: a university but an hour.
Kim and Lee find that a single sound-related direction learned from ordinary English cases generalizes to these spelling-sound exceptions, reaching 100.0% accuracy for Llama, 95.1% for Qwen, and 98.0% for Gemma.1 More importantly, this feature is not merely decodable. When researchers hold a synthetic nonce embedding fixed and change only its position along that direction, the models shift between preferring a and an.
The model also appears to anticipate the upcoming word before generating it. FutureLens recovers the future trigger substantially better in a/an contexts than in matched controls, while inverse steering of the relevant sound feature flips many article choices.
For evaluation teams, the operational lesson is bounded but important: what a model can explicitly say about language need not be the representation that actually controls generation. When that distinction matters, intervention-based tests can provide evidence that direct questioning alone does not.
The article depends on a sound the model has not generated yet
English indefinite articles expose a useful mechanistic puzzle. A model must choose a or an before it emits the noun that determines the choice.
If the behavior were primarily an orthographic shortcut, difficult cases should expose it. Hour begins with a consonant letter but a vowel sound; university begins with a vowel letter but a consonant sound.
The researchers fit a linear classifier using ordinary cases where spelling and sound agree, then tested it only on spelling-sound exceptions. The resulting direction classified the exception tokens with 100.0% accuracy for Llama, 95.1% for Qwen, and 98.0% for Gemma.
That matters because the test deliberately removes an easy explanation for the classifier: first-letter spelling.
The authors then project this first direction out of the embeddings and ask whether another classifier can recover independent phonological information. Residual linear performance falls into the paper’s estimated chance regime, and nonlinear MLP probes on the residual representation remain near chance. Bootstrap and cross-validation checks also show that the original direction is highly stable.
The evidence therefore supports a relatively narrow claim: for the tested English a/an distinction in these models, the conditioning feature is concentrated along one dominant direction in representation space. It does not establish that phonology in general is one-dimensional.
Causality appears when the word does not exist
High probe accuracy still leaves a major ambiguity. Information can be decodable from a representation without being used to produce an output.
The paper addresses that with a token-level version of a wug test. Instead of giving the model a familiar English word, the researchers construct synthetic nonce-token embeddings. For each nonce base, they first remove its component along the identified sound-related direction. They then add controlled amounts of that direction back while keeping the remaining embedding unchanged.
Across 500 nonce bases per model and 30 direction values, changing this single coordinate moves the models between preference for a and preference for an.
This is the stronger result. The manipulated tokens have no ordinary lexical history from which the model could retrieve a memorized article pairing, and the experiment changes the targeted feature while holding the rest of each nonce representation fixed.
The paper calls the resulting behavior rule-like generalization, but that wording should not be inflated into a claim that the models execute a symbolic phonological rule. The experiment demonstrates productive causal use of an internal feature. It does not identify a symbolic rule engine.
The model forecasts the word before choosing its article
Causal use creates the next puzzle: how can the article decision depend on the sound of a token that has not yet been generated?
The researchers use FutureLens, a learned linear map that asks what information about a later token is already recoverable from an earlier hidden state. They train it at the position where the model is predicting the article and test whether that state already contains a usable forecast of the following trigger word.
In late layers, the forecast is much stronger for a/an contexts than for matched random-token controls. Across the three main models, peak top-five agreement with the actual upcoming trigger reaches roughly 90%.
Forecasting alone would again be correlational. The authors therefore make the FutureLens map orthogonal so that its transpose can map the previously identified phonological direction back into the current hidden-state space. They replace only that component with the mean value associated with the competing article class.
Late-layer steering flips roughly 84–90% of article choices in the three main models.
A larger Llama-3.3-70B-Instruct replication, which does not rely on tied input and output embeddings, follows the same pattern: peak top-five trigger agreement reaches 81.8% versus 45.1% for controls, while steering at layer 78 flips an average of 86.1% of held-out article choices.
Taken together, the experiments support a specific mechanistic account: the model anticipates the upcoming lexical item, its forecast carries the phonological feature relevant to article selection, and that feature can causally influence the current decision.
FutureLens is still an auxiliary probe, however. The experiments expose a manipulable information pathway; they do not trace the full circuit implementing the computation.
Generation competence is not the same as sound explanation
The causal pattern is not confined to English. The paper applies analogous nonce-token interventions to Korean object-particle alternation, Turkish locative vowel harmony, Italian indefinite articles, and French definite-article contraction. The systems differ in language family, morphology, and whether the conditioning trigger appears before or after the allomorph.
For French and Italian, the researchers also repeat the forecasting-and-steering analysis. Peak trigger agreement reaches 94.8% versus 57.0% for the French control and 91.4% versus 36.9% for Italian. Final-layer average steering flip rates are 74.9% and 66.3%, respectively.
Yet the same English direction behaves differently when the model is explicitly asked whether a nonce word begins with a consonant or vowel sound. Effects on yes/no judgments are weaker, inconsistent, and sometimes non-monotonic.
That dissociation changes how the results should be used. A language model may possess information that reliably influences ordinary generation while failing to expose the same information cleanly when prompted to describe or classify it.
Evaluation should test the mechanism that controls the output
For teams deciding whether a language-sensitive model is ready for production, the paper suggests a practical addition to benchmark design.
| Evaluation question | Suitable evidence | What it adds |
|---|---|---|
| Can the model answer a linguistic question? | Direct prompts and classification benchmarks | Measures elicited reporting behavior |
| Is a linguistic feature internally represented? | Probing, including exception tests | Measures recoverable information |
| Does that feature actually influence generation? | Controlled representation interventions | Tests causal relevance |
| Is future information routed into an earlier decision? | Forecasting plus counterfactual steering | Tests a proposed decision pathway |
The Cognaptus inference is not that every deployment needs mechanistic interpretability tooling. It is narrower. When a production decision depends on language-specific structure and failures are costly to diagnose, direct-question performance alone may be an incomplete assurance signal.
A multilingual interface, language-learning system, assessment tool, or other language-sensitive application may therefore benefit from tests that perturb the representation thought to control generation and observe whether the intended behavior changes. That is particularly relevant when direct metalinguistic prompts and ordinary generation disagree.
The boundary is selected mechanisms, not phonology in general
The paper’s scope is substantial but defined. It studies decoder-only text models with conventional tokenization, not tokenization-free systems or multimodal models that receive speech directly. Its phonological labels are inferred from corpus allomorph patterns rather than observed pronunciations. The multilingual cases are selected systems rather than a representative survey of the world’s languages.
Most importantly, the proposed mechanism is inferred through probes and targeted interventions. It demonstrates causal leverage over model behavior without identifying every circuit-level computation that produces that behavior.
Within those boundaries, the paper advances the evidence beyond “the information is somewhere in the model.” For the tested allomorph decisions, the relevant feature can be isolated, changed, routed from a forecast of a future token, and shown to alter what the model generates.
That also explains why asking the model what it knows is not always the right measurement. The more consequential question may be what information the generation process actually uses.
Cognaptus: Automate the Present, Incubate the Future.
-
Sangwoo Kim and Sangah Lee (2026). How Do Language Models Represent and Use Phonological Information for Allomorph Selection?. arXiv:2609.04708. https://arxiv.org/abs/2609.04708 ↩︎