TL;DR for operators

A translation can be wrong in at least three different ways: the model can arrange words incorrectly, produce the wrong language, or choose the wrong lexical content. The paper examined here finds evidence that these are not merely different visible error types. In the tested multilingual LLMs, they often correspond to separable internal stages.

Across eight of nine main model-dataset comparisons, intervention effects peak in the order syntax first, then syntax plus surface language, then syntax plus language plus lexical content. The model can therefore commit to target-side word order before its intermediate token predictions look like the target language at all.

For multilingual evaluation, the practical implication is to stop treating translation quality as one indivisible score. Grammar, language identity, and lexical content can be tested separately. Mechanistic localization also opens a research path toward targeted debugging or steering, but the paper does not establish production-ready controls: its strongest evidence comes from three models, controlled word-order constructions, and prompt-sensitive experiments.

Translation errors may come from different internal decisions

Suppose a multilingual assistant returns a sentence with the correct meaning but the wrong word order. That failure is operationally different from returning fluent text in the wrong language, or choosing the wrong noun while preserving both language and grammar.

The harder question is whether the model itself makes those decisions separately.

Sonkin and colleagues investigate exactly that possibility in Separating Syntax from Language: A Mechanistic Account of Translation in Multilingual LLMs.1 Their proposed decomposition has three components: target-side syntactic structure, target surface language, and lexical or conceptual content.

The strongest result is not that these components can be described separately after the fact. It is that interventions inside the models suggest they are often resolved at different points during computation.

Across noun-phrase order, SVO/SOV order, and modal-verb order, the main layer-patching comparison shows a syntax → language → lexical-content sequence in eight of nine model-dataset combinations. The exception is Llama 3 on the SVO task, where syntax and language reach their maximum intervention effect at the same layer before lexical content changes later.

That pattern turns translation from one bundled output event into a sequence of partially separable commitments.

The model can have target grammar before target-language words

The most counterintuitive observation concerns intermediate predictions.

Using a technique that projects intermediate hidden states back into token probabilities, the researchers inspect what token-like prediction is visible at successive layers. In Aya Expanse and Llama 3, those intermediate predictions can remain English-like even after their position in the sentence already follows the target language’s word order.

This matters because an English-looking intermediate token is easy to overinterpret. It does not mean the model first constructs a complete English sentence and then translates that sentence into another language.

The evidence supports a different picture: lexical realization can still look English-like while syntactic planning has already shifted toward the target language.

That distinction separates two questions that are often collapsed:

  • What language does the current token representation resemble?
  • What grammatical structure is the model already preparing to produce?

The paper shows that the answers need not change together.

Intervening inside the model turns the pattern into causal evidence

Intermediate predictions alone would establish an interesting correlation, but not whether particular internal states actually carry syntax.

The authors therefore use activation patching. They construct paired translation prompts whose target syntax, surface language, and lexical content are systematically varied. Internal activations from one computation are then inserted into another, and the resulting changes in next-token probabilities reveal which properties are transferred.

The design produces candidate outputs representing all eight combinations of base-versus-plant syntax, language, and content. If patching one layer disproportionately increases the probability of a candidate with the plant prompt’s syntax while retaining the base prompt’s language and content, that layer is carrying information relevant to syntax before the other properties have switched.

The staged pattern survives this stronger test across most of the main comparisons. That is why the paper’s central contribution is mechanistic rather than merely descriptive.

The result is also model-dependent. Aya Expanse and Llama 3 show particularly concentrated syntax effects in individual attention heads: H15.0 in Aya Expanse and H14.25 in Llama 3. mGPT is less cleanly localized, with a more distributed effect and weaker separation between syntax and surface-language influence.

So “there is a syntax head” would be too broad a conclusion. The evidence instead shows that some architectures can contain highly concentrated components with disproportionate influence on target word order.

Syntax can be functionally shared without sharing one geometry

The authors then ask whether the syntax-sensitive components are really encoding grammatical order, or merely carrying language identity that happens to correlate with order.

They patch mean activations across languages that either share or differ in word order. In general, changing to a language with a different word order shifts the relevant syntactic contrasts much more than it shifts language-identity contrasts.

That supports functional separation: the identified components can influence grammatical arrangement without simply transferring the patched language itself.

There is an important qualification. When the researchers compare the mean activation geometry of these heads across languages, languages do not cluster cleanly according to word-order class.

The same functional role therefore does not require a single simple geometric representation shared across languages. Functional similarity and representational similarity are not equivalent.

The paper also reports a smaller asymmetry: when German is the base target language for Aya Expanse and Llama 3, the expected syntax effect disappears. The source package leaves that result unresolved rather than explaining it away.

Multilingual evaluation should split one score into three diagnoses

The paper does not test production monitoring systems, so the following is a Cognaptus inference rather than a demonstrated deployment result.

For multilingual products, evaluation can be organized around at least three failure dimensions:

Failure dimension Operational test What a failure suggests
Syntactic structure Does the output use the expected target-side word order? Grammar or structural planning failure
Surface language Is the realized output consistently in the requested language? Language-selection failure
Lexical content Are the intended entities, predicates, and concepts preserved? Semantic or lexical realization failure

This decomposition is more actionable than aggregate translation accuracy when a team needs to decide what to debug. Two models with similar overall translation scores could fail for very different reasons.

The localization results also suggest a research direction for targeted diagnostics or steering. If a model’s grammatical behavior is concentrated in identifiable internal components, intervention may eventually become more selective than globally modifying the model.

The paper does not establish that such interventions are safe, stable, or commercially useful. It establishes a mechanism worth testing.

The strongest evidence is controlled, not yet broadly naturalistic

The main experiments use 23,924 controlled translation prompts derived from three synthetic datasets covering noun-phrase order, SVO/SOV order, and modal-verb order. That control is an advantage for causal isolation: syntax, surface language, and content can be manipulated with relatively little confounding.

It is also the main generalization boundary.

Only three models are studied. The grammatical phenomena cover a narrow subset of syntax. The intervention method assumes that internal computations can be meaningfully exchanged between runs and are sufficiently modular for patching to reveal interpretable causal effects. The appendix also reports substantial sensitivity to prompt phrasing.

A natural-language check using 27 FLORES-200 sentences provides only limited extension beyond the synthetic setting. It shows a possible staged pattern for mGPT and Aya Expanse, but not clearly for Llama 3, and the authors explicitly treat the sample as too small for substantial conclusions.

The evidence is therefore strong for a mechanistic claim inside the tested settings, but not for a universal theory of multilingual translation.

A translation is not necessarily one decision

The useful shift from this paper is conceptual before it is technical.

A multilingual LLM does not have to resolve grammar, language identity, and lexical content at the same moment. In the tested systems, target syntax is usually committed first, and that commitment can already be present while intermediate lexical predictions remain English-like.

For evaluation teams, this makes translation errors more diagnosable. For interpretability researchers, it provides evidence that multilingual generation contains partially separable internal computations that can sometimes be localized surprisingly tightly.

The next step is not to assume that these components transfer cleanly into production controls. It is to test whether the same separation survives broader grammars, natural prompts, additional model families, and realistic deployment conditions.

Cognaptus: Automate the Present, Incubate the Future.


  1. Mikhail Sonkin and Tanja Baeumel and Daniil Gurgurov and Josef van Genabith and Simon Ostermann (2026). Separating Syntax from Language: A Mechanistic Account of Translation in Multilingual LLMs. arXiv:2609.01356. https://arxiv.org/abs/2609.01356 ↩︎