TL;DR for operators

When a translation is already mostly correct, another broad LLM rewrite can create new problems faster than it removes the remaining one. In Wang and Wu’s domain-specific machine-translation study,1 an initial EN→DE post-edit reached 89.93 COMET-22. Six of seven tested second-round editors failed to improve that result; five lowered it, one left it unchanged, and only reusing Claude Sonnet 4.5 produced a small increase to 90.07.

The proposed alternative is architectural rather than simply model-based. One LLM diagnoses specific translation errors using a structured rubric and retrieved Translation Memory examples. A separate LLM receives those diagnoses under an “edit contract” that tells it to repair the identified problem directly while preserving unaffected text.

For localization operations, three findings matter. First, the combination of structured diagnosis, constrained editing, and retrieval produced much stronger first-round results than a basic evaluation prompt in the paper’s main comparison. Second, the best post-editor depended on language; a larger model did not dominate every direction. Third, repeated repair was not reliably beneficial.

The practical interpretation is a pipeline built around diagnose → repair → stop, with model selection validated by language. That interpretation is bounded by the study’s internal e-commerce data, automatic evaluation metrics, absence of reported significance tests, and single-round P95 latency of 59.41 seconds.

Once the main error is fixed, another edit can become the error

A familiar localization problem begins with a translation that is not bad enough to regenerate from scratch. A product phrase uses the wrong domain term. A stylistic convention is slightly off. One expression does not match how the company’s approved translations normally phrase it.

The tempting response is to let a capable LLM revise the translation again. The paper’s second-round experiment shows why that assumption needs testing.

After Claude Sonnet 4.5 produced a first EN→DE repair with a COMET-22 score of 89.93, the authors applied seven possible second-round editors. Gemma 4B reduced the score to 89.18, Claude Sonnet 4 to 89.31, GPT OSS 20B to 89.30, and GPT OSS 120B to 88.07. Gemma 27B left the score unchanged. Claude Sonnet 4.5 was the sole model to improve it, and only by 0.14 points.

This is not evidence that iterative refinement is generally harmful. It is evidence that, in this tested setting, additional editing had diminishing and often negative returns once a strong targeted repair had already occurred. The paper attributes the risk to unnecessary changes and paraphrastic drift.

That changes the deployment question. The relevant control is not only which model should edit? It is also how much freedom should the editor have, and when should the system stop editing?

The architecture reduces editing freedom before it adds model capability

The paper’s central design separates diagnosis from repair.

First, an evaluator examines the source sentence and initial translation. Rather than returning a scalar quality score, it produces structured MQM-style issues: specific spans, error categories, severities, explanations, suggested corrections, and optional links to supporting evidence. MQM here functions as a localized error vocabulary rather than a single overall judgment.

The evaluator also receives semantically similar, previously approved source-target examples retrieved from the company’s Translation Memory. Those examples provide evidence about domain terminology and style without retraining the underlying translation model.

The resulting report is then sent to a different LLM acting as the post-editor. An explicit edit contract constrains that model to make direct, minimal repairs tied to diagnosed spans and to preserve unaffected text.

The mechanism is therefore not “use two LLM calls instead of one.” The relevant change is a narrower generation problem. The evaluator identifies where intervention is justified; retrieval supplies domain evidence; the contract reduces the editor’s permission to rewrite material that was not flagged.

That design is consistent with the first-round prompt experiment. With Claude Sonnet 4.5 fixed as the post-editor, the base GEMBA-MQM prompt produced almost no improvement: EN→DE COMET-22 moved from 86.39 to 86.47. Adding a descriptive rubric and edit contract increased it only slightly further, to 86.57. Adding retrieved domain exemplars produced 89.93.

For EN→CN, the same progression was 87.76 without editing, 87.97 with the base prompt, 88.86 with the rubric and contract, and 90.80 when retrieval was added.

Because these components are introduced cumulatively, the experiment does not isolate the independent causal contribution of each one. It does show that the complete package performed substantially better than the simpler evaluator prompts under the fixed post-editor.

Two stages usually beat one, but not on every metric or language

The architecture comparison gives the paper its strongest support for separating diagnosis from generation.

Against a one-stage GEMBA-MQM setup on EN→DE, the one-stage model improved COMET-22 by only 0.12 points and CometKiwi by 0.22. The two-stage configuration improved them by 3.54 and 1.84 respectively.

Across the broader six-language results, every reported two-stage post-editor improved COMET-22 over the no-post-edit baseline. The pattern is less uniform on CometKiwi. For Claude Sonnet 4.5, for example, COMET-22 rose by 7.92 points on EN→ES, while CometKiwi rose only 0.10; on EN→IT and EN→JP, COMET-22 improved while CometKiwi declined.

There is also a direct exception to any claim of universal two-stage superiority: on EN→FR, the one-stage configuration scored 90.83 COMET-22 versus 90.81 for Claude Sonnet 4.5 in the two-stage system.

The evidence therefore supports a narrower statement: structured two-stage repair produces broad COMET-22 gains in this benchmark and frequently outperforms the tested one-stage alternatives, but the advantage is neither metric-uniform nor universal across language directions.

The best editor changes with the language

A deployment team could still read the architecture correctly and make a second mistake: choosing the largest or most capable-looking post-editor globally.

The benchmark argues against that shortcut.

Claude-family models were strongest on EN→DE, where Claude Sonnet 4.5 reached 89.93 COMET-22. But on EN→CN, Gemma 3 4B achieved the highest reported COMET-22 score at 91.08, above Claude Sonnet 4.5’s 90.80 and Gemma 3 27B’s 90.83.

A separate evaluator ablation on EN→FR does show a clearer capacity pattern. With the P3 prompt and Claude Sonnet 4.5 post-editor fixed, COMET-22 improvement rose from +1.09 with Claude 3.5 Haiku to +2.61 with Sonnet 4, +3.18 with Sonnet 4.5, and +3.62 with Opus 4.5. This is best read as an evaluator-capacity sensitivity test, not proof that model scale predicts post-editor quality in general.

For operators, the resulting unit of model selection is closer to language × role × quality target than simply provider or parameter count.

Translation Memory becomes evidence rather than training data

The business opportunity comes from what the system does not require.

The base translation model is not fine-tuned or retrained. Existing Translation Memory is converted into a retrieval knowledge base, allowing approved historical translations to influence error diagnosis at inference time.

For teams already maintaining substantial TM assets, Cognaptus infers a potential deployment pattern: keep the existing translation engine, insert an evaluator for segments that merit inspection, retrieve nearby approved examples, and allow a constrained post-editor to repair only diagnosed defects.

That pattern could be particularly relevant to catalog, product-description, policy, documentation, and other batch localization workflows where domain consistency matters and processing need not be instantaneous.

It also introduces a more auditable intermediate artifact. A reviewer can inspect which span was flagged, how the problem was classified, what correction was proposed, and which retrieved example supported it. The paper does not test organizational audit outcomes directly, so that is an operational inference rather than an empirical result.

Latency and evaluation evidence keep this in the batch-workflow category

The strongest boundary is deployment speed. On 50 measured segments, the evaluator averaged 14.34 seconds, the post-editor 3.30 seconds, and the complete first round 17.64 seconds. End-to-end median latency was 10.90 seconds and P95 latency reached 59.41 seconds.

That profile is difficult to reconcile with instantaneous interactive translation without substantial engineering changes.

The evidence base also remains narrower than the architecture’s potential scope. All primary data come from one internal e-commerce Translation Memory, with English as the source language and six target directions. The paper reports 5,000 held-out test segments per direction, but its conclusions rely mainly on COMET-22 and CometKiwi point estimates without reported significance tests or uncertainty intervals.

The source package also records an unresolved documentation issue: the paper’s abstract claims strong agreement with human MQM annotations and human editor preferences, but the available body and appendices do not provide the human-evaluation protocol, sample size, agreement statistic, or preference results needed to assess that claim. The automatic-metric findings therefore carry more documented support than the human-validation claim.

The stopping rule may be as important as the editing model

The paper’s most transferable idea is not that every translation system needs two specific LLMs. It is that correction can benefit from separating evidence that a change is justified from permission to generate a change.

Within this e-commerce benchmark, the strongest workflow first localized the defect, grounded the diagnosis in approved domain examples, constrained the repair, and usually stopped after one pass. More model intervention did not reliably create more quality.

For localization teams, that suggests a concrete evaluation agenda: test the evaluator and editor separately, validate model-language pairings instead of assuming a universal winner, and measure whether another editing round repairs residual errors or merely reopens text that was already good enough.

The next optimization target may not be a stronger editor. It may be a better decision about when editing should end.

Cognaptus: Automate the Present, Incubate the Future.


  1. Ji Hun Wang and Siyu Wu (2026). Diagnose, Then Repair: A Two-Stage MQM-Guided Post-Editing Framework for Domain-Specific Machine Translation. arXiv:2609.22793. https://arxiv.org/abs/2609.22793 ↩︎