TL;DR for operators

An optimization assistant can generate executable solver programs for many business problems, but expert-verified formulations are expensive and a trajectory-level success signal does not reveal which modeling section was wrong. That makes solver execution useful not just as a final check, but as a potential source of training supervision. :chatgpt-content-reference{index=“0”}

The difficult part is deciding where that feedback should change the model.

The paper behind SOLID—Beyond Verified Answers: Solver-Informed Self-Distillation for Bootstrapping Operations Research Language Models1—finds that distilling privileged solver information across an entire response can improve performance initially and then deteriorate. The same pattern appears even when the privileged LP reference is correct, so better reference information alone does not remove the failure mode.

SOLID instead combines two levels of supervision. Executed solver outcomes provide a coarse group-level signal, while differences in solver artifacts restrict finer corrections to non-consensus trajectories and to modeling sections—variables, objectives, constraints, or generated code—whose structure disagrees with a provisional reference.

For organizations already running optimization solvers, this suggests a lower-supervision post-training architecture: reuse executable artifacts as selective training signals rather than immediately adding a separate learned evaluator. The evidence is promising but bounded. Consensus can be wrong, artifact extraction can fail, gains vary across benchmark-model combinations, and deployment outputs still require independent validation.

More solver information is not automatically better supervision

Suppose an optimization assistant produces a formulation, explanatory reasoning, and executable solver code. A solver can tell you whether the code executes and what objective value it reaches. What it does not directly tell you is which sentence or modeling decision should receive credit.

One response may define the correct variables but encode the wrong constraint. Another may use different prose while producing an equivalent optimization model. A reward assigned to the whole trajectory cannot distinguish these cases.

A natural response is to give the model richer solver information and let it teach itself. The paper tests that idea through whole-response on-policy self-distillation. The model re-scores its own trajectory while receiving privileged LP information unavailable to the deployed student.

The training-dynamics diagnostic is the warning sign. Whole-response distillation improves at first and then deteriorates under both a consensus-derived LP and a diagnostic correct LP. The experiment does not establish style mismatch as the unique cause. It does establish a narrower and operationally consequential result: correcting reference accuracy alone does not remove the failure.

That changes the design question. The choice is not simply whether solver information is informative. It is where that information should be permitted to exert gradient pressure.

SOLID turns execution artifacts into two levels of supervision

SOLID begins by sampling multiple solver-producing trajectories from the current policy for the same unlabeled optimization problem. Their programs are executed, and usable objective values are grouped. The largest objective cluster becomes the group’s provisional consensus.

This is not ground truth. It is a vote among executable outcomes.

From that vote, SOLID derives two different training channels.

First, trajectories receive sequence-level rewards based partly on whether their objective belongs to the majority cluster. Group-relative reinforcement learning therefore pushes consensus trajectories up and non-majority trajectories down.

Second, the system selects an executable LP artifact near the median of the majority cluster as a majority-LP pseudo-reference. Candidate and reference LPs are canonicalized and compared separately across decision variables, objective structure, constraints, and the complete generated LP code.

Those comparisons create the LP-structured mask. Self-distillation is enabled only when a trajectory is outside the majority cluster, and only inside solver-grounded response sections whose LP signatures disagree with the pseudo-reference.

The contextual teacher is not a separately trained, stronger model. It is the same rollout policy, frozen for re-scoring and supplied with additional solver-derived context. SOLID therefore uses the model-solver loop itself to produce both a coarse outcome signal and a localized token-level correction signal.

Outcome consensus anchors the correction

Localization could still be dangerous if privileged context were allowed to override the reinforcement signal arbitrarily. The paper’s gradient analysis shows a more constrained interaction.

In simplified form, the effective token-level advantage is

$$ A^{\mathrm{eff}}_{i,t} = A_i + \beta b_{i,t} \left(e^{\Delta_{i,t}}-1\right), $$

where $A_i$ is the trajectory’s group-relative advantage, $b_{i,t}$ determines whether the token is inside the structural mask, and $\Delta_{i,t}$ measures how much more likely the contextual teacher considers that token than the deployable student does.

For a trajectory with $A_i<0$, contextual information reverses the negative update only if its preference is sufficiently strong. Under the paper’s simplified assumptions, the threshold is

$$ \tau_i=\log\left(1-\frac{A_i}{\beta}\right). $$

The interpretation matters more than the algebra. Consensus remains the anchor. Privileged solver context supplies selective corrections, but weak contextual preferences cannot automatically rescue a trajectory that received a negative group-level signal.

This is mechanism analysis, not a performance guarantee. It characterizes how the two learning signals interact under stated assumptions.

The benchmark gains are meaningful, but concentrated

The main comparisons use matched training budgets and evaluate both a general-purpose Qwen3-4B-Instruct model and the OR-specialized StepORLM across OptMATH, MAMO-Complex, and IndustryOR.

For Qwen3-4B-Instruct on OptMATH, SOLID raises majority accuracy at 64 samples from 30.12% under TTRL to 39.76%. Pass@1 rises from 19.01% to 24.25%, pass@2 from 25.29% to 31.62%, and pass@4 from 31.32% to 38.70%.

The evidence is not uniform. On Qwen3-4B-Instruct, the differences on MAMO-Complex are mixed, while IndustryOR gains over TTRL are small.

The OR-specialized model shows a different pattern. StepORLM with SOLID improves or matches TTRL majority accuracy on all three benchmarks and exceeds it on eight of twelve reported dataset-metric combinations. MAMO-Complex is the clearest case: pass@1 rises from 61.22% to 66.43%, pass@2 from 68.08% to 71.58%, and pass@4 from 72.26% to 74.79%.

The accompanying experiments serve different purposes:

Evidence Likely purpose What it supports What it does not establish
Matched-budget SOLID vs. TTRL Main evidence Localized solver-informed learning can add value beyond outcome-only reinforcement Uniform superiority across models and datasets
Correct vs. majority vs. random LP references Diagnostic High-quality solver context can materially help the model That consensus-derived references are reliably correct
Structured, whole-response, and random masking Ablation Where contextual supervision is applied affects results That the structured mask wins every metric
Training dynamics with whole-response distillation Mechanism diagnostic Correct reference information alone does not prevent deterioration Style mismatch as the unique causal explanation
Anchored-correction derivation Mechanism analysis Explains how group and contextual signals combine Guaranteed improvements in solver accuracy

Reusing the solver changes the supervision architecture

For organizations building optimization assistants, the business question is concrete.

If proprietary problem descriptions are plentiful but expert-built formulations are scarce, can the existing solver stack generate enough structured feedback to support continued post-training?

SOLID suggests that it can contribute more than an end-of-pipeline pass/fail check. Executable objective values can generate group-level supervision, while LP artifacts can identify disagreements in specific modeling components. That may reduce dependence on separately trained process evaluators or large stores of verified formulations.

The affected workflow is narrow but recognizable: teams operating language-model systems that translate natural-language planning, scheduling, allocation, or related OR problems into executable optimization models, and that already have trustworthy solvers available during training.

The potential saving is therefore in supervision infrastructure, not in eliminating validation. The solver becomes part of the training signal, while expert effort can be reserved for cases where consensus and structural checks are insufficient.

Consensus is useful only while its errors remain containable

The majority-derived LP is a pseudo-reference. Multiple rollouts can agree on the same incorrect formulation or objective. A confident consensus is still capable of being wrong.

Artifact extraction introduces another failure surface. If generated LP structures are incomplete or unstable, the structural mask can activate the wrong response sections. The current mask is also binary rather than confidence-weighted.

Applicability is narrower than “self-training for any reasoning model.” SOLID depends on executable programs and comparable solver-readable artifacts. Its results cover three OR benchmarks and two model settings, with gains that vary materially across benchmark-model combinations.

There is also an engineering boundary: the implementation isolates generated solver processes and uses timeouts, but the source package explicitly notes that this is not a hardened security sandbox. Production systems executing model-generated optimization code need stronger controls.

So the paper does not remove the verification problem. It relocates part of the supervision burden from manually verified answers toward executable artifacts whose reliability can be checked, compared, and bounded.

That is the more durable design lesson. When a system can produce machine-verifiable intermediate artifacts, the next training signal need not be another model judging the entire response. It can be a narrower correction attached to the part of the output where the artifact shows disagreement.

Cognaptus: Automate the Present, Incubate the Future.


  1. Rui Zhu and Minglong Cao and Chenyu Zhou and Jianghao Lin and Dongdong Ge (2026). Beyond Verified Answers: Solver-Informed Self-Distillation for Bootstrapping Operations Research Language Models. arXiv:2609.09957. https://arxiv.org/abs/2609.09957 ↩︎