TL;DR for operators
A grid operator may see topology, live measurements, and an incident narrative all point to the same diagnosis. The decision is not only whether the answer is correct, but whether the model relied on evidence that the diagnostic task permits it to use.
In the study, shortcut incident text produced a mean signed utility effect of +0.062 even though its preregistered engineering importance was zero. The model therefore became more accurate by using evidence that should not have determined the answer. Accuracy and a plausible explanation cannot reveal that divergence on their own.
The proposed deployment gate first defines which evidence sources should matter, then removes them one at a time to observe what actually changes the answer. This comparison measures behavioral grounding rather than accepting the model’s stated rationale. A failed response must be corrected and independently re-ablated before approval. The result can guide operational clearance, checkpoint procurement, physics-tool or human escalation, and the audit record used to defend a recommendation after an incident.
A correct diagnosis can still rely on the wrong evidence
A grid operator may receive the same diagnosis from three sources: network topology, live operating measurements, and an incident narrative. Those sources can agree, but they do not necessarily deserve equal authority. A narrative may contain a label that happens to match the target event without providing the physical evidence required for that diagnostic task.
That failure appears directly in Task-Conditional Faithfulness Auditing of Multimodal LLMs for Grid Diagnosis. Across models, retaining shortcut incident text improved the score relative to removing it, even though the task contract said that modality should not determine the answer.
This changes the approval decision. A model can be accurate because it exploits an outcome-bearing cue, while an engineer would reject that cue as an admissible basis for action. Accuracy measures utility under the test distribution; it does not establish that the evidence pathway is legitimate.
The audit separates claims, behavior, and engineering requirements
The paper preregisters a task contract before testing. It fixes the query, parsers, admissible evidence, prohibited outcome-bearing fields, task score, intervention rules, thresholds, and engineering-reference weights. Each task therefore begins with an explicit view of which modalities should matter.
The audit compares three reliance vectors:
| Audit view | What it measures | Operational use |
|---|---|---|
| Self-reported reliance | What the model says it used | Detect unsupported or misleading explanations |
| Behavioral reliance | How much the parsed answer changes when each modality is removed | Measure marginal dependence under the registered intervention |
| Engineering reference importance | Which modalities the task contract says should matter | Define acceptable evidence use before evaluation |
Three similarities keep distinct questions from collapsing into one score. Claimed alignment compares the model’s report with engineering requirements. Behavioral grounding compares observed ablation behavior with those requirements. Explanation fidelity compares the model’s report with its observed reliance pattern.
The single-modality ablations are the main evidentiary mechanism. The query, target, nonintervened values, response schema, decoding settings, and physical invariants remain fixed while topology, measurements, or incident text is neutralized. Sham edits and positive controls serve as intervention-validity checks: they screen formatting artifacts and confirm that supposedly relevant modalities can perturb the output.
The causal boundary is explicit. These ablations estimate marginal behavioral dependence under a registered intervention. They do not reveal the model’s latent internal mechanism.
Stress testing exposes explanation mismatch
On jointly estimable stressed observations, self-reported alignment exceeded intervention-derived grounding by 0.024, with a 95% interval of [0.011, 0.036]. The gap is modest but systematic: the models described their evidence use as more aligned with engineering requirements than their behavior supported.
The modality effects add context. In aligned regimes, topology and measurements had mean signed effects of +0.552 and +0.458. Removing them materially reduced performance, consistent with their registered importance. Shortcut incident text also added utility despite having no registered importance. An accuracy-only benchmark would reward that behavior; the task-conditional audit flags it.
Conflicting text does not support a universal failure claim. Its mean effect was near zero overall, with positive and negative results across model-system pairs. The evidence indicates heterogeneous model behavior, not uniformly harmful narratives.
Correction is accepted only after another intervention test
The paper converts the audit from a report into a control loop. Depending on the diagnosed failure, correction may impose evidence-citation constraints, mask prohibited fields, link generation to evidence, invoke power-flow or contingency tools, use counterfactual prompting, or escalate to a human.
The repaired answer is then independently ablated again. A revised explanation can sound more convincing without changing the evidence driving the answer, so textual improvement alone cannot satisfy the gate.
Across 89 eligible failed stressed cases, performance increased from 0.452 to 0.893. The paired gain was 0.442, with a 95% interval of [0.324, 0.543], based on 75 valid performance pairs. Behavioral grounding increased from 0.816 to 0.979, a paired gain of 0.164 with a 95% interval of [0.124, 0.213], based on 66 valid grounding pairs.
The denominators constrain the claim. Missing pairs were not imputed, so the gains apply to valid paired cases rather than all 89 failures. The acceptance gate also requires noninferior performance; better grounding cannot compensate for an unacceptable loss of task value.
Auditability is a model-quality variable
The controlled-failure gate achieved balanced accuracy between 0.869 and 0.943 and specificity between 0.779 and 0.938. More consequential for model selection was the difference in usable audit coverage.
Q4 and M8 achieved complete intervention validity, responsiveness, and re-audit pair coverage. G12 had a strict structured-output rate of 0.887 and retained only 31 of 45 performance pairs and 22 of 45 grounding pairs. A checkpoint can therefore remain competitive on aggregate performance while producing too many malformed outputs for consistent governance.
For procurement teams, schema reliability and adapter behavior become selection criteria rather than implementation cleanup. For model-risk teams, malformed outputs must remain invalid instead of being silently repaired or scored. For incident review, the execution manifest—prompts, interventions, settings, tool outputs, and scenario identifiers—becomes part of the defensible record.
How utilities could turn the audit into a deployment gate
The paper directly demonstrates a controlled audit and correction process. The following operating model is a Cognaptus inference from that evidence:
| User | Decision | Condition for action | Boundary |
|---|---|---|---|
| Grid operator | Allow a recommendation into an operational workflow | Output and interventions are valid; responsiveness, performance, and grounding clear preregistered thresholds | Thresholds require local calibration and workflow validation |
| Model-risk team | Correct, tool-route, or escalate a response | Claimed reliance, behavioral reliance, and engineering importance materially diverge | Ablations diagnose observable dependence, not hidden reasoning |
| Procurement team | Select a checkpoint and adapter | Compare validity, responsiveness, signed modality effects, and re-audit coverage alongside accuracy | Evidence covers three quantized checkpoints, not the wider model population |
| Compliance or incident-review team | Defend or reconstruct a recommendation | Retain the execution manifest and intervention results | Reproducibility does not itself establish field safety |
This design is most credible as a predeployment or controlled-production gate for well-specified tasks. It is less applicable when admissible evidence cannot be preregistered, interventions cannot preserve physical invariants, or task outputs cannot be parsed reliably.
The evidence stops before field reliability
The study is a methodological proof of concept using three locally served quantized models, five diagnostic tasks, and two IEEE test systems. Its 450 model-regime observations and 2,606 registered calls provide a controlled demonstration, but the bootstrap intervals capture variation across scenario families, not uncertainty across the population of available models.
Utility-scale workflows introduce further dependencies: telemetry quality, operator procedures, changing network states, tool latency, cybersecurity controls, and accountability across human and automated decisions. None is evaluated here. The paper supports a stronger validation architecture; it does not provide operational clearance.
Trust requires an auditable evidence pathway
The central contribution is not another explanation score. It is a way to convert evidence use into an acceptance condition.
A grid-diagnosis recommendation should not pass because it is correct and accompanied by a plausible rationale. It should pass because the output is valid, the model responds to the evidence that engineers preregistered as relevant, corrective action survives a new ablation test, and the remaining uncertainty is visible in the audit record.
That standard is stricter than accuracy. In safety-critical operations, it is also more defensible.
Cognaptus: Automate the Present, Incubate the Future.