TL;DR for operators

A high-stakes AI system can fail in two very different ways: it can act on something that is not true, or it can fail to act when danger is real. Those errors need not have comparable consequences.

Martino Maggetti’s Reciprocal Trust and Distrust in Artificial Intelligence Systems: The Hard Problem of Regulation1 argues that this asymmetry should influence who receives final decision authority. In nuclear launch and strategic-warning settings, where a false positive could trigger catastrophic action, the paper favors protected human authority, independent corroboration, and explicit uncertainty. In reactor, chemical-process, and flight-control settings, where failing to intervene can be catastrophic, it allows for bounded AI authority through mechanisms such as non-overridable shutdown logic.

For organizations, the useful unit of analysis is therefore not “human in the loop” versus “AI autonomy” in general. It is the specific decision point: which error is more dangerous, who is better positioned to block it, and what verification and override mechanisms prevent that allocation from becoming a new single point of failure.

The boundary matters. The paper develops a conceptual regulatory heuristic through two historical counterfactuals and qualitative LLM plausibility checks. It does not empirically establish that a particular authority allocation will improve safety in deployed systems.

Human oversight can reduce risk—or preserve the wrong failure mode

A common safety instinct is to keep humans in final control. That works when automation may convert uncertain evidence into irreversible action. It is less convincing when the human is the source of the dangerous decision and the machine is permitted only to advise.

The paper makes this tension concrete with two contrasting counterfactual cases.

In its Serpukhov-15 scenario, the relevant danger is a catastrophic false positive: an automated system interprets faulty or ambiguous warning evidence as grounds for action. Under that configuration, more automation can magnify the consequences of bad inputs. Independent human judgment, corroboration requirements, uncertainty displays, and protected veto authority create deliberate barriers between an alert and an irreversible decision.

The paper then reverses the error structure with Chernobyl. Here, the critical danger is a false negative in the operational sense of failing to prevent an unsafe action. If a system correctly identifies a dangerous condition but humans can simply disregard it, advisory AI may add information without adding protection. Bounded machine authority—such as enforcing a hard safety constraint—can become the safer allocation.

These cases are illustrations, not observations of AI systems actually making those historical decisions. Their role is to expose why one permanent rule about human control cannot resolve both error configurations.

The governing variable is the cost of the wrong error

The paper condenses the two cases into a simple regulatory heuristic: assign greater or final authority according to which error is more consequential.

Decision context Costlier error Authority favored by the paper Illustrative safeguard
Nuclear launch, strategic warning False positive Human Redundancy, uncertainty display, protected veto
Reactor, chemical process, flight control False negative AI system Non-overridable shutdown logic, audited control mechanisms

This reframes authority as a failure-mode decision rather than a status decision about whether humans or machines are inherently more trustworthy.

That distinction matters because “human override” bundles together several different controls. A person may have authority to inspect a recommendation, delay execution, veto an action, alter a threshold, disable an interlock, or change the system configuration. Those rights do not carry the same risk. Likewise, “AI autonomy” can mean anything from ranking options to enforcing a narrow operating invariant.

The paper’s argument therefore points toward allocating specific decision rights rather than assigning blanket control to one side.

Reciprocal trust means machines may also need to challenge humans

Only after establishing the authority problem does the paper introduce its broader conceptual move: trust should be analyzed in both directions.

This does not mean that an AI system experiences trust as a person does. The paper uses the term functionally. A system can rely on human-provided data, instructions, feedback, or interventions; it can also weight them differently, challenge them, or refuse to follow them. Those behaviors can be treated as trust-like or distrust-like without claiming consciousness or emotional judgment.

The familiar governance question is whether humans should trust an AI system. Maggetti adds the reverse question: when should an AI system accept the human as a reliable source of instruction?

That reverse direction is operationally consequential. If a human operator is mistaken, negligent, biased, or deliberately violating a safety constraint, unconditional machine obedience can preserve the hazard. Yet giving the machine authority to reject human instructions creates a second-order control problem: the system itself can also be wrong.

The paper’s target is therefore neither maximum trust nor maximum autonomy. It is alignment between reliance and actual trustworthiness on both sides.

Cognaptus inference: design authority at the decision-point level

For firms deploying AI in consequential workflows, Cognaptus infers a concrete governance exercise from this framework.

Start by decomposing a workflow into decisions that can materially change system state. For each one, identify the consequences of an unnecessary intervention and of a missed intervention. The purpose is not to calculate a universal score, but to expose whether the two error directions are materially asymmetric.

That classification can then inform four control choices:

  • Veto rights: who can stop an action once the system recommends it?
  • Override rights: can the AI block or reverse a human instruction when a defined safety condition is violated?
  • Verification requirements: what independent evidence must exist before either side can exercise final authority?
  • Change authority: who can alter thresholds, disable interlocks, modify control logic, or approve exceptions?

In a high-cost false-positive workflow, an operator might require independent confirmation and a protected review window before automated execution. In a high-cost false-negative workflow, the safer architecture may include a narrowly specified safety invariant that cannot be bypassed by a single operator.

The affected users are risk owners, safety engineers, AI governance teams, and operational leaders deciding who receives control over consequential actions. The boundary is equally specific: the framework is most relevant where error consequences are severe enough that veto and override design deserves explicit governance rather than being inherited from a generic approval workflow.

Watchful trust requires more than choosing the right veto holder once

The paper does not treat the costlier-error rule as a complete operating formula. Humans and AI systems can both become unreliable, and the conditions that justified an authority allocation can change.

Its proposed answer is a form of reciprocal watchful trust: reliance combined with institutionalized skepticism. The supporting mechanisms include redundancy, transparent uncertainty, independent audits, periodic trust audits, incident sharing, stakeholder participation, overlapping oversight, and adaptive regulation.

For an organization, this implies that decision-right allocation should itself be auditable. A trust audit could ask whether users have begun over-relying on an AI system, whether they routinely disregard reliable warnings, whether the system places excessive weight on poor-quality human inputs, or whether an override mechanism is being used outside the conditions for which it was designed.

The governance object is therefore not just model behavior. It is the evolving relationship among model outputs, human instructions, technical interlocks, institutional authority, and the procedures for changing any of them.

The framework is a heuristic, not an empirical control law

The evidence supports the conceptual argument more strongly than any claim about real-world effectiveness.

The two historical cases are structured counterfactuals. They clarify opposing error regimes but do not show that a deployed AI system would necessarily have escalated the Serpukhov-15 incident or prevented Chernobyl.

The paper also reports responses from ChatGPT 5.2, Gemini 3, and Claude Sonnet 4.5 to two counterfactual prompts. Their likely purpose is qualitative plausibility checking: they show that the imagined scenarios are intelligible to contemporary frontier models. They are not a benchmark, controlled experiment, or causal validation of the proposed authority rule.

Implementation remains unresolved as well. The paper explicitly ties feasibility to institutional capacity, sector conditions, political choices, distributional effects, democratic legitimacy, and the ability of oversight mechanisms to keep pace with technological change.

For operators, that leaves a bounded conclusion. Error asymmetry is a disciplined way to begin allocating decision rights. It is not sufficient evidence that a particular organization has allocated them correctly.

Move the governance question from “who is in control?” to “which failure are we controlling?”

The paper’s strongest contribution is to make the human-versus-AI control debate more granular.

When acting on a false alarm is the dominant catastrophic risk, preserving independent human judgment can be protective. When failing to stop a real danger is the dominant risk, restricting AI to advice can leave the critical failure path untouched. Both prescriptions follow from the same underlying question: which error must the system be designed to resist?

That shift changes what organizations should document. Instead of recording only whether a workflow includes human review, governance teams can record why a particular actor holds veto or override authority, which error that authority is intended to contain, what safeguards constrain it, and what evidence would justify changing the allocation later.

The paper does not settle those decisions. It provides a stronger way to frame them.

Cognaptus: Automate the Present, Incubate the Future.


  1. Martino Maggetti (2026). Reciprocal Trust and Distrust in Artificial Intelligence Systems: The Hard Problem of Regulation. arXiv:2604.05826. https://arxiv.org/abs/2604.05826 ↩︎