TL;DR for operators

When an autonomous robot brakes, reroutes, or stops near a person, the operational question after an incident is larger than “which input features influenced the model?” Investigators may need the supporting sensor evidence, the state the system inferred, the alternatives it considered, why one action won, and whether execution matched the intended command.

TRACE is an architecture designed to preserve that chain while the decision is being made. The paper reports high completeness of its audit records across 500 simulated warehouse decision cycles: 98.6% Evidence Traceability, 99.0% Temporal Continuity, and 98.1% Decision Reconstructability, with about 0.12 ms of auditability computation per cycle.

Those numbers measure whether decision artifacts were recorded, not whether the robot made a safe or correct decision. For manufacturers and operators, the larger implication is architectural: auditability requires interfaces that expose intermediate representations and infrastructure that can retain and retrieve them. Hardware performance, investigator usefulness, cross-domain reliability, and fleet-scale storage costs remain unresolved.

An incident review needs more than a prediction explanation

Consider a warehouse robot that stops because a pedestrian enters its path. A post-hoc explanation might show which visual features or sensor inputs influenced a prediction. That can help interpret one component of the system, but it does not automatically answer the incident-review questions that follow.

What object did the robot think it detected? Which observations supported that belief? What alternative actions were considered? Why did braking outrank steering around the person? What would have needed to change for another action to be selected? Did the robot actually execute the intended maneuver?

Cagri Temel’s Towards Trustworthy Autonomous Robots: An Explainable AI-Based Decision Framework1 treats those questions as an architecture problem. Its TRACE framework—Transparent Reasoning Architecture for Credible Execution—records structured artifacts across the path from perception to execution rather than attempting to reconstruct that path after the event.

That distinction matters operationally. An explanation of a model prediction and an audit trail of an autonomous decision solve related but different problems.

TRACE makes auditability part of the decision pipeline

TRACE divides the decision process into four layers.

The Semantic Perception layer links detected entities to contributing sensor observations. The Belief Reasoning layer records probabilistic relationships among evidence, inferred scene states, and downstream decisions. Action Synthesis retains candidate actions, the selection rationale, and counterfactual conditions. Execution Verification then records what happened through timestamped audit entries connecting entities, beliefs, actions, deviations, and flags.

The counterfactual component is especially relevant to incident analysis. Recording that the robot chose “stop” is weaker than recording that it evaluated “stop,” “slow,” and “detour,” together with the conditions under which one of those alternatives would have become preferable.

This design also creates an integration requirement. Existing perception and planning modules cannot expose only final outputs if the surrounding system is expected to reconstruct their contribution later. Detected entities, belief updates, candidate actions, rationales, and execution outcomes have to remain accessible across module boundaries.

Auditability therefore becomes partly an API contract between robotic subsystems.

The metrics measure whether the trail is complete

The paper turns this architectural goal into explicit measurements rather than leaving “explainability” as a qualitative property.

Metric Operational question What a high score does not establish
Evidence Traceability (ET) Can the factors influencing a decision be followed back to recorded sensor evidence? That the evidence or resulting inference was correct
Temporal Continuity (TC) Does the audit record remain continuous across the evaluation period? That recorded decisions were safe
Decision Reconstructability (DR) Are the chosen action, alternatives, and selection rationale present? That the rationale was causally valid or useful to a human investigator

The distinction is important because these are documentation-completeness measures. A perfectly reconstructed bad decision is still a bad decision.

The paper also defines partial Evidence Traceability for cases where a complete evidence-to-decision chain is missing. Instead of treating such cases as entirely opaque, it measures how much of the required causal depth remains documented. That can make incomplete records diagnostically useful without pretending that partial reconstruction is equivalent to full traceability.

The warehouse experiment shows simulated feasibility, not safer robots

TRACE is evaluated in a Python-based warehouse simulation using LiDAR, camera, and ultrasonic observations. Five scenarios cover a clear path, pedestrian encounters at different distances, multiple pedestrians, and an approaching forklift. Each scenario contains 100 decision cycles, giving 500 cycles at a 10 Hz decision rate.

The evaluation deliberately injects failures including sensor unavailability, evidence-link corruption, missing causal edges, rationale truncation, counterfactual-generation timeouts, and audit-record write failures. This makes the experiment a test of whether the audit trail survives modeled faults rather than merely whether the architecture works under ideal logging conditions.

The headline reported results are:

  • 98.6% Evidence Traceability
  • 99.0% Temporal Continuity
  • 98.1% Decision Reconstructability
  • approximately 0.12 ms auditability overhead per decision cycle

Per-scenario Evidence Traceability ranges from 96.4% for the clear-path scenario to 100% for the forklift scenario.

These figures support a bounded conclusion: under the simulated warehouse conditions and injected failure model, TRACE can produce highly complete decision-audit artifacts with low reported computation time.

They do not show that TRACE improves navigation accuracy, prevents collisions, makes explanations causally correct, or helps investigators reach better conclusions.

The comparison with attention, SHAP, and LIME also needs careful reading. Those methods were not implemented as directly comparable decision systems in the experiment; their reported values are architectural capability estimates. The comparison therefore supports the design argument that post-hoc prediction explanations lack some of TRACE’s decision-level artifacts, not a conventional benchmark claim that TRACE experimentally outperforms those methods.

There is an additional reporting issue. Some numerical statements do not reconcile internally: the source notes that 23 unreconstructable records out of 500 would imply 95.4% rather than 98.1% Decision Reconstructability, while scenario-level Temporal Continuity values do not straightforwardly reproduce the reported 99.0% aggregate. The percentages should therefore be treated as paper-reported results pending reconciliation with corrected materials or source code.

Auditability becomes an infrastructure decision at fleet scale

For production systems, Cognaptus draws a broader implication from the architecture: richer auditability changes more than the robot’s reasoning code.

A manufacturer would need to decide which intermediate representations become durable system interfaces. An operator would need policies for retention, indexing, access control, retrieval, and incident preservation. Certification and risk teams would need definitions of what constitutes a sufficiently reconstructable event record.

Storage can become material. The paper projects roughly 1.7–4.2 GB per robot per day before compression and more than 30 PB for a fleet of 100,000 robots under a six-month retention assumption. It proposes hot, warm, and cold storage tiers plus lossless and delta compression that could reduce the projected requirement to roughly 2–5 PB.

Those are planning estimates, not measurements from an operational fleet. Still, they expose the design trade-off: retaining more context can make decisions easier to reconstruct while increasing the cost of keeping that context available.

Auditability is therefore not a logging switch added near deployment. It affects interfaces, storage architecture, governance rules, and potentially the economics of fleet operation.

What still has to be demonstrated

The next validation step is not another percentage from the same simulator.

TRACE has not been tested here on physical robotic hardware, where embedded compute constraints, real sensor faults, timing variation, and communication failures may change both completeness and overhead. The evaluation also does not test whether human investigators can use the resulting records efficiently or whether reconstructed rationales correspond to the true causal basis of a decision.

The warehouse setting leaves transfer to autonomous vehicles, surgical robots, agricultural systems, and other safety-critical environments unresolved. And incomplete evidence chains may matter most precisely when the operating situation is unusual or safety-critical.

For an engineering organization, that sets a clear validation sequence: first verify that the architecture can preserve the required artifacts on target hardware; then test whether those artifacts improve real incident investigation and governance workflows; only after that can traceability results support stronger claims about deployed operational value.

Auditability starts before the incident

The paper’s more durable contribution is not the exact value of its reported 98.6% traceability score. It is the decision to treat reconstructability as something the system must prepare for before an incident occurs.

Once that requirement is accepted, several design choices follow. Intermediate states cannot disappear behind module boundaries. Rejected actions may need to be retained. Rationales and counterfactual conditions become structured records. Logging continuity needs measurable service levels. Storage and retention become part of autonomous-system governance.

TRACE shows that such an architecture can be built and exercised under simulated warehouse conditions. Whether it remains complete, usable, and affordable on deployed robots is the harder question still ahead.

Cognaptus: Automate the Present, Incubate the Future.


  1. Cagri Temel (2026). Towards Trustworthy Autonomous Robots: An Explainable AI-Based Decision Framework. arXiv:2609.02861. https://arxiv.org/abs/2609.02861 ↩︎