TL;DR for operators
A policy-aware LLM has to do more than retrieve the right document. It must know what evidence is mandatory, when evidence is insufficient, when a case requires human review, which policy version governed the decision, and what information must be retained so the decision can later be reconstructed.
Policy-as-Skill: Governed LLM Decision Support with Evidence, Deterministic Control, and Audit1 turns those obligations into explicit runtime components rather than leaving them inside prompts or application glue code.
The strongest operational result is also the most cautionary. Adding deterministic policy control raised aggregate exact accuracy from 53.8% for the audit-only variant to 61.2% for the full system, but the average conceals sharply different task effects. The controller improved conflict detection by 59.3 percentage points while reducing risk-classification accuracy by 23.3 points and policy-QA accuracy by 8.7 points.
For organizations building governed decision support, that shifts the design problem. The useful question is which obligations deserve deterministic authority, which require model-assisted interpretation, and which conditions should end in human review.
Governance starts where retrieval stops
Suppose an internal assistant is asked whether a proposed action complies with company policy. Supplying the relevant documents is necessary, but it does not determine what happens when a required document is missing, two rules conflict, the request concerns a high-impact context, or the policy changed last week.
Those are runtime governance decisions.
The paper formalizes them through a policy skill
where the components specify skill identity and version, retrieval scope, required evidence, decision schema, human-review triggers, audit fields, failure behavior, prompt structure, and contextual boundaries.
This is the main architectural distinction. Policy-as-Skill is not primarily a more elaborate way to phrase instructions. A policy becomes a versioned capability whose operational requirements can be materialized separately.
That separation matters for change management. A system can record which policy version applied, restrict retrieval to authorized sources, require particular evidence, constrain available decision labels, escalate specified cases, and preserve the evidence IDs, hashes, model version, timestamps, and other trace information needed for later reconstruction.
The model still produces a probabilistic interpretation. The surrounding skill defines what the organization permits the system to do with it.
The ablations separate governance from override
The benchmark evaluates 600 instances across policy question answering, compliance checking, risk classification, and policy conflict detection. Thirteen methods use common substantive metrics, while model-based methods share the same Gemma4 backend.
The PaS variants separate three mechanisms that are often bundled together: skill-scoped retrieval, deterministic control, and audit or validation.
That decomposition produces a useful result. PaS Retrieval reaches 53.3% exact accuracy and 0.855 review F1. PaS+Audit reaches 53.8% exact accuracy and 0.854 review F1. LLM+RAG reaches 49.8% exact accuracy and 0.695 review F1.
PaS+Audit also records complete native audit fields in the reported evaluation. Its citation precision is 1.000 and policy-reference recall is 0.984.
These metrics describe different properties. Audit completeness means the system preserved the fields required to reconstruct its operation. It does not mean the substantive decision was correct. Likewise, LLM+RAG already achieves citation precision of 0.994 and policy-reference recall of 0.965, yet its review-routing F1 remains materially lower than the PaS retrieval and audit variants.
The paper therefore gives little support to treating “well cited,” “auditable,” “correct,” and “appropriately escalated” as interchangeable labels.
Deterministic control works differently across tasks
The controller is the part of the architecture that can replace or redirect the model’s candidate judgment according to explicit rules.
Across all 600 instances, it changes 401 decisions relative to PaS+Audit. Of those changes, 181 correct previously incorrect decisions, 137 turn previously correct decisions into incorrect ones, and 83 move from one incorrect class to another. The resulting net improvement is 44 tasks, or 7.3 percentage points.
The task-family decomposition explains why the aggregate rises:
| Task family | PaS+Audit | PaS Full | Controller effect |
|---|---|---|---|
| Policy QA | 40.7% | 32.0% | -8.7 pp |
| Compliance | 63.3% | 65.3% | +2.0 pp |
| Risk classification | 70.7% | 47.3% | -23.3 pp |
| Conflict detection | 40.7% | 100.0% | +59.3 pp |
The conflict-detection subset is unusual: all 150 cases have needs review as the reference decision. That makes it well suited to testing escalation behavior, but weak as a discriminative multi-class benchmark. A controller encoding conflict escalation has a structural advantage there.
Remove that subset and the interpretation changes. PaS+Audit reaches 58.2% exact accuracy, compared with 50.4% for LLM+RAG. PaS Full falls to 48.2%.
The evidence supports deterministic enforcement where the policy condition itself is precise. It does not support turning every interpretive policy judgment into a rule override.
The business design is a division of decision rights
Cognaptus’ inference is that the architectural value lies in assigning different forms of authority to different components.
An explicit prohibition, mandatory-document requirement, or mandatory escalation condition can plausibly become deterministic runtime logic. The rule is narrow, testable, and its intervention can be logged.
A risk classification that depends on contextual interpretation is different. Here, replacing the model’s judgment with a coarse rule can discard information the model was using productively. The reported 23.3-point decline on risk classification is a concrete example of that failure mode within this benchmark.
Human review then becomes more than a fallback after automation fails. In the PaS design, missing evidence, unresolved conflicts, ambiguity, or high-impact contexts can themselves be governed outcomes that intentionally route a case to a person.
For compliance teams, model-risk functions, internal audit, or operators maintaining policy-driven workflows, this suggests a practical separation:
- encode unambiguous obligations as deterministic controls;
- use scoped retrieval and structured evidence requirements for interpretive work;
- make escalation conditions explicit rather than relying on model confidence alone;
- preserve policy, model, evidence, and trace versions so decisions can be replayed and inspected;
- evaluate correctness, evidence quality, escalation behavior, and auditability separately.
The paper’s composite governance-readiness analysis reinforces the last point. Across 286 alternative weight combinations, the top-ranked method changes materially. A weighted readiness score therefore reflects engineering priorities as well as measured system properties; it is not a neutral substitute for the underlying metrics.
The benchmark supports architecture, not production certification
The evidence is useful but bounded.
The 600-instance benchmark is a development benchmark rather than a frozen final-system test. The underlying model did not train on these examples, but earlier benchmark trace analysis influenced controller and system development. Model-level unfamiliarity is therefore not the same as system-level held-out evaluation.
The policy corpus also consists of synthetic, illustrative, or real-world-inspired fixtures rather than authoritative legal or enterprise policies. Reference decisions were curated by one expert, only one model backend was tested, and the automatic evidence-faithfulness diagnostic lacks human validation in the reported run.
Reported latency differences should likewise remain descriptive because hardware conditions, caching, and run order were not controlled for performance attribution.
These constraints limit claims about legal correctness, production reliability, and cross-model generalization. They do not erase the architectural result: governance mechanisms can be separated, tested independently, and assigned different decision rights.
Govern the handoff, not only the answer
The paper’s most durable idea is not that stricter rules produce better LLM decisions. Its own benchmark shows that they sometimes do the opposite.
The stronger contribution is a way to make policy operationally explicit. Retrieval scope, evidence requirements, review triggers, failure behavior, audit obligations, policy versions, and deterministic authority become inspectable parts of the runtime system rather than implicit assumptions surrounding a prompt.
That gives organizations a more precise governance problem to solve. They can decide where interpretation belongs, where policy authority belongs, and where neither the model nor a deterministic rule should make the final call.
Cognaptus: Automate the Present, Incubate the Future.
-
Kabeh Mohsenzadegan and Vahid Tavakkoli and Kyandoghere Kyamakya (2026). Policy-as-Skill: Governed LLM Decision Support with Evidence, Deterministic Control, and Audit. arXiv:2609.27087. https://arxiv.org/abs/2609.27087 ↩︎