Cover image

A Proof Can Pass and Still Mean the Wrong Thing: What AxQM Changes About Formal AI Evaluation

AxQM shows where kernel-checked proof evaluation can remove grader ambiguity—and where semantic review still has to carry the risk.

September 22, 2026 · 6 min · Zelina
Cover image

Commit First, Fail First: Qualify the Judge Before It Steers Optimization

Commit-first LLM judging can stop evaluator gaming when the judge is right—and amplify it when the judge is wrong.

September 22, 2026 · 7 min · Zelina
Cover image

Give the Generator Less to Leak: KFS-RAG Moves Privacy to the Retrieval Boundary

KFS-RAG shows how replacing raw retrieved passages with query-relevant facts can reduce prompt-injection leakage while preserving useful RAG performance.

September 22, 2026 · 8 min · Zelina
Cover image

State Before Action: OODA-Tool Puts a Control Layer Between Context and Execution

OODA-Tool shows that reliable multi-turn tool use depends on checking state, readiness, action structure, and argument grounding before an agent is allowed to execute.

September 22, 2026 · 8 min · Zelina
Cover image

State Is the Workflow: AstronOS Moves Long-Horizon Agents Beyond Transcript Replay

AstronOS tests whether long-running agent workflows work better when accepted state is versioned and governed instead of reconstructed from conversation history.

September 22, 2026 · 7 min · Zelina
Cover image

The Benchmark Is in the Trace: Reusing Agent Trajectories to Shrink SWE Evaluation

PTA-IRT shows how historical agent trajectories can help smaller SWE-agent evaluation subsets preserve full-benchmark scores and rankings.

September 22, 2026 · 8 min · Zelina
Cover image

When AUROC Agrees and the Decision Still Changes

CRS-Bench shows why medical image encoder screening needs calibration, label efficiency, and robustness alongside clean-test AUROC.

September 22, 2026 · 7 min · Zelina
Cover image

A Refusal Is Only One Turn: PsychJail Tests Safety Under Adaptive Persuasion

PsychJail shows why release testing should examine how model safety changes when an attacker adapts persuasion strategy across turns, not only whether a harmful prompt is refused once.

September 21, 2026 · 7 min · Zelina
Cover image

Before the First Token: Put the Jailbreak Gate Inside the Model

GUARD-SLM shows how compact language models can use their own hidden activations to reject jailbreak prompts before generation, trading extra model calls for a model-specific calibration burden.

September 21, 2026 · 7 min · Zelina
Cover image

Don’t Rank the Guardrails: Map Prompt-Injection Defenses to the Stack

A systematic review of 88 prompt-injection defenses suggests security teams should design layered safeguards around their application architecture instead of ranking methods by isolated benchmark scores.

September 21, 2026 · 8 min · Zelina