A Proof Can Pass and Still Mean the Wrong Thing: What AxQM Changes About Formal AI Evaluation
AxQM shows where kernel-checked proof evaluation can remove grader ambiguity—and where semantic review still has to carry the risk.
AxQM shows where kernel-checked proof evaluation can remove grader ambiguity—and where semantic review still has to carry the risk.
Commit-first LLM judging can stop evaluator gaming when the judge is right—and amplify it when the judge is wrong.
KFS-RAG shows how replacing raw retrieved passages with query-relevant facts can reduce prompt-injection leakage while preserving useful RAG performance.
OODA-Tool shows that reliable multi-turn tool use depends on checking state, readiness, action structure, and argument grounding before an agent is allowed to execute.
AstronOS tests whether long-running agent workflows work better when accepted state is versioned and governed instead of reconstructed from conversation history.
PTA-IRT shows how historical agent trajectories can help smaller SWE-agent evaluation subsets preserve full-benchmark scores and rankings.
CRS-Bench shows why medical image encoder screening needs calibration, label efficiency, and robustness alongside clean-test AUROC.
PsychJail shows why release testing should examine how model safety changes when an attacker adapts persuasion strategy across turns, not only whether a harmful prompt is refused once.
GUARD-SLM shows how compact language models can use their own hidden activations to reject jailbreak prompts before generation, trading extra model calls for a model-specific calibration burden.
A systematic review of 88 prompt-injection defenses suggests security teams should design layered safeguards around their application architecture instead of ranking methods by isolated benchmark scores.