Cover image

The Answer Looked Clinical. The Critical Steps Were Missing.

A clinician-built evaluation shows why polished medical AI responses can pass presentation checks while omitting the reasoning steps most closely tied to patient harm.

August 1, 2026 · 9 min · Zelina
Cover image

The Smart Part Was the Memory, Not the Controller

Neural Subspace Reallocation shows that continual-learning recovery depends more on storing and retrieving compact task adapters than on training a sophisticated allocation controller.

August 1, 2026 · 9 min · Zelina
Cover image

Fewer Extreme Costs, Higher Average Cost: Risk-Aware Planning Beyond Expected Reward

A practical reading of how risk-aware GUMDPs trade average performance for fewer severe trajectory-level outcomes, and where that framework remains unproven.

July 31, 2026 · 9 min · Zelina
Cover image

One Correction, Every Case: When LLMs Actually Update the Rule

A controlled reversal-learning study shows how to distinguish genuine rule transfer from gradual case-by-case recovery in large language models.

July 31, 2026 · 10 min · Zelina
Cover image

Stocked but Not Synthesizable: URSA Tests the Chemistry Inside the Route

URSA shows why route completion is a weak procurement metric and offers a staged way to screen retrosynthesis systems before chemists commit laboratory effort.

July 31, 2026 · 8 min · Zelina
Cover image

A Full Distribution Is Not a Risk Certificate

A state-level audit shows why distributional reinforcement-learning heads should not drive safety decisions until their strongest risk claims are independently validated.

July 30, 2026 · 8 min · Zelina
Cover image

Five Answers, One Bad Retrieval: When RAG Agreement Misleads

A diagnostic framework for detecting RAG failures that remain invisible when repeated answers agree.

July 30, 2026 · 10 min · Zelina
Cover image

Same Agent, Different Audience: When Social Pressure Changes the Recommendation

A dual-channel benchmark shows how authority, sponsorship, and future dependence can alter an agent’s public recommendation without changing its model or stated task.

July 30, 2026 · 7 min · Zelina
Cover image

Proof, Then Trust: Lean Certifies an AI-Generated QAOA Result

A machine-checked QAOA proof shows how generative systems can search broadly while deterministic verification controls what gets accepted.

July 29, 2026 · 8 min · Zelina
Cover image

Rank the Work, Not the Model: Meta-Benchmarks for Bank LLM Screening

A financial-services meta-benchmark turns public LLM results into domain-specific screening evidence—without pretending that a ranking is deployment approval.

July 29, 2026 · 10 min · Zelina