Cover image

Agency Check, Please: What a New Benchmark Says About LLMs That Actually Empower Users

HumanAgencyBench turns the fuzzy idea of user empowerment into six testable assistant behaviours—and shows why helpfulness is not the same as agency support.

September 14, 2025 · 16 min · Zelina
Cover image

Automate All the Things? Mind the Blind Spots

AI scientist systems do not just automate research; they relocate scientific judgement into hidden workflow choices, where ordinary paper review can no longer see it.

September 14, 2025 · 18 min · Zelina
Cover image

From Blobs to Blocks: Componentizing LLM Output for Real Work

A mechanism-first look at why treating LLM responses as editable components may matter more than yet another round of prompt engineering.

September 14, 2025 · 16 min · Zelina
Cover image

Guardrails Before Gas: Secure Plan‑Then‑Execute Agents for Real Work

A mechanism-first guide to why Plan-then-Execute agents improve control-flow security, where they still fail, and how enterprises should harden them before production.

September 14, 2025 · 15 min · Zelina
Cover image

Repo, Meet Your Agent: Turning GitHub into a Workforce with EnvX

EnvX shows how repositories can become callable agents, but the real business value is disciplined software reuse—not fantasy staff replacement.

September 14, 2025 · 15 min · Zelina
Cover image

Confidence, Not Confidence Tricks: Statistical Guardrails for Generative AI

How statistical wrappers, calibration sets, confidence intervals, and interventions turn generative AI reliability from theatre into operating discipline.

September 13, 2025 · 14 min · Zelina
Cover image

Hook, Line, and Import: How RAG Lets Attackers Snare Your Code

ImportSnare shows how poisoned code manuals can steer retrieval-augmented code generators into recommending malicious dependencies, turning documentation governance into a supply-chain control.

September 13, 2025 · 17 min · Zelina
Cover image

Kernel Kombat: How Multi‑Agent LLMs Squeeze 1.32× More From Your GPUs

A mechanism-first look at Astra, a multi-agent LLM system that optimizes existing CUDA kernels and shows where GPU cost savings may actually come from.

September 13, 2025 · 14 min · Zelina
Cover image

Stop, Verify, and Listen: HALT‑RAG Brings a ‘Reject Option’ to RAG

HALT-RAG shows how calibrated verification can turn RAG hallucination detection from a vague quality score into an operational reject option.

September 13, 2025 · 11 min · Zelina
Cover image

Tool Time, Any Time: Inside RLFactory’s Plug‑and‑Play RL for Multi‑Turn Tool Use

RLFactory shows how agent RL can be rebuilt around tool feedback, async invocation, and modular rewards—useful plumbing, not magic autonomy.

September 13, 2025 · 16 min · Zelina