Agency Check, Please: What a New Benchmark Says About LLMs That Actually Empower Users
HumanAgencyBench turns the fuzzy idea of user empowerment into six testable assistant behaviours—and shows why helpfulness is not the same as agency support.
HumanAgencyBench turns the fuzzy idea of user empowerment into six testable assistant behaviours—and shows why helpfulness is not the same as agency support.
AI scientist systems do not just automate research; they relocate scientific judgement into hidden workflow choices, where ordinary paper review can no longer see it.
A mechanism-first look at why treating LLM responses as editable components may matter more than yet another round of prompt engineering.
A mechanism-first guide to why Plan-then-Execute agents improve control-flow security, where they still fail, and how enterprises should harden them before production.
EnvX shows how repositories can become callable agents, but the real business value is disciplined software reuse—not fantasy staff replacement.
How statistical wrappers, calibration sets, confidence intervals, and interventions turn generative AI reliability from theatre into operating discipline.
ImportSnare shows how poisoned code manuals can steer retrieval-augmented code generators into recommending malicious dependencies, turning documentation governance into a supply-chain control.
A mechanism-first look at Astra, a multi-agent LLM system that optimizes existing CUDA kernels and shows where GPU cost savings may actually come from.
HALT-RAG shows how calibrated verification can turn RAG hallucination detection from a vague quality score into an operational reject option.
RLFactory shows how agent RL can be rebuilt around tool feedback, async invocation, and modular rewards—useful plumbing, not magic autonomy.