Blame Isn’t a Bug: Turning Agent ‘Whodunits’ into Fixable Systems
A practical reading of incident analysis for AI agents: why serious failures need causal evidence, not just public anecdotes and model blame.
A practical reading of incident analysis for AI agents: why serious failures need causal evidence, not just public anecdotes and model blame.
A practical reading of the APCP framework as a maturity ladder for designing AI that supports learning without pretending software has a soul.
A practical reading of why AI introspection should mean privileged self-access, not merely clever self-reporting from visible output.
MCP-Universe shows that connecting agents to real tools is easy; making them reliable across messy, live workflows is still the hard part.
Structured planner-derived examples help LLM agents with simple shared-visibility filtering, but the harder business problem is belief tracking and pricing the cost of information.
ComputerRL shows that useful desktop agents may depend less on prettier clicking and more on machine-friendly APIs, scalable online RL, and training schedules that keep exploration alive.
A mechanism-first reading of an agentic AI science system that ran an online human-participant experiment, wrote three manuscripts, and showed where research automation is useful—and where it still needs adult supervision.
A comparison-driven operator guide to why active memory management may matter more than longer context windows for enterprise AI agents.
A mechanism-first reading of why benign agent fine-tuning can erode refusals, and why PING works by steering the first response tokens rather than rewriting the model.
TS-Agent shows that financial modelling agents improve when they are constrained by curated model banks, refinement knowledge, feedback loops, and auditable code-edit trails.