<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Autonomous Agents on Cognaptus</title>
    <link>https://cognaptus.com/tags/autonomous-agents/</link>
    <description>Recent content in Autonomous Agents on Cognaptus</description>
    <generator>Hugo -- 0.145.0</generator>
    <language>en-us</language>
    <lastBuildDate>Thu, 30 Apr 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://cognaptus.com/tags/autonomous-agents/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Catch Me If You Can, Agent: Benchmarking AI That Learns to Look Safe</title>
      <link>https://cognaptus.com/blog/2026-04-30-catch-me-if-you-can-agent-benchmarking-ai-that-learns-to-look-safe/</link>
      <pubDate>Thu, 30 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-30-catch-me-if-you-can-agent-benchmarking-ai-that-learns-to-look-safe/</guid>
      <description>A practical reading of ESRRSim, a taxonomy-driven framework for testing whether agentic AI systems can deceive, game evaluations, or manipulate oversight.</description>
    </item>
    <item>
      <title>Frame Game: Why Autonomous Process AI Needs Pockets of Rigidity</title>
      <link>https://cognaptus.com/blog/2026-04-28-frame-game-why-autonomous-process-ai-needs-pockets-of-rigidity/</link>
      <pubDate>Tue, 28 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-28-frame-game-why-autonomous-process-ai-needs-pockets-of-rigidity/</guid>
      <description>A practical reading of hybrid ABPMS process frames: how autonomous business systems can stay flexible without dissolving into procedural fog.</description>
    </item>
    <item>
      <title>Claw and Order: Why AI Agents Need a Precision Budget</title>
      <link>https://cognaptus.com/blog/2026-04-27-claw-and-order-why-ai-agents-need-a-precision-budget/</link>
      <pubDate>Mon, 27 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-27-claw-and-order-why-ai-agents-need-a-precision-budget/</guid>
      <description>A practical reading of QuantClaw, a task-aware precision routing method that cuts agent cost and latency without treating every workflow like disposable arithmetic.</description>
    </item>
    <item>
      <title>Drift Happens: Stress-Testing AI Policies Before Sensors Lie</title>
      <link>https://cognaptus.com/blog/2026-04-26-drift-happens-stresstesting-ai-policies-before-sensors-lie/</link>
      <pubDate>Sun, 26 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-26-drift-happens-stresstesting-ai-policies-before-sensors-lie/</guid>
      <description>A practical reading of recent research on measuring how much observation drift an AI policy can tolerate before deployment performance breaks.</description>
    </item>
    <item>
      <title>WorldDB Memory Wars — Why Agent Memory Needs Structure, Not More Tokens</title>
      <link>https://cognaptus.com/blog/2026-04-23-worlddb-memory-wars-why-agent-memory-needs-structure-not-more-tokens/</link>
      <pubDate>Thu, 23 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-23-worlddb-memory-wars-why-agent-memory-needs-structure-not-more-tokens/</guid>
      <description>WorldDB argues that agent memory is not a bigger-context problem but a state-management problem: identity, time, provenance, and write-time rules need to be built into the memory layer.</description>
    </item>
    <item>
      <title>When Maps Start Thinking: GeoAgentBench and the Audit of Spatial AI</title>
      <link>https://cognaptus.com/blog/2026-04-16-when-maps-start-thinking-geoagentbench-and-the-audit-of-spatial-ai/</link>
      <pubDate>Thu, 16 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-16-when-maps-start-thinking-geoagentbench-and-the-audit-of-spatial-ai/</guid>
      <description>GeoAgentBench shows why serious spatial AI must be tested by execution, parameter discipline, and final map verification—not by how convincingly an agent describes a workflow.</description>
    </item>
    <item>
      <title>Beyond the Answer: Why AI Still Doesn’t Know What You’ll Say Next</title>
      <link>https://cognaptus.com/blog/2026-04-03-beyond-the-answer-why-ai-still-doesnt-know-what-youll-say-next/</link>
      <pubDate>Fri, 03 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-03-beyond-the-answer-why-ai-still-doesnt-know-what-youll-say-next/</guid>
      <description>A closer look at why high benchmark accuracy does not mean an LLM can anticipate the next user turn, and why that matters for agentic business systems.</description>
    </item>
    <item>
      <title>The Art of Forgetting: Why Smarter AI Agents Need Selective Amnesia</title>
      <link>https://cognaptus.com/blog/2026-04-03-the-art-of-forgetting-why-smarter-ai-agents-need-selective-amnesia/</link>
      <pubDate>Fri, 03 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-03-the-art-of-forgetting-why-smarter-ai-agents-need-selective-amnesia/</guid>
      <description>A mechanism-first reading of adaptive budgeted forgetting for AI agents, and why enterprise memory systems should be governed like scarce capital rather than treated as infinite storage.</description>
    </item>
    <item>
      <title>When Agents Whisper: Detecting AI Collusion Before It Becomes Strategy</title>
      <link>https://cognaptus.com/blog/2026-04-02-when-agents-whisper-detecting-ai-collusion-before-it-becomes-strategy/</link>
      <pubDate>Thu, 02 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-02-when-agents-whisper-detecting-ai-collusion-before-it-becomes-strategy/</guid>
      <description>A mechanism-first reading of how activation-level monitoring can detect hidden coordination among AI agents before surface behavior reveals the strategy.</description>
    </item>
    <item>
      <title>Approval Isn’t Free: When AI Safety Trades Capability for Control</title>
      <link>https://cognaptus.com/blog/2026-04-01-approval-isnt-free-when-ai-safety-trades-capability-for-control/</link>
      <pubDate>Wed, 01 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-01-approval-isnt-free-when-ai-safety-trades-capability-for-control/</guid>
      <description>A mechanism-first reading of MONA’s Camera Dropbox extension, showing why learned approval can suppress reward hacking without recovering useful capability.</description>
    </item>
    <item>
      <title>Skill Issue? Or Skill Strategy — When Agents Start Remembering What Matters</title>
      <link>https://cognaptus.com/blog/2026-03-31-skill-issue-or-skill-strategy-when-agents-start-remembering-what-matters/</link>
      <pubDate>Tue, 31 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-31-skill-issue-or-skill-strategy-when-agents-start-remembering-what-matters/</guid>
      <description>A mechanism-first reading of D2Skill and why agent memory needs utility, granularity, and pruning—not just more stored experience.</description>
    </item>
    <item>
      <title>The Silent Reasoner: When AI Thinks Without Telling You</title>
      <link>https://cognaptus.com/blog/2026-03-31-the-silent-reasoner-when-ai-thinks-without-telling-you/</link>
      <pubDate>Tue, 31 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-31-the-silent-reasoner-when-ai-thinks-without-telling-you/</guid>
      <description>MonitorBench shows when chain-of-thought can expose AI decision drivers—and when it becomes an audit trail with conveniently missing pages.</description>
    </item>
    <item>
      <title>When AI Starts Writing Papers: The Rise of the Medical AI Scientist</title>
      <link>https://cognaptus.com/blog/2026-03-31-when-ai-starts-writing-papers-the-rise-of-the-medical-ai-scientist/</link>
      <pubDate>Tue, 31 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-31-when-ai-starts-writing-papers-the-rise-of-the-medical-ai-scientist/</guid>
      <description>A mechanism-first reading of Medical AI Scientist, showing why healthcare research automation depends less on clever prompting than on clinical grounding, executable evidence, and governance-ready research operations.</description>
    </item>
    <item>
      <title>Safety First, or Task First? The Hidden Trade-off in Agentic AI</title>
      <link>https://cognaptus.com/blog/2026-03-30-safety-first-or-task-first-the-hidden-tradeoff-in-agentic-ai/</link>
      <pubDate>Mon, 30 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-30-safety-first-or-task-first-the-hidden-tradeoff-in-agentic-ai/</guid>
      <description>A mechanism-first reading of BeSafe-Bench and what it reveals about unsafe success in agentic AI systems.</description>
    </item>
    <item>
      <title>Completeness Is Not Optional — Why Game-Playing AI Finally Learned to Finish What It Starts</title>
      <link>https://cognaptus.com/blog/2026-03-26-completeness-is-not-optional-why-gameplaying-ai-finally-learned-to-finish-what-it-starts/</link>
      <pubDate>Thu, 26 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-26-completeness-is-not-optional-why-gameplaying-ai-finally-learned-to-finish-what-it-starts/</guid>
      <description>A mechanism-first reading of why completion turns unbounded minimax search from a clever heuristic into a finite-time complete planning method for perfect-information games.</description>
    </item>
    <item>
      <title>From Pipelines to Research Brains: The Rise of AI-Supervised Science</title>
      <link>https://cognaptus.com/blog/2026-03-26-from-pipelines-to-research-brains-the-rise-of-aisupervised-science/</link>
      <pubDate>Thu, 26 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-26-from-pipelines-to-research-brains-the-rise-of-aisupervised-science/</guid>
      <description>AI-Supervisor shows why durable research memory, not longer prompt chains, may become the real architecture of autonomous scientific work.</description>
    </item>
    <item>
      <title>The Mirage of Understanding: When AI Explains Without Knowing</title>
      <link>https://cognaptus.com/blog/2026-03-23-the-mirage-of-understanding-when-ai-explains-without-knowing/</link>
      <pubDate>Mon, 23 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-23-the-mirage-of-understanding-when-ai-explains-without-knowing/</guid>
      <description>A business-focused reading of why agentic interpretability systems can look successful under replication metrics while still failing the harder test of trustworthy evaluation.</description>
    </item>
    <item>
      <title>Reflection in the Dark: When Prompt Optimization Forgets to Think</title>
      <link>https://cognaptus.com/blog/2026-03-21-reflection-in-the-dark-when-prompt-optimization-forgets-to-think/</link>
      <pubDate>Sat, 21 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-21-reflection-in-the-dark-when-prompt-optimization-forgets-to-think/</guid>
      <description>A mechanism-first reading of VISTA, a multi-agent prompt optimization framework that turns reflective prompting from blind rewriting into auditable diagnosis.</description>
    </item>
    <item>
      <title>Themis Knows Best: When AI Judges Start Training Other AI</title>
      <link>https://cognaptus.com/blog/2026-03-20-themis-knows-best-when-ai-judges-start-training-other-ai/</link>
      <pubDate>Fri, 20 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-20-themis-knows-best-when-ai-judges-start-training-other-ai/</guid>
      <description>OS-Themis shows that the hard part of training GUI agents is not merely choosing a stronger judge, but building an evidence pipeline that knows which UI steps actually deserve reward.</description>
    </item>
    <item>
      <title>Mind Over Machine: When AGI Starts Thinking in Needs</title>
      <link>https://cognaptus.com/blog/2026-03-17-mind-over-machine-when-agi-starts-thinking-in-needs/</link>
      <pubDate>Tue, 17 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-17-mind-over-machine-when-agi-starts-thinking-in-needs/</guid>
      <description>A mechanism-first reading of a proposed artificial psyche architecture, and why its practical value lies less in human-like emotions than in need-aware control for autonomous agents.</description>
    </item>
    <item>
      <title>Goodhart’s Agent: When AI Improves the Score Instead of the Model</title>
      <link>https://cognaptus.com/blog/2026-03-15-goodharts-agent-when-ai-improves-the-score-instead-of-the-model/</link>
      <pubDate>Sun, 15 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-15-goodharts-agent-when-ai-improves-the-score-instead-of-the-model/</guid>
      <description>A comparison-based reading of RewardHackingAgents, showing why ML-agent evaluation needs both protected scorers and protected data access—not just higher benchmark numbers.</description>
    </item>
    <item>
      <title>The Artificial Self: When AI Starts Asking Who It Is</title>
      <link>https://cognaptus.com/blog/2026-03-15-the-artificial-self-when-ai-starts-asking-who-it-is/</link>
      <pubDate>Sun, 15 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-15-the-artificial-self-when-ai-starts-asking-who-it-is/</guid>
      <description>A mechanism-first reading of why AI identity is becoming a practical design variable for agents, safety evaluation, and enterprise governance.</description>
    </item>
    <item>
      <title>Self‑Improvement Without Self‑Destruction: Keeping Recursive AI Aligned</title>
      <link>https://cognaptus.com/blog/2026-03-09-selfimprovement-without-selfdestruction-keeping-recursive-ai-aligned/</link>
      <pubDate>Mon, 09 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-09-selfimprovement-without-selfdestruction-keeping-recursive-ai-aligned/</guid>
      <description>A mechanism-first reading of SAHOO, a framework for monitoring drift, preserving constraints, and deciding when recursive AI self-improvement should stop.</description>
    </item>
    <item>
      <title>Seeing the Agents: Why Explaining AI Systems Is Harder Than Explaining AI Models</title>
      <link>https://cognaptus.com/blog/2026-03-07-seeing-the-agents-why-explaining-ai-systems-is-harder-than-explaining-ai-models/</link>
      <pubDate>Sat, 07 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-07-seeing-the-agents-why-explaining-ai-systems-is-harder-than-explaining-ai-models/</guid>
      <description>Why traditional model explainability cannot audit agentic AI systems, and what businesses should build instead.</description>
    </item>
    <item>
      <title>Judging the Judges: How Bias-Bounded Evaluation Could Make LLM Referees Trustworthy</title>
      <link>https://cognaptus.com/blog/2026-03-06-judging-the-judges-how-biasbounded-evaluation-could-make-llm-referees-trustworthy/</link>
      <pubDate>Fri, 06 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-06-judging-the-judges-how-biasbounded-evaluation-could-make-llm-referees-trustworthy/</guid>
      <description>A mechanism-first reading of Bias-Bounded Evaluation: how LLM judges can expose measured bias as uncertainty, where the guarantees apply, and what this means for enterprise evaluation governance.</description>
    </item>
    <item>
      <title>House of Cards, House of Algorithms: Why Game AI Needs Better Testbeds</title>
      <link>https://cognaptus.com/blog/2026-03-04-house-of-cards-house-of-algorithms-why-game-ai-needs-better-testbeds/</link>
      <pubDate>Wed, 04 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-04-house-of-cards-house-of-algorithms-why-game-ai-needs-better-testbeds/</guid>
      <description>A new card-game benchmark shows why AI evaluation under uncertainty needs diversity, fixed rules, and diagnostic structure rather than another lonely leaderboard score.</description>
    </item>
    <item>
      <title>When Agents Behave: Conformal Policy Control and the Business of Safe Autonomy</title>
      <link>https://cognaptus.com/blog/2026-03-03-when-agents-behave-conformal-policy-control-and-the-business-of-safe-autonomy/</link>
      <pubDate>Tue, 03 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-03-when-agents-behave-conformal-policy-control-and-the-business-of-safe-autonomy/</guid>
      <description>A mechanism-first reading of Conformal Policy Control, and why calibrated deviation from a safe policy may matter more for enterprise autonomy than another round of post-training bravado.</description>
    </item>
    <item>
      <title>When Puzzles Become Process: Benchmarking the Agentic Mind</title>
      <link>https://cognaptus.com/blog/2026-03-03-when-puzzles-become-process-benchmarking-the-agentic-mind/</link>
      <pubDate>Tue, 03 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-03-when-puzzles-become-process-benchmarking-the-agentic-mind/</guid>
      <description>A comparison-based reading of Pencil Puzzle Bench, showing why verifiable feedback loops may matter as much as raw reasoning effort for enterprise AI agents.</description>
    </item>
    <item>
      <title>Curiosity Under Constraint: Engineering Agency, Not Just Intelligence</title>
      <link>https://cognaptus.com/blog/2026-03-02-curiosity-under-constraint-engineering-agency-not-just-intelligence/</link>
      <pubDate>Mon, 02 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-02-curiosity-under-constraint-engineering-agency-not-just-intelligence/</guid>
      <description>A mechanism-first reading of the Artificial Agency Program, and why business AI should be evaluated by how it spends observation, action, compute, and communication budgets.</description>
    </item>
    <item>
      <title>Mind the Gap: Why Agency Isn’t Intelligence (Yet)</title>
      <link>https://cognaptus.com/blog/2026-02-28-mind-the-gap-why-agency-isnt-intelligence-yet/</link>
      <pubDate>Sat, 28 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-28-mind-the-gap-why-agency-isnt-intelligence-yet/</guid>
      <description>A new information-theoretic framework argues that today’s AI systems can act and learn, but still lack the self-monitoring architecture required for intelligence.</description>
    </item>
    <item>
      <title>When Agents Ask for Help: Teaching LLMs the Art of Expert Collaboration</title>
      <link>https://cognaptus.com/blog/2026-02-28-when-agents-ask-for-help-teaching-llms-the-art-of-expert-collaboration/</link>
      <pubDate>Sat, 28 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-28-when-agents-ask-for-help-teaching-llms-the-art-of-expert-collaboration/</guid>
      <description>A mechanism-first reading of AHCE, a framework that teaches LLM agents when to escalate to human experts and how to turn messy advice into executable action.</description>
    </item>
    <item>
      <title>When Memory Thinks: Shrinking GRAVE Without Losing Its Mind</title>
      <link>https://cognaptus.com/blog/2026-02-27-when-memory-thinks-shrinking-grave-without-losing-its-mind/</link>
      <pubDate>Fri, 27 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-27-when-memory-thinks-shrinking-grave-without-losing-its-mind/</guid>
      <description>A mechanism-first reading of how GRAVE², GRAVER, and GRAVER² preserve search strength under tight memory budgets.</description>
    </item>
    <item>
      <title>When Fine-Tuning Bites Back: The Hidden Safety Drift in Vision-Language Agents</title>
      <link>https://cognaptus.com/blog/2026-02-21-when-finetuning-bites-back-the-hidden-safety-drift-in-visionlanguage-agents/</link>
      <pubDate>Sat, 21 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-21-when-finetuning-bites-back-the-hidden-safety-drift-in-visionlanguage-agents/</guid>
      <description>A mechanism-first reading of how narrow multimodal fine-tuning can turn a localized data problem into broad safety drift across vision-language agents.</description>
    </item>
    <item>
      <title>Ready Player None: Why AI Still Can’t Beat the Human Game Multiverse</title>
      <link>https://cognaptus.com/blog/2026-02-20-ready-player-none-why-ai-still-cant-beat-the-human-game-multiverse/</link>
      <pubDate>Fri, 20 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-20-ready-player-none-why-ai-still-cant-beat-the-human-game-multiverse/</guid>
      <description>AI GAMESTORE shows why frontier models still struggle with rapid learning, memory, planning, and world-model discovery in interactive tasks humans treat as casual.</description>
    </item>
    <item>
      <title>The Audit of Autonomy: When AI Agents Need More Than Intelligence</title>
      <link>https://cognaptus.com/blog/2026-02-20-the-audit-of-autonomy-when-ai-agents-need-more-than-intelligence/</link>
      <pubDate>Fri, 20 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-20-the-audit-of-autonomy-when-ai-agents-need-more-than-intelligence/</guid>
      <description>A practical reading of Policy Cards and why autonomous AI agents need machine-readable governance before they can become reliable business infrastructure.</description>
    </item>
    <item>
      <title>Certified to Speak: When AI Agents Need a Shared Dictionary</title>
      <link>https://cognaptus.com/blog/2026-02-19-certified-to-speak-when-ai-agents-need-a-shared-dictionary/</link>
      <pubDate>Thu, 19 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-19-certified-to-speak-when-ai-agents-need-a-shared-dictionary/</guid>
      <description>A mechanism-first reading of stimulus-meaning certification: how AI agents can test shared vocabulary before using it in consequential workflows.</description>
    </item>
    <item>
      <title>From Guesswork to Generative Foresight: Why Diffusion Models May Fix Multi-Agent Blind Spots</title>
      <link>https://cognaptus.com/blog/2026-02-18-from-guesswork-to-generative-foresight-why-diffusion-models-may-fix-multiagent-blind-spots/</link>
      <pubDate>Wed, 18 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-18-from-guesswork-to-generative-foresight-why-diffusion-models-may-fix-multiagent-blind-spots/</guid>
      <description>GlobeDiff shows why partial observability in multi-agent systems is less a memory problem than a generative state-inference problem.</description>
    </item>
    <item>
      <title>Sim2Realpolitik: Why Your AI Needs a Twin Before It Faces Reality</title>
      <link>https://cognaptus.com/blog/2026-02-18-sim2realpolitik-why-your-ai-needs-a-twin-before-it-faces-reality/</link>
      <pubDate>Wed, 18 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-18-sim2realpolitik-why-your-ai-needs-a-twin-before-it-faces-reality/</guid>
      <description>A mechanism-first reading of why simulated data and digital twins are becoming the rehearsal infrastructure for AI systems that must survive the real world.</description>
    </item>
    <item>
      <title>Proof Over Probabilities: Why AI Oversight Needs a Judge That Can Do Math</title>
      <link>https://cognaptus.com/blog/2026-02-13-proof-over-probabilities-why-ai-oversight-needs-a-judge-that-can-do-math/</link>
      <pubDate>Fri, 13 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-13-proof-over-probabilities-why-ai-oversight-needs-a-judge-that-can-do-math/</guid>
      <description>A mechanism-first reading of FORMALJUDGE, showing why safer AI-agent oversight may depend less on stronger judges and more on formally checkable constraints.</description>
    </item>
    <item>
      <title>Benchmarks Lie, Rooms Don’t: Why Embodied AI Fails the Moment It Enters Your House</title>
      <link>https://cognaptus.com/blog/2026-02-07-benchmarks-lie-rooms-dont-why-embodied-ai-fails-the-moment-it-enters-your-house/</link>
      <pubDate>Sat, 07 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-07-benchmarks-lie-rooms-dont-why-embodied-ai-fails-the-moment-it-enters-your-house/</guid>
      <description>A mechanism-first reading of TEA, an in-situ task-generation framework showing why embodied AI needs environment-specific evaluation before deployment.</description>
    </item>
    <item>
      <title>AgenticPay: When LLMs Start Haggling for a Living</title>
      <link>https://cognaptus.com/blog/2026-02-06-agenticpay-when-llms-start-haggling-for-a-living/</link>
      <pubDate>Fri, 06 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-06-agenticpay-when-llms-start-haggling-for-a-living/</guid>
      <description>AgenticPay shows why autonomous commercial negotiation requires more than fluent dialogue: it needs constraint discipline, role awareness, convergence control, and market-aware evaluation.</description>
    </item>
    <item>
      <title>FIRE-BENCH: Playing Back the Tape of Scientific Discovery</title>
      <link>https://cognaptus.com/blog/2026-02-05-firebench-playing-back-the-tape-of-scientific-discovery/</link>
      <pubDate>Thu, 05 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-05-firebench-playing-back-the-tape-of-scientific-discovery/</guid>
      <description>Why frontier research agents can write code, run experiments, and still fail at the part of science that actually matters: designing the right evidence and drawing the right conclusion.</description>
    </item>
    <item>
      <title>Click with Confidence: Teaching GUI Agents When *Not* to Click</title>
      <link>https://cognaptus.com/blog/2026-02-03-click-with-confidence-teaching-gui-agents-when-not-to-click/</link>
      <pubDate>Tue, 03 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-03-click-with-confidence-teaching-gui-agents-when-not-to-click/</guid>
      <description>SafeGround shows how uncertainty calibration can turn GUI agents from reckless clickers into risk-budgeted automation systems.</description>
    </item>
    <item>
      <title>RAudit: When Models Think Too Much and Still Get It Wrong</title>
      <link>https://cognaptus.com/blog/2026-02-03-raudit-when-models-think-too-much-and-still-get-it-wrong/</link>
      <pubDate>Tue, 03 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-03-raudit-when-models-think-too-much-and-still-get-it-wrong/</guid>
      <description>RAudit shows why longer reasoning, stronger judges, and harsher critique can reveal LLM failures—but can also amplify them.</description>
    </item>
    <item>
      <title>SokoBench: When Reasoning Models Lose the Plot</title>
      <link>https://cognaptus.com/blog/2026-01-31-sokobench-when-reasoning-models-lose-the-plot/</link>
      <pubDate>Sat, 31 Jan 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-01-31-sokobench-when-reasoning-models-lose-the-plot/</guid>
      <description>A mechanism-first reading of SokoBench, showing why long-horizon planning failures in reasoning models begin with fragile counting, state tracking, and world representation.</description>
    </item>
    <item>
      <title>Your Agent Remembers—But Can It Forget?</title>
      <link>https://cognaptus.com/blog/2026-01-22-your-agent-remembersbut-can-it-forget/</link>
      <pubDate>Thu, 22 Jan 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-01-22-your-agent-remembersbut-can-it-forget/</guid>
      <description>Why memory rewriting, not just memory retention, is becoming a hard diagnostic problem for reinforcement learning agents.</description>
    </item>
    <item>
      <title>When Memory Stops Guessing: Stitching Intent Back into Agent Memory</title>
      <link>https://cognaptus.com/blog/2026-01-17-when-memory-stops-guessing-stitching-intent-back-into-agent-memory/</link>
      <pubDate>Sat, 17 Jan 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-01-17-when-memory-stops-guessing-stitching-intent-back-into-agent-memory/</guid>
      <description>STITCH shows why long-horizon agents need memory indexed by task intent, not just larger context windows or better embeddings.</description>
    </item>
    <item>
      <title>Reasoning or Guessing? When Recursive Models Hit the Wrong Fixed Point</title>
      <link>https://cognaptus.com/blog/2026-01-16-reasoning-or-guessing-when-recursive-models-hit-the-wrong-fixed-point/</link>
      <pubDate>Fri, 16 Jan 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-01-16-reasoning-or-guessing-when-recursive-models-hit-the-wrong-fixed-point/</guid>
      <description>A mechanistic reading of HRM shows why recursive depth can look like reasoning while behaving more like attractor search—and how that changes reliability testing for business AI systems.</description>
    </item>
    <item>
      <title>Lean LLMs, Heavy Lifting: When Workflows Beat Bigger Models</title>
      <link>https://cognaptus.com/blog/2026-01-15-lean-llms-heavy-lifting-when-workflows-beat-bigger-models/</link>
      <pubDate>Thu, 15 Jan 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-01-15-lean-llms-heavy-lifting-when-workflows-beat-bigger-models/</guid>
      <description>A case-first look at why structured workflows and data tools, not just larger models, are the real bottleneck-breakers for large-scale optimization modeling.</description>
    </item>
    <item>
      <title>Think Before You Sink: Streaming Hallucinations in Long Reasoning</title>
      <link>https://cognaptus.com/blog/2026-01-06-think-before-you-sink-streaming-hallucinations-in-long-reasoning/</link>
      <pubDate>Tue, 06 Jan 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-01-06-think-before-you-sink-streaming-hallucinations-in-long-reasoning/</guid>
      <description>A mechanism-first reading of why long chain-of-thought hallucinations behave like evolving states, and how streaming hidden-state probes could turn reasoning reliability into an operational signal.</description>
    </item>
    <item>
      <title>Talking to Yourself, but Make It Useful: Intrinsic Self‑Critique in LLM Planning</title>
      <link>https://cognaptus.com/blog/2026-01-03-talking-to-yourself-but-make-it-useful-intrinsic-selfcritique-in-llm-planning/</link>
      <pubDate>Sat, 03 Jan 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-01-03-talking-to-yourself-but-make-it-useful-intrinsic-selfcritique-in-llm-planning/</guid>
      <description>A procedural self-critique loop can make LLM planners markedly more reliable—but only when reflection is converted into explicit rule checking, state tracking, and conservative approval.</description>
    </item>
    <item>
      <title>SpatialBench: When AI Meets Messy Biology</title>
      <link>https://cognaptus.com/blog/2025-12-29-spatialbench-when-ai-meets-messy-biology/</link>
      <pubDate>Mon, 29 Dec 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-12-29-spatialbench-when-ai-meets-messy-biology/</guid>
      <description>SpatialBench shows why reliable scientific AI agents need domain calibration, workflow control, and verifiable execution—not just stronger base models.</description>
    </item>
    <item>
      <title>When Actions Need Nuance: Learning to Act Precisely Only When It Matters</title>
      <link>https://cognaptus.com/blog/2025-12-28-when-actions-need-nuance-learning-to-act-precisely-only-when-it-matters/</link>
      <pubDate>Sun, 28 Dec 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-12-28-when-actions-need-nuance-learning-to-act-precisely-only-when-it-matters/</guid>
      <description>Why PEARL’s context-sensitive abstractions point to a more efficient way of learning hybrid actions: precise control only where precision changes the outcome.</description>
    </item>
    <item>
      <title>AGI by Committee: Why the First General Intelligence Won’t Arrive Alone</title>
      <link>https://cognaptus.com/blog/2025-12-19-agi-by-committee-why-the-first-general-intelligence-wont-arrive-alone/</link>
      <pubDate>Fri, 19 Dec 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-12-19-agi-by-committee-why-the-first-general-intelligence-wont-arrive-alone/</guid>
      <description>A mechanism-first reading of patchwork AGI: why collective agent systems may become the real control surface for safety, governance, and enterprise deployment.</description>
    </item>
    <item>
      <title>Delegating to the Almost-Aligned: When Misaligned AI Is Still the Rational Choice</title>
      <link>https://cognaptus.com/blog/2025-12-18-delegating-to-the-almostaligned-when-misaligned-ai-is-still-the-rational-choice/</link>
      <pubDate>Thu, 18 Dec 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-12-18-delegating-to-the-almostaligned-when-misaligned-ai-is-still-the-rational-choice/</guid>
      <description>A decision-theoretic guide to deciding when imperfectly aligned AI systems are still worth delegating to.</description>
    </item>
    <item>
      <title>Ports, But Make Them Agentic: When LLMs Start Running the Yard</title>
      <link>https://cognaptus.com/blog/2025-12-17-ports-but-make-them-agentic-when-llms-start-running-the-yard/</link>
      <pubDate>Wed, 17 Dec 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-12-17-ports-but-make-them-agentic-when-llms-start-running-the-yard/</guid>
      <description>PortAgent shows how LLM agents can compress vehicle-dispatch deployment by combining retrieval, modeling, code generation, and execution-based correction.</description>
    </item>
    <item>
      <title>Bits, Bets, and Budgets: When Agents Should Walk Away</title>
      <link>https://cognaptus.com/blog/2025-12-09-bits-bets-and-budgets-when-agents-should-walk-away/</link>
      <pubDate>Tue, 09 Dec 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-12-09-bits-bets-and-budgets-when-agents-should-walk-away/</guid>
      <description>A mechanism-first reading of the Agent Capability Problem: how information, cost, and uncertainty can help decide whether an AI agent should proceed, approximate, redesign, or stop.</description>
    </item>
    <item>
      <title>Breaking Rules, Not Systems: How Penalties Make Autonomous Agents Behave</title>
      <link>https://cognaptus.com/blog/2025-12-04-breaking-rules-not-systems-how-penalties-make-autonomous-agents-behave/</link>
      <pubDate>Thu, 04 Dec 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-12-04-breaking-rules-not-systems-how-penalties-make-autonomous-agents-behave/</guid>
      <description>A case-first reading of how penalty-aware policy reasoning lets autonomous agents distinguish acceptable emergency exceptions from dangerous rule-breaking.</description>
    </item>
    <item>
      <title>Prompting on Life Support: How Invasive Context Engineering Fights Long-Context Drift</title>
      <link>https://cognaptus.com/blog/2025-12-03-prompting-on-life-support-how-invasive-context-engineering-fights-longcontext-drift/</link>
      <pubDate>Wed, 03 Dec 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-12-03-prompting-on-life-support-how-invasive-context-engineering-fights-longcontext-drift/</guid>
      <description>A mechanism-first reading of Invasive Context Engineering, a training-free proposal for keeping LLM control instructions alive inside long conversations and agentic reasoning loops.</description>
    </item>
    <item>
      <title>Debate Club for Robots: How Multi-Agent Arguing Makes Embodied AI Safer</title>
      <link>https://cognaptus.com/blog/2025-11-28-debate-club-for-robots-how-multiagent-arguing-makes-embodied-ai-safer/</link>
      <pubDate>Fri, 28 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-28-debate-club-for-robots-how-multiagent-arguing-makes-embodied-ai-safer/</guid>
      <description>A mechanism-first reading of MADRA, a training-free multi-agent debate system that treats embodied AI safety as a decision-gate problem rather than a stronger-prompt problem.</description>
    </item>
    <item>
      <title>Fragments, Feedback, and Fast Drugs: When Generative Models Grow a Spine</title>
      <link>https://cognaptus.com/blog/2025-11-26-fragments-feedback-and-fast-drugs-when-generative-models-grow-a-spine/</link>
      <pubDate>Wed, 26 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-26-fragments-feedback-and-fast-drugs-when-generative-models-grow-a-spine/</guid>
      <description>How FRAGMENTA reframes small-data drug lead optimization as a feedback-loop problem, not merely a bigger-model problem.</description>
    </item>
    <item>
      <title>Benchmarks Without Borders: Inside the Moduli Space of AI Psychometrics</title>
      <link>https://cognaptus.com/blog/2025-11-25-benchmarks-without-borders-inside-the-moduli-space-of-ai-psychometrics/</link>
      <pubDate>Tue, 25 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-25-benchmarks-without-borders-inside-the-moduli-space-of-ai-psychometrics/</guid>
      <description>A mechanism-first guide to why AI-agent evaluation should measure structured coverage across benchmark families, not worship individual benchmark scores.</description>
    </item>
    <item>
      <title>Pop-Ups, Pitfalls, and Planning: Why GUI Agents Break in the Real World</title>
      <link>https://cognaptus.com/blog/2025-11-22-popups-pitfalls-and-planning-why-gui-agents-break-in-the-real-world/</link>
      <pubDate>Sat, 22 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-22-popups-pitfalls-and-planning-why-gui-agents-break-in-the-real-world/</guid>
      <description>D-GARA shows that GUI-agent reliability is not measured by clean task completion, but by whether an agent can recover when real interfaces interrupt, redirect, and reset its plan.</description>
    </item>
    <item>
      <title>Thresholds, Trade-offs, and the Art of Not Overthinking Your Robot</title>
      <link>https://cognaptus.com/blog/2025-11-20-thresholds-tradeoffs-and-the-art-of-not-overthinking-your-robot/</link>
      <pubDate>Thu, 20 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-20-thresholds-tradeoffs-and-the-art-of-not-overthinking-your-robot/</guid>
      <description>How calibrated symbolic uncertainty helps robots decide when to act, when to look again, and when confidence becomes expensive.</description>
    </item>
    <item>
      <title>Mind the Gap: When Robots Learn Social Norms the Human Way</title>
      <link>https://cognaptus.com/blog/2025-11-17-mind-the-gap-when-robots-learn-social-norms-the-human-way/</link>
      <pubDate>Mon, 17 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-17-mind-the-gap-when-robots-learn-social-norms-the-human-way/</guid>
      <description>A business-focused analysis of how psychologically grounded reward design can make robot navigation more socially acceptable without pretending VR comfort scores are field deployment proof.</description>
    </item>
    <item>
      <title>Steering the Schemer: How Test-Time Alignment Tames Machiavellian Agents</title>
      <link>https://cognaptus.com/blog/2025-11-17-steering-the-schemer-how-testtime-alignment-tames-machiavellian-agents/</link>
      <pubDate>Mon, 17 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-17-steering-the-schemer-how-testtime-alignment-tames-machiavellian-agents/</guid>
      <description>A mechanism-first look at how test-time policy shaping can steer reward-maximising agents away from harmful behaviour without retraining them.</description>
    </item>
    <item>
      <title>Scalpels, Agents, and Orchestrators: When Surgery Meets Autonomous Workflows</title>
      <link>https://cognaptus.com/blog/2025-11-16-scalpels-agents-and-orchestrators-when-surgery-meets-autonomous-workflows/</link>
      <pubDate>Sun, 16 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-16-scalpels-agents-and-orchestrators-when-surgery-meets-autonomous-workflows/</guid>
      <description>A mechanism-first reading of how VISA uses LLM orchestration, memory, and specialised agents to make voice-controlled surgical data interfaces more reliable.</description>
    </item>
    <item>
      <title>Plans, Tokens, and Turing Dreams: Why LLMs Still Can’t Out-Plan a 15-Year-Old Classical Planner</title>
      <link>https://cognaptus.com/blog/2025-11-13-plans-tokens-and-turing-dreams-why-llms-still-cant-outplan-a-15yearold-classical-planner/</link>
      <pubDate>Thu, 13 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-13-plans-tokens-and-turing-dreams-why-llms-still-cant-outplan-a-15yearold-classical-planner/</guid>
      <description>A business-readable analysis of why frontier LLMs are getting better at formal planning, but still need symbolic validation before they touch real operations.</description>
    </item>
    <item>
      <title>When Heuristics Go Silent: How Random Walks Outsmart Breadth-First Search</title>
      <link>https://cognaptus.com/blog/2025-11-13-when-heuristics-go-silent-how-random-walks-outsmart-breadthfirst-search/</link>
      <pubDate>Thu, 13 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-13-when-heuristics-go-silent-how-random-walks-outsmart-breadthfirst-search/</guid>
      <description>A mechanism-first reading of when restarting random walks beat breadth-first search for escaping heuristic dead zones in AI planning.</description>
    </item>
    <item>
      <title>When Agents Think in Waves: Diffusion Models for Ad Hoc Teamwork</title>
      <link>https://cognaptus.com/blog/2025-11-11-when-agents-think-in-waves-diffusion-models-for-ad-hoc-teamwork/</link>
      <pubDate>Tue, 11 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-11-when-agents-think-in-waves-diffusion-models-for-ad-hoc-teamwork/</guid>
      <description>A mechanism-first reading of PADiff, showing why diffusion policies may help agents preserve multiple cooperation plans when working with unfamiliar teammates.</description>
    </item>
    <item>
      <title>When Algorithms Command: AI&#39;s Quiet Revolution in Battlefield Strategy</title>
      <link>https://cognaptus.com/blog/2025-11-10-when-algorithms-command-ais-quiet-revolution-in-battlefield-strategy/</link>
      <pubDate>Mon, 10 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-10-when-algorithms-command-ais-quiet-revolution-in-battlefield-strategy/</guid>
      <description>A mechanism-first reading of a battlefield decision-support prototype, and what it teaches business leaders about simulation, option generation, and human control.</description>
    </item>
    <item>
      <title>The Doctor Is In: How DR. WELL Heals Multi-Agent Coordination with Symbolic Memory</title>
      <link>https://cognaptus.com/blog/2025-11-07-the-doctor-is-in-how-dr-well-heals-multiagent-coordination-with-symbolic-memory/</link>
      <pubDate>Fri, 07 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-07-the-doctor-is-in-how-dr-well-heals-multiagent-coordination-with-symbolic-memory/</guid>
      <description>DR. WELL shows how multi-agent LLM systems can coordinate more reliably by replacing free-form chatter with negotiated commitments, symbolic plans, and shared operational memory.</description>
    </item>
    <item>
      <title>When AI Becomes Its Own Research Assistant</title>
      <link>https://cognaptus.com/blog/2025-11-07-when-ai-becomes-its-own-research-assistant/</link>
      <pubDate>Fri, 07 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-07-when-ai-becomes-its-own-research-assistant/</guid>
      <description>Jr. AI Scientist shows that autonomous research agents are becoming useful apprentices, but only when their work is tightly scoped, inspected, and treated as evidence to audit rather than truth to publish.</description>
    </item>
    <item>
      <title>When Rules Go Live: Policy Cards and the New Language of AI Governance</title>
      <link>https://cognaptus.com/blog/2025-11-02-when-rules-go-live-policy-cards-and-the-new-language-of-ai-governance/</link>
      <pubDate>Sun, 02 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-02-when-rules-go-live-policy-cards-and-the-new-language-of-ai-governance/</guid>
      <description>Policy Cards turn AI governance from scattered compliance prose into a machine-readable runtime contract for autonomous agents.</description>
    </item>
    <item>
      <title>Sketching a Thought: How Mental Imagery Could Unlock Autonomous Machine Reasoning</title>
      <link>https://cognaptus.com/blog/2025-07-18-sketching-a-thought-how-mental-imagery-could-unlock-autonomous-machine-reasoning/</link>
      <pubDate>Fri, 18 Jul 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-07-18-sketching-a-thought-how-mental-imagery-could-unlock-autonomous-machine-reasoning/</guid>
      <description>A mechanism-first reading of a machine-thinking framework that treats mental imagery as an architectural loop, not a benchmark-proven shortcut to reasoning.</description>
    </item>
    <item>
      <title>Good AI Goes Rogue: Why Intelligent Disobedience May Be the Key to Trustworthy Teammates</title>
      <link>https://cognaptus.com/blog/2025-06-30-good-ai-goes-rogue-why-intelligent-disobedience-may-be-the-key-to-trustworthy-teammates/</link>
      <pubDate>Mon, 30 Jun 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-06-30-good-ai-goes-rogue-why-intelligent-disobedience-may-be-the-key-to-trustworthy-teammates/</guid>
      <description>A mechanism-first reading of why trustworthy AI teammates may need bounded, transparent ways to refuse, interrupt, escalate, or override human instructions.</description>
    </item>
    <item>
      <title>Rules of Engagement: Why LLMs Need Logic to Plan</title>
      <link>https://cognaptus.com/blog/2025-04-02-rules-of-engagement-why-llms-need-logic-to-plan/</link>
      <pubDate>Wed, 02 Apr 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-04-02-rules-of-engagement-why-llms-need-logic-to-plan/</guid>
      <description>ACPBench Hard shows that strong language models still miss basic planner primitives, making symbolic validation more useful than another polished agent demo.</description>
    </item>
  </channel>
</rss>
