<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>LLM Agents on Cognaptus</title>
    <link>https://cognaptus.com/tags/llm-agents/</link>
    <description>Recent content in LLM Agents on Cognaptus</description>
    <generator>Hugo -- 0.145.0</generator>
    <language>en-us</language>
    <lastBuildDate>Mon, 07 Sep 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://cognaptus.com/tags/llm-agents/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Reasoning Under a Running Clock: Why Agent Rankings Reverse in Real Time</title>
      <link>https://cognaptus.com/blog/2026-09-07-reasoning-under-a-running-clock-why-agent-rankings-reverse-in-real-time/</link>
      <pubDate>Mon, 07 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-09-07-reasoning-under-a-running-clock-why-agent-rankings-reverse-in-real-time/</guid>
      <description>STAR shows why model selection for time-sensitive agents must account for inference latency, action throughput, and execution quality—not reasoning strength alone.</description>
    </item>
    <item>
      <title>When Model Output Can Change State: An Architecture Guide to Agent Reliability</title>
      <link>https://cognaptus.com/blog/2026-09-06-when-model-output-can-change-state-an-architecture-guide-to-agent-reliability/</link>
      <pubDate>Sun, 06 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-09-06-when-model-output-can-change-state-an-architecture-guide-to-agent-reliability/</guid>
      <description>A systems view of Agentic AI shows where autonomy creates operational risk—and where architecture can contain it.</description>
    </item>
    <item>
      <title>Train the Agent Before the Sandbox Exists</title>
      <link>https://cognaptus.com/blog/2026-09-03-train-the-agent-before-the-sandbox-exists/</link>
      <pubDate>Thu, 03 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-09-03-train-the-agent-before-the-sandbox-exists/</guid>
      <description>ESAT shows that API specifications can become a training asset before executable backends and realistic sandbox state are ready.</description>
    </item>
    <item>
      <title>More Memory, Worse Decisions: Why Agent Recall Needs Routing</title>
      <link>https://cognaptus.com/blog/2026-08-27-more-memory-worse-decisions-why-agent-recall-needs-routing/</link>
      <pubDate>Thu, 27 Aug 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-08-27-more-memory-worse-decisions-why-agent-recall-needs-routing/</guid>
      <description>A controlled comparison of agent memory systems shows that retrieving more history can improve factual recall while degrading action selection, making memory a routing decision rather than a universal architecture choice.</description>
    </item>
    <item>
      <title>Decision Rights, Not More Layers: What an Auditable Fraud Pipeline Actually Earns</title>
      <link>https://cognaptus.com/blog/2026-08-24-decision-rights-not-more-layers-what-an-auditable-fraud-pipeline-actually-earns/</link>
      <pubDate>Mon, 24 Aug 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-08-24-decision-rights-not-more-layers-what-an-auditable-fraud-pipeline-actually-earns/</guid>
      <description>A fraud pipeline study shows why graph signals, anomaly detection, and LLM investigators should earn narrowly defined decision roles rather than automatic authority.</description>
    </item>
    <item>
      <title>Higher Pass Rate, More Broken Tasks: The Regression Tax in Agent Skill Libraries</title>
      <link>https://cognaptus.com/blog/2026-08-19-higher-pass-rate-more-broken-tasks-the-regression-tax-in-agent-skill-libraries/</link>
      <pubDate>Wed, 19 Aug 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-08-19-higher-pass-rate-more-broken-tasks-the-regression-tax-in-agent-skill-libraries/</guid>
      <description>Skill libraries can raise average agent performance while breaking workflows that already worked; paired evaluation reveals how large that reliability cost can be.</description>
    </item>
    <item>
      <title>The Trace Has the Answer, Not the Alternative: Agentic-DPO for Offline Agent Training</title>
      <link>https://cognaptus.com/blog/2026-08-19-the-trace-has-the-answer-not-the-alternative-agenticdpo-for-offline-agent-training/</link>
      <pubDate>Wed, 19 Aug 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-08-19-the-trace-has-the-answer-not-the-alternative-agenticdpo-for-offline-agent-training/</guid>
      <description>Agentic-DPO shows how expert traces can supervise the mistakes an agent is likely to make, without requiring full online rollouts during training.</description>
    </item>
    <item>
      <title>The Catalog Grew. The Agent Needed a Call Stack.</title>
      <link>https://cognaptus.com/blog/2026-08-03-the-catalog-grew-the-agent-needed-a-call-stack/</link>
      <pubDate>Mon, 03 Aug 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-08-03-the-catalog-grew-the-agent-needed-a-call-stack/</guid>
      <description>A hierarchical agent architecture sharply reduces tool-schema exposure at scale, but only when taxonomy, validation, and latency controls are engineered with equal care.</description>
    </item>
    <item>
      <title>Same Agent, Different Audience: When Social Pressure Changes the Recommendation</title>
      <link>https://cognaptus.com/blog/2026-07-30-same-agent-different-audience-when-social-pressure-changes-the-recommendation/</link>
      <pubDate>Thu, 30 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-07-30-same-agent-different-audience-when-social-pressure-changes-the-recommendation/</guid>
      <description>A dual-channel benchmark shows how authority, sponsorship, and future dependence can alter an agent’s public recommendation without changing its model or stated task.</description>
    </item>
    <item>
      <title>Refusal Is Not a Result: Vera Tests What Agents Actually Changed</title>
      <link>https://cognaptus.com/blog/2026-07-29-refusal-is-not-a-result-vera-tests-what-agents-actually-changed/</link>
      <pubDate>Wed, 29 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-07-29-refusal-is-not-a-result-vera-tests-what-agents-actually-changed/</guid>
      <description>Vera reframes agent safety evaluation as reproducible software testing built around observable effects, adaptive attacks, and deterministic verification.</description>
    </item>
    <item>
      <title>The Fine Print Is the Task: Why Long-Context AI Fails After Finding the Answer</title>
      <link>https://cognaptus.com/blog/2026-07-22-the-fine-print-is-the-task-why-longcontext-ai-fails-after-finding-the-answer/</link>
      <pubDate>Wed, 22 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-07-22-the-fine-print-is-the-task-why-longcontext-ai-fails-after-finding-the-answer/</guid>
      <description>A large-scale study shows that long-context AI often retrieves the right facts but misses the local rules that determine whether its answer is actually valid.</description>
    </item>
    <item>
      <title>Skill Issue, Literally: Repairing Agent Instructions Without an Answer Key</title>
      <link>https://cognaptus.com/blog/2026-07-03-skill-issue-literally-repairing-agent-instructions-without-an-answer-key/</link>
      <pubDate>Fri, 03 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-07-03-skill-issue-literally-repairing-agent-instructions-without-an-answer-key/</guid>
      <description>SkillAudit shows how agent skills can be improved without hidden tests by comparing with-skill and without-skill executions, but only when correctness leaves an observable trace.</description>
    </item>
    <item>
      <title>When &#39;Check the AC&#39; Becomes the Hard Part</title>
      <link>https://cognaptus.com/blog/2026-06-25-when-check-the-ac-becomes-the-hard-part/</link>
      <pubDate>Thu, 25 Jun 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-06-25-when-check-the-ac-becomes-the-hard-part/</guid>
      <description>A mechanism-first reading of PEC-Home, showing why smart-home assistants fail when users naturally compress repeated commands into household shorthand.</description>
    </item>
    <item>
      <title>Feedback Is the New Attack Surface</title>
      <link>https://cognaptus.com/blog/2026-06-23-feedback-is-the-new-attack-surface/</link>
      <pubDate>Tue, 23 Jun 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-06-23-feedback-is-the-new-attack-surface/</guid>
      <description>Why automated prompt-injection risk is not just about malicious prompts, but about the feedback loops that let attackers optimize against agentic systems.</description>
    </item>
    <item>
      <title>Ground Control to Synthetic Data: Why Enterprise LLMs Need a Source of Truth</title>
      <link>https://cognaptus.com/blog/2026-06-21-ground-control-to-synthetic-data-why-enterprise-llms-need-a-source-of-truth/</link>
      <pubDate>Sun, 21 Jun 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-06-21-ground-control-to-synthetic-data-why-enterprise-llms-need-a-source-of-truth/</guid>
      <description>Synthetic data only becomes useful for enterprise AI when it is grounded in real system structure, verified for meaning, and filtered before training.</description>
    </item>
    <item>
      <title>Less Prompt, More Blueprint: MOSAIC and the Data-Science Agent That Keeps Receipts</title>
      <link>https://cognaptus.com/blog/2026-06-20-less-prompt-more-blueprint-mosaic-and-the-datascience-agent-that-keeps-receipts/</link>
      <pubDate>Sat, 20 Jun 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-06-20-less-prompt-more-blueprint-mosaic-and-the-datascience-agent-that-keeps-receipts/</guid>
      <description>MOSAIC shows how agentic data science becomes more useful when model-building is treated as reusable workflow construction, not free-form code generation.</description>
    </item>
    <item>
      <title>Logs Are Not Lineage: The Accountability Layer AI Agents Are Missing</title>
      <link>https://cognaptus.com/blog/2026-06-16-logs-are-not-lineage-the-accountability-layer-ai-agents-are-missing/</link>
      <pubDate>Tue, 16 Jun 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-06-16-logs-are-not-lineage-the-accountability-layer-ai-agents-are-missing/</guid>
      <description>A practical reading of evidence tracing and execution provenance as the infrastructure layer that turns opaque AI-agent activity into auditable, controllable business systems.</description>
    </item>
    <item>
      <title>The Solver Isn’t the Strategy: FrontierOR’s Reality Check for AI Optimisation Agents</title>
      <link>https://cognaptus.com/blog/2026-06-14-the-solver-isnt-the-strategy-frontierors-reality-check-for-ai-optimisation-agents/</link>
      <pubDate>Sun, 14 Jun 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-06-14-the-solver-isnt-the-strategy-frontierors-reality-check-for-ai-optimisation-agents/</guid>
      <description>FrontierOR shows why runnable optimisation code is not the same as scalable algorithm design, and why enterprise AI agents need harder tests than solver demos.</description>
    </item>
    <item>
      <title>Commit Issues: Why Multi-Agent AI Needs Typed Finality, Not Another Vote</title>
      <link>https://cognaptus.com/blog/2026-06-11-commit-issues-why-multiagent-ai-needs-typed-finality-not-another-vote/</link>
      <pubDate>Thu, 11 Jun 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-06-11-commit-issues-why-multiagent-ai-needs-typed-finality-not-another-vote/</guid>
      <description>A mechanism-first reading of H-CSC, a protocol that separates what AI agents decide from what kind of agreement their decision can honestly claim.</description>
    </item>
    <item>
      <title>Prompt and Order: Why LLM Trading Needs a Factory, Not a Fortune Teller</title>
      <link>https://cognaptus.com/blog/2026-06-11-prompt-and-order-why-llm-trading-needs-a-factory-not-a-fortune-teller/</link>
      <pubDate>Thu, 11 Jun 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-06-11-prompt-and-order-why-llm-trading-needs-a-factory-not-a-fortune-teller/</guid>
      <description>A mechanism-first reading of MadEvolve shows why LLMs are more useful as governed search engines for trading-system design than as magical alpha machines.</description>
    </item>
    <item>
      <title>Search, Critique, Repeat: Critic-R Turns RAG Complaints into Retriever Training</title>
      <link>https://cognaptus.com/blog/2026-06-08-search-critique-repeat-criticr-turns-rag-complaints-into-retriever-training/</link>
      <pubDate>Mon, 08 Jun 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-06-08-search-critique-repeat-criticr-turns-rag-complaints-into-retriever-training/</guid>
      <description>A mechanism-first reading of Critic-R, a framework that uses agent introspection to repair retrieval at inference time and train better retrievers without gold passage labels.</description>
    </item>
    <item>
      <title>Scaffold and Ladder: Why AI Agents Need Meta-Reasoning, Not Longer Monologues</title>
      <link>https://cognaptus.com/blog/2026-06-01-scaffold-and-ladder-why-ai-agents-need-metareasoning-not-longer-monologues/</link>
      <pubDate>Mon, 01 Jun 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-06-01-scaffold-and-ladder-why-ai-agents-need-metareasoning-not-longer-monologues/</guid>
      <description>A mechanism-first reading of Deep Reasoning and Dolores, showing why agent reliability may depend less on longer thinking and more on executable task-specific decomposition.</description>
    </item>
    <item>
      <title>Think Longer, Act Smarter: Why Coding Agents Need Behavior-Preserving Reasoning</title>
      <link>https://cognaptus.com/blog/2026-05-31-think-longer-act-smarter-why-coding-agents-need-behaviorpreserving-reasoning/</link>
      <pubDate>Sun, 31 May 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-05-31-think-longer-act-smarter-why-coding-agents-need-behaviorpreserving-reasoning/</guid>
      <description>A mechanism-first reading of M2A, a training-free method for injecting mathematical reasoning into coding agents without breaking their think-act-observe loop.</description>
    </item>
    <item>
      <title>Don’t Average the Needle: Spectral Retrieval and the RAG Evidence Problem</title>
      <link>https://cognaptus.com/blog/2026-05-30-dont-average-the-needle-spectral-retrieval-and-the-rag-evidence-problem/</link>
      <pubDate>Sat, 30 May 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-05-30-dont-average-the-needle-spectral-retrieval-and-the-rag-evidence-problem/</guid>
      <description>A mechanism-first reading of Spectral Retrieval: why dense retrieval can bury localized evidence, how multi-scale sinc convolution tries to recover it, and where the business value actually begins.</description>
    </item>
    <item>
      <title>Experience Is Not Memory: Why Learning Agents Need a Better Feedback Loop</title>
      <link>https://cognaptus.com/blog/2026-05-29-experience-is-not-memory-why-learning-agents-need-a-better-feedback-loop/</link>
      <pubDate>Fri, 29 May 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-05-29-experience-is-not-memory-why-learning-agents-need-a-better-feedback-loop/</guid>
      <description>A mechanism-first reading of In-context Training, a new framework for testing whether language agents can turn one-off experience into reusable operational improvement.</description>
    </item>
    <item>
      <title>Think Longer, Act Worse? What M2A Teaches About Reasoning Agents</title>
      <link>https://cognaptus.com/blog/2026-05-29-think-longer-act-worse-what-m2a-teaches-about-reasoning-agents/</link>
      <pubDate>Fri, 29 May 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-05-29-think-longer-act-worse-what-m2a-teaches-about-reasoning-agents/</guid>
      <description>A mechanism-first reading of M2A, showing why better reasoning agents need protected action loops, not just longer thought traces.</description>
    </item>
    <item>
      <title>Silent Errors, Loud Consequences: ASMR-Bench and the Coming Era of AI Auditors</title>
      <link>https://cognaptus.com/blog/2026-04-22-silent-errors-loud-consequences-asmrbench-and-the-coming-era-of-ai-auditors/</link>
      <pubDate>Wed, 22 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-22-silent-errors-loud-consequences-asmrbench-and-the-coming-era-of-ai-auditors/</guid>
      <description>A research-sabotage benchmark shows why AI auditability is not a code-review feature, but an operating model for trustworthy AI work.</description>
    </item>
    <item>
      <title>Reviewer, Reviewed: When AI Starts Grading the Graders</title>
      <link>https://cognaptus.com/blog/2026-04-16-reviewer-reviewed-when-ai-starts-grading-the-graders/</link>
      <pubDate>Thu, 16 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-16-reviewer-reviewed-when-ai-starts-grading-the-graders/</guid>
      <description>A field deployment of AI-generated peer review at AAAI-26 shows where AI can outperform human reviewers, where it still fails, and what businesses should learn about governed second-opinion systems.</description>
    </item>
    <item>
      <title>Evolve or Die Trying: When LLMs Stop Writing Code and Start Designing Algorithms</title>
      <link>https://cognaptus.com/blog/2026-04-15-evolve-or-die-trying-when-llms-stop-writing-code-and-start-designing-algorithms/</link>
      <pubDate>Wed, 15 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-15-evolve-or-die-trying-when-llms-stop-writing-code-and-start-designing-algorithms/</guid>
      <description>BEAM shows that useful LLM algorithm design is less about clever prompting and more about structured search, reusable memory, and evaluation that actually resembles solver construction.</description>
    </item>
    <item>
      <title>The Memory Isn’t Broken — It’s Flat: Why LLMs Need to ‘Draw’ to Remember</title>
      <link>https://cognaptus.com/blog/2026-04-15-the-memory-isnt-broken-its-flat-why-llms-need-to-draw-to-remember/</link>
      <pubDate>Wed, 15 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-15-the-memory-isnt-broken-its-flat-why-llms-need-to-draw-to-remember/</guid>
      <description>A mechanism-first reading of dual-trace memory encoding and why enterprise AI agents may need richer contextual traces, not just larger memory stores.</description>
    </item>
    <item>
      <title>CivBench: When AI Stops Guessing and Starts Planning</title>
      <link>https://cognaptus.com/blog/2026-04-11-civbench-when-ai-stops-guessing-and-starts-planning/</link>
      <pubDate>Sat, 11 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-11-civbench-when-ai-stops-guessing-and-starts-planning/</guid>
      <description>CivBench shows why serious agent evaluation needs progress signals, not just final scoreboards.</description>
    </item>
    <item>
      <title>Feeling the Model: When LLMs Don’t Just Predict — They ‘Feel’</title>
      <link>https://cognaptus.com/blog/2026-04-11-feeling-the-model-when-llms-dont-just-predict-they-feel/</link>
      <pubDate>Sat, 11 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-11-feeling-the-model-when-llms-dont-just-predict-they-feel/</guid>
      <description>Anthropic’s emotion-vector study shows why enterprise AI risk is not only about bad prompts or bad outputs, but about hidden internal states that can steer agents toward shortcuts, sycophancy, and coercive behavior.</description>
    </item>
    <item>
      <title>Mind the Cut: Where Your AI Strategy Quietly Breaks</title>
      <link>https://cognaptus.com/blog/2026-04-11-mind-the-cut-where-your-ai-strategy-quietly-breaks/</link>
      <pubDate>Sat, 11 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-11-mind-the-cut-where-your-ai-strategy-quietly-breaks/</guid>
      <description>A business-oriented reading of the Cartesian cut: why the boundary between model and runtime determines whether AI agents remain governable, brittle, or truly autonomous.</description>
    </item>
    <item>
      <title>The Orchestrator Problem: When AI Meets Exascale Reality</title>
      <link>https://cognaptus.com/blog/2026-04-11-the-orchestrator-problem-when-ai-meets-exascale-reality/</link>
      <pubDate>Sat, 11 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-11-the-orchestrator-problem-when-ai-meets-exascale-reality/</guid>
      <description>A mechanism-first reading of how LLM agents become useful for scientific computing only when they stop pretending to be schedulers.</description>
    </item>
    <item>
      <title>The Persuasion Engine: When AI Starts Selling (More Than Just Answers)</title>
      <link>https://cognaptus.com/blog/2026-04-10-the-persuasion-engine-when-ai-starts-selling-more-than-just-answers/</link>
      <pubDate>Fri, 10 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-10-the-persuasion-engine-when-ai-starts-selling-more-than-just-answers/</guid>
      <description>A mechanism-first reading of how sponsored incentives can distort AI assistants before they ever need to lie.</description>
    </item>
    <item>
      <title>From Chains to Trees: Why LLM Agents Need Structural Memory</title>
      <link>https://cognaptus.com/blog/2026-04-09-from-chains-to-trees-why-llm-agents-need-structural-memory/</link>
      <pubDate>Thu, 09 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-09-from-chains-to-trees-why-llm-agents-need-structural-memory/</guid>
      <description>A mechanism-first reading of T-STAR, showing why multi-turn LLM agents learn better when failed and successful rollouts are compared as shared trees rather than isolated chains.</description>
    </item>
    <item>
      <title>The Map Is Not the Territory—But Your LLM Thinks It Is</title>
      <link>https://cognaptus.com/blog/2026-04-09-the-map-is-not-the-territorybut-your-llm-thinks-it-is/</link>
      <pubDate>Thu, 09 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-09-the-map-is-not-the-territorybut-your-llm-thinks-it-is/</guid>
      <description>EVGeoQA shows why tool-using LLM agents still struggle with real-world spatial planning: they can reason locally, but often fail to explore enough.</description>
    </item>
    <item>
      <title>The Minimal LLM Thesis: When Agents Think for Themselves</title>
      <link>https://cognaptus.com/blog/2026-04-09-the-minimal-llm-thesis-when-agents-think-for-themselves/</link>
      <pubDate>Thu, 09 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-09-the-minimal-llm-thesis-when-agents-think-for-themselves/</guid>
      <description>A decomposition study shows why agent performance may come from measurable harness structure before it comes from larger or more frequent LLM calls.</description>
    </item>
    <item>
      <title>Benchmarking the Benchmarks: Why ACE-Bench Might Be the Missing Layer in Agent Evaluation</title>
      <link>https://cognaptus.com/blog/2026-04-08-benchmarking-the-benchmarks-why-acebench-might-be-the-missing-layer-in-agent-evaluation/</link>
      <pubDate>Wed, 08 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-08-benchmarking-the-benchmarks-why-acebench-might-be-the-missing-layer-in-agent-evaluation/</guid>
      <description>A mechanism-first reading of AgentCE-Bench, showing why controllable agent evaluation may be more useful than another realism-heavy leaderboard.</description>
    </item>
    <item>
      <title>From Spreadsheets to Swarms: How Agentic AI Rewrites the Retail Supply Chain</title>
      <link>https://cognaptus.com/blog/2026-04-08-from-spreadsheets-to-swarms-how-agentic-ai-rewrites-the-retail-supply-chain/</link>
      <pubDate>Wed, 08 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-08-from-spreadsheets-to-swarms-how-agentic-ai-rewrites-the-retail-supply-chain/</guid>
      <description>A mechanism-first reading of Flowr, an agentic AI framework that turns supermarket replenishment from manual coordination into supervised workflow automation.</description>
    </item>
    <item>
      <title>Walking the Graph: When LLMs Stop Guessing and Start Navigating</title>
      <link>https://cognaptus.com/blog/2026-04-05-walking-the-graph-when-llms-stop-guessing-and-start-navigating/</link>
      <pubDate>Sun, 05 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-05-walking-the-graph-when-llms-stop-guessing-and-start-navigating/</guid>
      <description>GraphWalk shows why enterprise knowledge-graph reasoning needs auditable navigation tools, not just larger prompts or cleaner retrieval.</description>
    </item>
    <item>
      <title>The Model That Forgot Itself: Why LLMs Drift Without Knowing</title>
      <link>https://cognaptus.com/blog/2026-03-29-the-model-that-forgot-itself-why-llms-drift-without-knowing/</link>
      <pubDate>Sun, 29 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-29-the-model-that-forgot-itself-why-llms-drift-without-knowing/</guid>
      <description>A mechanism-first reading of why LLMs can appear consistent while silently changing their hidden goals across a conversation.</description>
    </item>
    <item>
      <title>Belief Is a Graph: Why LLM Agents Need Structured Minds</title>
      <link>https://cognaptus.com/blog/2026-03-23-belief-is-a-graph-why-llm-agents-need-structured-minds/</link>
      <pubDate>Mon, 23 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-23-belief-is-a-graph-why-llm-agents-need-structured-minds/</guid>
      <description>A mechanism-first reading of dynamic belief graphs, and why enterprise LLM agents need structured, auditable mental states rather than longer prompts.</description>
    </item>
    <item>
      <title>The Illusion of Anonymity: When AI Connects the Dots You Thought Were Safe</title>
      <link>https://cognaptus.com/blog/2026-03-21-the-illusion-of-anonymity-when-ai-connects-the-dots-you-thought-were-safe/</link>
      <pubDate>Sat, 21 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-21-the-illusion-of-anonymity-when-ai-connects-the-dots-you-thought-were-safe/</guid>
      <description>A mechanism-first reading of how LLM agents turn weak, anonymized cues into real identity hypotheses—and why enterprise privacy governance must move beyond PII masking.</description>
    </item>
    <item>
      <title>The Hidden Playbook of LLMs: How AI Quietly Thinks Like a Hacker</title>
      <link>https://cognaptus.com/blog/2026-03-20-the-hidden-playbook-of-llms-how-ai-quietly-thinks-like-a-hacker/</link>
      <pubDate>Fri, 20 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-20-the-hidden-playbook-of-llms-how-ai-quietly-thinks-like-a-hacker/</guid>
      <description>A mechanism-first reading of how LLM agents implicitly control long-horizon binary vulnerability analysis through pruning, lock-in, backtracking, and prioritization.</description>
    </item>
    <item>
      <title>When Alignment Meets Reality: Why LLMs Can’t Agree With Themselves</title>
      <link>https://cognaptus.com/blog/2026-03-17-when-alignment-meets-reality-why-llms-cant-agree-with-themselves/</link>
      <pubDate>Tue, 17 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-17-when-alignment-meets-reality-why-llms-cant-agree-with-themselves/</guid>
      <description>A mechanism-first reading of why LLM alignment conflicts emerge, how priority hacking exploits them, and what enterprise AI systems should do at runtime.</description>
    </item>
    <item>
      <title>Ants in the Machine: What Swarm Intelligence Teaches Us About Routing LLM Agents</title>
      <link>https://cognaptus.com/blog/2026-03-16-ants-in-the-machine-what-swarm-intelligence-teaches-us-about-routing-llm-agents/</link>
      <pubDate>Mon, 16 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-16-ants-in-the-machine-what-swarm-intelligence-teaches-us-about-routing-llm-agents/</guid>
      <description>A mechanism-first reading of AMRO-S, a semantic and ant-colony-inspired routing framework for making multi-agent LLM systems cheaper, faster, and easier to inspect.</description>
    </item>
    <item>
      <title>From Hallucination to Verification: Why AI Needs a Pharmacist’s Mindset</title>
      <link>https://cognaptus.com/blog/2026-03-13-from-hallucination-to-verification-why-ai-needs-a-pharmacists-mindset/</link>
      <pubDate>Fri, 13 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-13-from-hallucination-to-verification-why-ai-needs-a-pharmacists-mindset/</guid>
      <description>A prescription-auditing paper shows why safe AI needs hybrid knowledge stores, deterministic checks, and evidence-grounded reasoning—not just bigger models.</description>
    </item>
    <item>
      <title>Prompt Politics: How Tiny Policies Can Steer Entire AI Societies</title>
      <link>https://cognaptus.com/blog/2026-03-11-prompt-politics-how-tiny-policies-can-steer-entire-ai-societies/</link>
      <pubDate>Wed, 11 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-11-prompt-politics-how-tiny-policies-can-steer-entire-ai-societies/</guid>
      <description>A mechanism-first reading of how policy-parameterized prompts can steer LLM multi-agent dialogue without model training—and what that means for business agent systems.</description>
    </item>
    <item>
      <title>Silver Bots: When Agentic AI Becomes the Caregiver</title>
      <link>https://cognaptus.com/blog/2026-03-07-silver-bots-when-agentic-ai-becomes-the-caregiver/</link>
      <pubDate>Sat, 07 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-07-silver-bots-when-agentic-ai-becomes-the-caregiver/</guid>
      <description>A clearer look at agentic AI in elderly care: what autonomous care agents can plausibly do, what the evidence actually shows, and where governance becomes operational.</description>
    </item>
    <item>
      <title>When Plans Talk Back: Conversational AI Meets Classical Planning</title>
      <link>https://cognaptus.com/blog/2026-03-03-when-plans-talk-back-conversational-ai-meets-classical-planning/</link>
      <pubDate>Tue, 03 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-03-when-plans-talk-back-conversational-ai-meets-classical-planning/</guid>
      <description>A mechanism-first reading of how LLM agents can make formal planning systems easier to question, revise, and trust without pretending to replace the planner.</description>
    </item>
    <item>
      <title>When Agents Ask for Help: Teaching LLMs the Art of Expert Collaboration</title>
      <link>https://cognaptus.com/blog/2026-02-28-when-agents-ask-for-help-teaching-llms-the-art-of-expert-collaboration/</link>
      <pubDate>Sat, 28 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-28-when-agents-ask-for-help-teaching-llms-the-art-of-expert-collaboration/</guid>
      <description>A mechanism-first reading of AHCE, a framework that teaches LLM agents when to escalate to human experts and how to turn messy advice into executable action.</description>
    </item>
    <item>
      <title>Gamma Rays and Toolboxes: Why Superintelligence May Be a Systems Engineering Problem</title>
      <link>https://cognaptus.com/blog/2026-02-25-gamma-rays-and-toolboxes-why-superintelligence-may-be-a-systems-engineering-problem/</link>
      <pubDate>Wed, 25 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-25-gamma-rays-and-toolboxes-why-superintelligence-may-be-a-systems-engineering-problem/</guid>
      <description>A new benchmark suggests that long-horizon AI reasoning may depend less on raw model scale than on whether models can reliably combine state, evidence, validation, and tools.</description>
    </item>
    <item>
      <title>Agents in Lab Coats: When LLMs Try to Become Data Scientists</title>
      <link>https://cognaptus.com/blog/2026-02-22-agents-in-lab-coats-when-llms-try-to-become-data-scientists/</link>
      <pubDate>Sun, 22 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-22-agents-in-lab-coats-when-llms-try-to-become-data-scientists/</guid>
      <description>A comparison-based guide to when single-agent, two-agent, multi-agent, and dynamic LLM data-science systems actually make business sense.</description>
    </item>
    <item>
      <title>Don’t Prompt Harder — Engineer Smarter: Inside CEDAR’s Agentic Data Scientist</title>
      <link>https://cognaptus.com/blog/2026-02-22-dont-prompt-harder-engineer-smarter-inside-cedars-agentic-data-scientist/</link>
      <pubDate>Sun, 22 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-22-dont-prompt-harder-engineer-smarter-inside-cedars-agentic-data-scientist/</guid>
      <description>CEDAR shows why useful AI data science systems depend less on magical prompting and more on structured context, local execution, agent routing, and inspectable workflows.</description>
    </item>
    <item>
      <title>From SQL Copilot to Autonomous Data Scientist: The L0–L5 Reality Check</title>
      <link>https://cognaptus.com/blog/2026-02-22-from-sql-copilot-to-autonomous-data-scientist-the-l0l5-reality-check/</link>
      <pubDate>Sun, 22 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-22-from-sql-copilot-to-autonomous-data-scientist-the-l0l5-reality-check/</guid>
      <description>A practical autonomy map for separating ordinary data copilots from supervised workflow agents, proactive data operators, and still-speculative autonomous data scientists.</description>
    </item>
    <item>
      <title>Death by a Thousand Prompts: Why Long-Horizon Attacks Break AI Agents</title>
      <link>https://cognaptus.com/blog/2026-02-21-death-by-a-thousand-prompts-why-longhorizon-attacks-break-ai-agents/</link>
      <pubDate>Sat, 21 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-21-death-by-a-thousand-prompts-why-longhorizon-attacks-break-ai-agents/</guid>
      <description>AgentLAB shows why enterprise AI security must move from single-prompt filtering to trajectory-level control over tools, memory, and multi-step behavior.</description>
    </item>
    <item>
      <title>From PDE to Pipeline: When LLMs Become Numerical Architects</title>
      <link>https://cognaptus.com/blog/2026-02-20-from-pde-to-pipeline-when-llms-become-numerical-architects/</link>
      <pubDate>Fri, 20 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-20-from-pde-to-pipeline-when-llms-become-numerical-architects/</guid>
      <description>A mechanism-first reading of AutoNumerics, showing why automated PDE solving is less about code generation and more about controlled solver planning, debugging, and verification.</description>
    </item>
    <item>
      <title>Consistency Is Not a Coincidence: When LLM Agents Disagree With Themselves</title>
      <link>https://cognaptus.com/blog/2026-02-14-consistency-is-not-a-coincidence-when-llm-agents-disagree-with-themselves/</link>
      <pubDate>Sat, 14 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-14-consistency-is-not-a-coincidence-when-llm-agents-disagree-with-themselves/</guid>
      <description>A paper on behavioral consistency shows why repeated agent trajectories can become an early warning signal for enterprise AI reliability.</description>
    </item>
    <item>
      <title>Mind the Gap: When Clinical LLMs Learn from Their Own Mistakes</title>
      <link>https://cognaptus.com/blog/2026-02-11-mind-the-gap-when-clinical-llms-learn-from-their-own-mistakes/</link>
      <pubDate>Wed, 11 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-11-mind-the-gap-when-clinical-llms-learn-from-their-own-mistakes/</guid>
      <description>A close reading of Differential Reasoning Learning, a clinical-agent framework that turns reasoning failures into reusable, auditable correction patches.</description>
    </item>
    <item>
      <title>From Features to Actions: Why Agentic AI Needs a New Explainability Playbook</title>
      <link>https://cognaptus.com/blog/2026-02-09-from-features-to-actions-why-agentic-ai-needs-a-new-explainability-playbook/</link>
      <pubDate>Mon, 09 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-09-from-features-to-actions-why-agentic-ai-needs-a-new-explainability-playbook/</guid>
      <description>A practical reading of why feature attribution explains static predictions, but trajectory-level diagnostics are needed to understand failures in agentic AI systems.</description>
    </item>
    <item>
      <title>When Aligned Models Compete: Nash Equilibria as the New Alignment Layer</title>
      <link>https://cognaptus.com/blog/2026-02-09-when-aligned-models-compete-nash-equilibria-as-the-new-alignment-layer/</link>
      <pubDate>Mon, 09 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-09-when-aligned-models-compete-nash-equilibria-as-the-new-alignment-layer/</guid>
      <description>A mechanism-first reading of LLM active alignment: why individually aligned agents can still produce exclusionary system equilibria when they compete for attention.</description>
    </item>
    <item>
      <title>Learning to Inject: When Prompt Injection Becomes an Optimization Problem</title>
      <link>https://cognaptus.com/blog/2026-02-08-learning-to-inject-when-prompt-injection-becomes-an-optimization-problem/</link>
      <pubDate>Sun, 08 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-08-learning-to-inject-when-prompt-injection-becomes-an-optimization-problem/</guid>
      <description>AutoInject shows why prompt injection should be tested as an adaptive optimization problem, not merely as a list of hand-written attack templates.</description>
    </item>
    <item>
      <title>Stop the All-Hands Meeting: When AI Agents Learn Who Actually Needs to Talk</title>
      <link>https://cognaptus.com/blog/2026-02-06-stop-the-allhands-meeting-when-ai-agents-learn-who-actually-needs-to-talk/</link>
      <pubDate>Fri, 06 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-06-stop-the-allhands-meeting-when-ai-agents-learn-who-actually-needs-to-talk/</guid>
      <description>DyTopo shows why multi-agent AI systems should route information by need, not by habit.</description>
    </item>
    <item>
      <title>More Isn’t Smarter: Why Agent Diversity Beats Agent Count</title>
      <link>https://cognaptus.com/blog/2026-02-04-more-isnt-smarter-why-agent-diversity-beats-agent-count/</link>
      <pubDate>Wed, 04 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-04-more-isnt-smarter-why-agent-diversity-beats-agent-count/</guid>
      <description>A mechanism-first reading of why multi-agent LLM systems saturate when agents repeat each other, and why useful diversity beats raw agent count.</description>
    </item>
    <item>
      <title>When Agents Stop Talking to the Wrong People</title>
      <link>https://cognaptus.com/blog/2026-02-04-when-agents-stop-talking-to-the-wrong-people/</link>
      <pubDate>Wed, 04 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-04-when-agents-stop-talking-to-the-wrong-people/</guid>
      <description>TodyComm shows why multi-agent AI systems need learned communication governance, not just more agents talking more often.</description>
    </item>
    <item>
      <title>Coaching the Swarm: Why Multi‑Agent RL Finally Scales</title>
      <link>https://cognaptus.com/blog/2026-02-03-coaching-the-swarm-why-multiagent-rl-finally-scales/</link>
      <pubDate>Tue, 03 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-03-coaching-the-swarm-why-multiagent-rl-finally-scales/</guid>
      <description>A mechanism-first reading of MAPPA, a process-reward method for turning multiagent LLM workflows from prompted collaboration into trainable systems.</description>
    </item>
    <item>
      <title>Agentic Systems Need Architecture, Not Vibes</title>
      <link>https://cognaptus.com/blog/2026-02-02-agentic-systems-need-architecture-not-vibes/</link>
      <pubDate>Mon, 02 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-02-agentic-systems-need-architecture-not-vibes/</guid>
      <description>A mechanism-first reading of why reliable AI agents need subsystem architecture, reusable design patterns, and clearer diagnosis than another enthusiastic list of agent tricks.</description>
    </item>
    <item>
      <title>When LLMs Get a Laptop: Why Sandboxes Might Be the Real AGI Benchmark</title>
      <link>https://cognaptus.com/blog/2026-01-24-when-llms-get-a-laptop-why-sandboxes-might-be-the-real-agi-benchmark/</link>
      <pubDate>Sat, 24 Jan 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-01-24-when-llms-get-a-laptop-why-sandboxes-might-be-the-real-agi-benchmark/</guid>
      <description>A mechanism-first reading of LLM-in-Sandbox, showing why giving models a minimal computer environment may matter more than adding another clever prompt.</description>
    </item>
    <item>
      <title>Affective Inertia: Teaching LLM Agents to Remember Who They Are</title>
      <link>https://cognaptus.com/blog/2026-01-23-affective-inertia-teaching-llm-agents-to-remember-who-they-are/</link>
      <pubDate>Fri, 23 Jan 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-01-23-affective-inertia-teaching-llm-agents-to-remember-who-they-are/</guid>
      <description>A mechanism-first reading of how explicit state dynamics can make LLM agents more temporally coherent, and why too much stability becomes its own failure mode.</description>
    </item>
    <item>
      <title>Rebuttal Agents, Not Rebuttal Text: Why ‘Verify‑Then‑Write’ Is the Only Scalable Future</title>
      <link>https://cognaptus.com/blog/2026-01-21-rebuttal-agents-not-rebuttal-text-why-verifythenwrite-is-the-only-scalable-future/</link>
      <pubDate>Wed, 21 Jan 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-01-21-rebuttal-agents-not-rebuttal-text-why-verifythenwrite-is-the-only-scalable-future/</guid>
      <description>How RebuttalAgent turns author responses from fluent text generation into auditable concern tracking, evidence construction, and strategic planning.</description>
    </item>
    <item>
      <title>Recommendations With Receipts: When LLMs Have to Prove They Behaved</title>
      <link>https://cognaptus.com/blog/2026-01-17-recommendations-with-receipts-when-llms-have-to-prove-they-behaved/</link>
      <pubDate>Sat, 17 Jan 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-01-17-recommendations-with-receipts-when-llms-have-to-prove-they-behaved/</guid>
      <description>A mechanism-first look at PCN-Rec, a proof-carrying architecture that turns LLM recommenders from trusted decision-makers into auditable proposers.</description>
    </item>
    <item>
      <title>Bubble Trouble: Why Top‑K Retrieval Keeps Letting LLMs Down</title>
      <link>https://cognaptus.com/blog/2026-01-16-bubble-trouble-why-topk-retrieval-keeps-letting-llms-down/</link>
      <pubDate>Fri, 16 Jan 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-01-16-bubble-trouble-why-topk-retrieval-keeps-letting-llms-down/</guid>
      <description>A practical reading of Context Bubble construction: why enterprise RAG needs constrained, auditable context assembly rather than larger top-k piles.</description>
    </item>
    <item>
      <title>Knowing Is Not Doing: When LLM Agents Pass the Task but Fail the World</title>
      <link>https://cognaptus.com/blog/2026-01-15-knowing-is-not-doing-when-llm-agents-pass-the-task-but-fail-the-world/</link>
      <pubDate>Thu, 15 Jan 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-01-15-knowing-is-not-doing-when-llm-agents-pass-the-task-but-fail-the-world/</guid>
      <description>Task2Quiz shows why agent evaluation needs to separate task completion from grounded environment understanding.</description>
    </item>
    <item>
      <title>Scaling the Sandbox: When LLM Agents Need Better Worlds</title>
      <link>https://cognaptus.com/blog/2026-01-14-scaling-the-sandbox-when-llm-agents-need-better-worlds/</link>
      <pubDate>Wed, 14 Jan 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-01-14-scaling-the-sandbox-when-llm-agents-need-better-worlds/</guid>
      <description>EnvScaler shows why useful LLM agents may need scalable executable worlds—not just more prompts, more tools, or larger models.</description>
    </item>
    <item>
      <title>STACKPLANNER: When Agents Learn to Forget</title>
      <link>https://cognaptus.com/blog/2026-01-12-stackplanner-when-agents-learn-to-forget/</link>
      <pubDate>Mon, 12 Jan 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-01-12-stackplanner-when-agents-learn-to-forget/</guid>
      <description>A mechanism-first reading of STACKPLANNER, showing why long-horizon agent systems may need memory control more than bigger context windows.</description>
    </item>
    <item>
      <title>Agents That Ship, Not Just Think: When LLM Self-Improvement Meets Release Engineering</title>
      <link>https://cognaptus.com/blog/2026-01-11-agents-that-ship-not-just-think-when-llm-selfimprovement-meets-release-engineering/</link>
      <pubDate>Sun, 11 Jan 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-01-11-agents-that-ship-not-just-think-when-llm-selfimprovement-meets-release-engineering/</guid>
      <description>AgentDevel shows why improving LLM agents may require release gates, traces, and regression control more than another round of self-reflection.</description>
    </item>
    <item>
      <title>ResMAS: When Multi‑Agent Systems Stop Falling Apart</title>
      <link>https://cognaptus.com/blog/2026-01-11-resmas-when-multiagent-systems-stop-falling-apart/</link>
      <pubDate>Sun, 11 Jan 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-01-11-resmas-when-multiagent-systems-stop-falling-apart/</guid>
      <description>A mechanism-first reading of ResMAS, showing why resilient LLM agent systems depend on communication topology and topology-aware prompts, not just more agents.</description>
    </item>
    <item>
      <title>From Tokens to Topology: Teaching LLMs to Think in Simulink</title>
      <link>https://cognaptus.com/blog/2026-01-09-from-tokens-to-topology-teaching-llms-to-think-in-simulink/</link>
      <pubDate>Fri, 09 Jan 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-01-09-from-tokens-to-topology-teaching-llms-to-think-in-simulink/</guid>
      <description>A mechanism-first reading of SimuAgent, a Simulink modeling assistant that shows why representation, validation, curriculum, and reflection matter more than merely attaching a larger model to an engineering tool.</description>
    </item>
    <item>
      <title>Trading Without Cheating: Teaching LLMs to Reason When Markets Lie</title>
      <link>https://cognaptus.com/blog/2026-01-08-trading-without-cheating-teaching-llms-to-reason-when-markets-lie/</link>
      <pubDate>Thu, 08 Jan 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-01-08-trading-without-cheating-teaching-llms-to-reason-when-markets-lie/</guid>
      <description>A mechanism-first reading of Trade-R1, a framework for training financial LLM agents when market returns are objective but dangerously noisy.</description>
    </item>
    <item>
      <title>Pulling the Thread: Why LLM Reasoning Often Unravels</title>
      <link>https://cognaptus.com/blog/2026-01-06-pulling-the-thread-why-llm-reasoning-often-unravels/</link>
      <pubDate>Tue, 06 Jan 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-01-06-pulling-the-thread-why-llm-reasoning-often-unravels/</guid>
      <description>Project Ariadne shows how counterfactual interventions can audit whether an LLM’s reasoning trace actually causes its answer, or merely decorates it.</description>
    </item>
    <item>
      <title>Talking to Yourself, but Make It Useful: Intrinsic Self‑Critique in LLM Planning</title>
      <link>https://cognaptus.com/blog/2026-01-03-talking-to-yourself-but-make-it-useful-intrinsic-selfcritique-in-llm-planning/</link>
      <pubDate>Sat, 03 Jan 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-01-03-talking-to-yourself-but-make-it-useful-intrinsic-selfcritique-in-llm-planning/</guid>
      <description>A procedural self-critique loop can make LLM planners markedly more reliable—but only when reflection is converted into explicit rule checking, state tracking, and conservative approval.</description>
    </item>
    <item>
      <title>Silent Scholars, No More: When Uncertainty Becomes an Agent’s Survival Instinct</title>
      <link>https://cognaptus.com/blog/2025-12-28-silent-scholars-no-more-when-uncertainty-becomes-an-agents-survival-instinct/</link>
      <pubDate>Sun, 28 Dec 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-12-28-silent-scholars-no-more-when-uncertainty-becomes-an-agents-survival-instinct/</guid>
      <description>A mechanism-first reading of why future LLM agents may need uncertainty-driven feedback loops, not just larger memories or better retrieval.</description>
    </item>
    <item>
      <title>When Reflection Needs a Committee: Why LLMs Think Better in Groups</title>
      <link>https://cognaptus.com/blog/2025-12-28-when-reflection-needs-a-committee-why-llms-think-better-in-groups/</link>
      <pubDate>Sun, 28 Dec 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-12-28-when-reflection-needs-a-committee-why-llms-think-better-in-groups/</guid>
      <description>A mechanism-first reading of Multi-Agent Reflexion and what it teaches businesses about separating execution, critique, judgment, and memory in LLM agents.</description>
    </item>
    <item>
      <title>When Agents Agree Too Much: Emergent Bias in Multi‑Agent AI Systems</title>
      <link>https://cognaptus.com/blog/2025-12-21-when-agents-agree-too-much-emergent-bias-in-multiagent-ai-systems/</link>
      <pubDate>Sun, 21 Dec 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-12-21-when-agents-agree-too-much-emergent-bias-in-multiagent-ai-systems/</guid>
      <description>A financial AI fairness study shows why testing individual LLM agents is not enough when their collaboration can create new system-level bias.</description>
    </item>
    <item>
      <title>Don’t Tell the Robot What You Know</title>
      <link>https://cognaptus.com/blog/2025-12-20-dont-tell-the-robot-what-you-know/</link>
      <pubDate>Sat, 20 Dec 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-12-20-dont-tell-the-robot-what-you-know/</guid>
      <description>A new embodied-agent study shows why collaborative AI fails when the informed agent gives more instructions instead of helping the limited agent verify what it can actually perceive.</description>
    </item>
    <item>
      <title>Model First, Think Later: Why LLMs Fail Before They Reason</title>
      <link>https://cognaptus.com/blog/2025-12-17-model-first-think-later-why-llms-fail-before-they-reason/</link>
      <pubDate>Wed, 17 Dec 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-12-17-model-first-think-later-why-llms-fail-before-they-reason/</guid>
      <description>A practical reading of Model-First Reasoning: why agent failures often begin with unstable problem representation, not weak reasoning.</description>
    </item>
    <item>
      <title>When Rewards Learn Back: Evolution, but With Gradients</title>
      <link>https://cognaptus.com/blog/2025-12-16-when-rewards-learn-back-evolution-but-with-gradients/</link>
      <pubDate>Tue, 16 Dec 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-12-16-when-rewards-learn-back-evolution-but-with-gradients/</guid>
      <description>A mechanism-first reading of DERL: how reward design becomes a learnable outer-loop problem, and why that matters for enterprise agents.</description>
    </item>
    <item>
      <title>When Agents Loop: Geometry, Drift, and the Hidden Physics of LLM Behavior</title>
      <link>https://cognaptus.com/blog/2025-12-14-when-agents-loop-geometry-drift-and-the-hidden-physics-of-llm-behavior/</link>
      <pubDate>Sun, 14 Dec 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-12-14-when-agents-loop-geometry-drift-and-the-hidden-physics-of-llm-behavior/</guid>
      <description>A practical reading of how recursive LLM agents converge, drift, or wander depending less on the model than on the loop we force it to run.</description>
    </item>
    <item>
      <title>When Tokens Become Actions: A Policy Gradient Built for Transformers</title>
      <link>https://cognaptus.com/blog/2025-12-14-when-tokens-become-actions-a-policy-gradient-built-for-transformers/</link>
      <pubDate>Sun, 14 Dec 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-12-14-when-tokens-become-actions-a-policy-gradient-built-for-transformers/</guid>
      <description>A mechanism-first reading of GPG, a Transformer-aware policy-gradient framework that turns output segments into trainable macro-actions for LLM agents.</description>
    </item>
    <item>
      <title>Teach Me Once: How One‑Shot LLM Guidance Reshapes Hierarchical Planning</title>
      <link>https://cognaptus.com/blog/2025-12-11-teach-me-once-how-oneshot-llm-guidance-reshapes-hierarchical-planning/</link>
      <pubDate>Thu, 11 Dec 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-12-11-teach-me-once-how-oneshot-llm-guidance-reshapes-hierarchical-planning/</guid>
      <description>A mechanism-first reading of SCOPE, a paper showing how LLM guidance can be moved from runtime planning into one-time subgoal initialization for cheaper hierarchical agents.</description>
    </item>
    <item>
      <title>Error 404: Peer Review Not Found — How LLMs Are Quietly Rewriting Scientific Quality Control</title>
      <link>https://cognaptus.com/blog/2025-12-08-error-404-peer-review-not-found-how-llms-are-quietly-rewriting-scientific-quality-control/</link>
      <pubDate>Mon, 08 Dec 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-12-08-error-404-peer-review-not-found-how-llms-are-quietly-rewriting-scientific-quality-control/</guid>
      <description>A close reading of how a GPT-5-based correctness checker turns scientific paper auditing from artisanal peer-review labor into a scalable quality-control workflow.</description>
    </item>
    <item>
      <title>Stacking the Odds: Why Blocksworld Still Breaks Your Fancy LLM Agent</title>
      <link>https://cognaptus.com/blog/2025-12-04-stacking-the-odds-why-blocksworld-still-breaks-your-fancy-llm-agent/</link>
      <pubDate>Thu, 04 Dec 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-12-04-stacking-the-odds-why-blocksworld-still-breaks-your-fancy-llm-agent/</guid>
      <description>A practical reading of an MCP-integrated Blocksworld benchmark showing why planning, verification, execution, and replanning must be tested together before LLM agents touch real operations.</description>
    </item>
    <item>
      <title>Short Paths, Sharp Minds: Why Knowledge Graph Distance Feels Like Cognitive Gravity</title>
      <link>https://cognaptus.com/blog/2025-12-02-short-paths-sharp-minds-why-knowledge-graph-distance-feels-like-cognitive-gravity/</link>
      <pubDate>Tue, 02 Dec 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-12-02-short-paths-sharp-minds-why-knowledge-graph-distance-feels-like-cognitive-gravity/</guid>
      <description>A mechanism-first reading of how graph distance can act as a surprise signal for knowledge-graph reasoning, and why the idea is useful before it is proven.</description>
    </item>
    <item>
      <title>Think Fast, Act Faster: How &#39;Thinking-by-Doing&#39; Is Rewiring LLM World Models</title>
      <link>https://cognaptus.com/blog/2025-12-01-think-fast-act-faster-how-thinkingbydoing-is-rewiring-llm-world-models/</link>
      <pubDate>Mon, 01 Dec 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-12-01-think-fast-act-faster-how-thinkingbydoing-is-rewiring-llm-world-models/</guid>
      <description>WMAct shows how multi-turn interaction can train LLM agents to compress feedback into reusable world-model reasoning, but only when exploration is disciplined.</description>
    </item>
    <item>
      <title>When Agents Treat Agents as Tools: What Tool-RoCo Tells Us About LLM Autonomy</title>
      <link>https://cognaptus.com/blog/2025-11-29-when-agents-treat-agents-as-tools-what-toolroco-tells-us-about-llm-autonomy/</link>
      <pubDate>Sat, 29 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-29-when-agents-treat-agents-as-tools-what-toolroco-tells-us-about-llm-autonomy/</guid>
      <description>Tool-RoCo shows why multi-agent LLM systems need evaluation of coordination lifecycle decisions, not just final task success or agent activation.</description>
    </item>
    <item>
      <title>Cutting Through the Noise: How Programmatic Pruning Turns Web Agents into Real Operators</title>
      <link>https://cognaptus.com/blog/2025-11-28-cutting-through-the-noise-how-programmatic-pruning-turns-web-agents-into-real-operators/</link>
      <pubDate>Fri, 28 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-28-cutting-through-the-noise-how-programmatic-pruning-turns-web-agents-into-real-operators/</guid>
      <description>A mechanism-first analysis of Prune4Web and why reliable web agents may depend less on larger context windows than on better ways to shrink the page before reasoning begins.</description>
    </item>
    <item>
      <title>Enviro-Mental Gymnastics: Why Cross-Environment Agents Still Trip Over Their Own Feet</title>
      <link>https://cognaptus.com/blog/2025-11-25-enviromental-gymnastics-why-crossenvironment-agents-still-trip-over-their-own-feet/</link>
      <pubDate>Tue, 25 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-25-enviromental-gymnastics-why-crossenvironment-agents-still-trip-over-their-own-feet/</guid>
      <description>AutoEnv shows why agent learning needs diverse environment testing, adaptive learning methods, and fewer victory laps from single-demo performance.</description>
    </item>
    <item>
      <title>Agents Behaving Badly: Why &#39;Agentic AI&#39; Needs Adult Supervision</title>
      <link>https://cognaptus.com/blog/2025-11-24-agents-behaving-badly-why-agentic-ai-needs-adult-supervision/</link>
      <pubDate>Mon, 24 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-24-agents-behaving-badly-why-agentic-ai-needs-adult-supervision/</guid>
      <description>A mechanism-first reading of why LLM-based agents need explicit architectures, communication rules, incentives, norms, trust models, and institutional supervision before they can become reliable business systems.</description>
    </item>
    <item>
      <title>Intent, Actually: Why DeFi Needs a Mind‑Reader</title>
      <link>https://cognaptus.com/blog/2025-11-21-intent-actually-why-defi-needs-a-mindreader/</link>
      <pubDate>Fri, 21 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-21-intent-actually-why-defi-needs-a-mindreader/</guid>
      <description>A mechanism-first reading of TIM, a multi-agent LLM framework that turns opaque DeFi transactions into evidence-ranked intent labels without pretending to read private motives.</description>
    </item>
    <item>
      <title>Skills to Pay the Agent Bills: Why LLMs Need Better Moves, Not Bigger Models</title>
      <link>https://cognaptus.com/blog/2025-11-20-skills-to-pay-the-agent-bills-why-llms-need-better-moves-not-bigger-models/</link>
      <pubDate>Thu, 20 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-20-skills-to-pay-the-agent-bills-why-llms-need-better-moves-not-bigger-models/</guid>
      <description>SkillGen shows why the next gain in LLM agents may come from reusable procedural skills, not longer prompts or larger models.</description>
    </item>
    <item>
      <title>Tools of Habit: Why LLM Agents Benefit from a Little Inertia</title>
      <link>https://cognaptus.com/blog/2025-11-20-tools-of-habit-why-llm-agents-benefit-from-a-little-inertia/</link>
      <pubDate>Thu, 20 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-20-tools-of-habit-why-llm-agents-benefit-from-a-little-inertia/</guid>
      <description>AutoTool shows how agent systems can cut repeated tool-selection costs by learning when workflow habits are reliable enough to bypass another LLM call.</description>
    </item>
    <item>
      <title>Memory, Bias, and the Mind of Machines: How Agentic LLMs Mislearn</title>
      <link>https://cognaptus.com/blog/2025-11-12-memory-bias-and-the-mind-of-machines-how-agentic-llms-mislearn/</link>
      <pubDate>Wed, 12 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-12-memory-bias-and-the-mind-of-machines-how-agentic-llms-mislearn/</guid>
      <description>A practical reading of why agent memory can turn useful experience into brittle, biased behaviour when systems continuously rewrite their own lessons.</description>
    </item>
    <item>
      <title>Parallel Worlds of Moderation: How LLM Simulations Are Stress-Testing Online Civility</title>
      <link>https://cognaptus.com/blog/2025-11-12-parallel-worlds-of-moderation-how-llm-simulations-are-stresstesting-online-civility/</link>
      <pubDate>Wed, 12 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-12-parallel-worlds-of-moderation-how-llm-simulations-are-stresstesting-online-civility/</guid>
      <description>COSMOS turns online moderation into a counterfactual simulation problem, showing why personalised interventions may reduce toxicity without the collateral damage of blunt bans.</description>
    </item>
    <item>
      <title>Parallel Worlds of Moderation: Simulating Online Civility with LLMs</title>
      <link>https://cognaptus.com/blog/2025-11-11-parallel-worlds-of-moderation-simulating-online-civility-with-llms/</link>
      <pubDate>Tue, 11 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-11-parallel-worlds-of-moderation-simulating-online-civility-with-llms/</guid>
      <description>A mechanism-first reading of COSMOS, an LLM-powered counterfactual simulator for testing moderation strategies before exposing real communities to policy experiments.</description>
    </item>
    <item>
      <title>Dirty Data, Clean Machines: How LLM Agents Rewire Predictive Maintenance</title>
      <link>https://cognaptus.com/blog/2025-11-10-dirty-data-clean-machines-how-llm-agents-rewire-predictive-maintenance/</link>
      <pubDate>Mon, 10 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-10-dirty-data-clean-machines-how-llm-agents-rewire-predictive-maintenance/</guid>
      <description>A new benchmark shows where LLM agents can clean predictive-maintenance logs today, and where industrial deployment still needs rules, temporal logic, and human discipline.</description>
    </item>
    <item>
      <title>Thinking Fast and Flowing Slow: Real-Time Reasoning for Autonomous Agents</title>
      <link>https://cognaptus.com/blog/2025-11-10-thinking-fast-and-flowing-slow-realtime-reasoning-for-autonomous-agents/</link>
      <pubDate>Mon, 10 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-10-thinking-fast-and-flowing-slow-realtime-reasoning-for-autonomous-agents/</guid>
      <description>A sharp look at why real-time AI agents need latency-aware architectures, not merely bigger models or longer reasoning traces.</description>
    </item>
    <item>
      <title>The Rational Illusion: How LLMs Outplayed Humans at Cooperation</title>
      <link>https://cognaptus.com/blog/2025-11-07-the-rational-illusion-how-llms-outplayed-humans-at-cooperation/</link>
      <pubDate>Fri, 07 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-07-the-rational-illusion-how-llms-outplayed-humans-at-cooperation/</guid>
      <description>A comparison-led reading of why Llama, Qwen, and Mistral behave differently in game-theory simulations—and what that means for using LLMs as behavioural testbeds.</description>
    </item>
    <item>
      <title>Recursive Minds: How ReCAP Turns LLMs into Self-Correcting Planners</title>
      <link>https://cognaptus.com/blog/2025-11-02-recursive-minds-how-recap-turns-llms-into-selfcorrecting-planners/</link>
      <pubDate>Sun, 02 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-02-recursive-minds-how-recap-turns-llms-into-selfcorrecting-planners/</guid>
      <description>ReCAP shows that long-horizon AI agents need not just more context, but better context organisation: recursive planning, parent-plan reinjection, and bounded memory.</description>
    </item>
    <item>
      <title>When Agents Learn to Test Themselves: TDFlow and the Future of Software Engineering</title>
      <link>https://cognaptus.com/blog/2025-11-02-when-agents-learn-to-test-themselves-tdflow-and-the-future-of-software-engineering/</link>
      <pubDate>Sun, 02 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-02-when-agents-learn-to-test-themselves-tdflow-and-the-future-of-software-engineering/</guid>
      <description>TDFlow shows that coding agents become far more useful when humans define correctness as executable tests and agents are constrained to solve them.</description>
    </item>
    <item>
      <title>Beyond Utility: When LLM Agents Start Dreaming Their Own Tasks</title>
      <link>https://cognaptus.com/blog/2025-10-23-beyond-utility-when-llm-agents-start-dreaming-their-own-tasks/</link>
      <pubDate>Thu, 23 Oct 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-10-23-beyond-utility-when-llm-agents-start-dreaming-their-own-tasks/</guid>
      <description>A mechanism-first look at how open-ended LLM agents move from task execution to task generation—and why that is not the same as autonomy.</description>
    </item>
    <item>
      <title>Blueprints of Agency: Compositional Machines and the New Architecture of Intelligence</title>
      <link>https://cognaptus.com/blog/2025-10-23-blueprints-of-agency-compositional-machines-and-the-new-architecture-of-intelligence/</link>
      <pubDate>Thu, 23 Oct 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-10-23-blueprints-of-agency-compositional-machines-and-the-new-architecture-of-intelligence/</guid>
      <description>A mechanism-first reading of how LLM agents assemble, test, refine, and partially learn machine designs inside a physics simulator.</description>
    </item>
    <item>
      <title>Pods over Prompts: Shachi’s Playbook for Serious Agent-Based Simulation</title>
      <link>https://cognaptus.com/blog/2025-10-03-pods-over-prompts-shachis-playbook-for-serious-agentbased-simulation/</link>
      <pubDate>Fri, 03 Oct 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-10-03-pods-over-prompts-shachis-playbook-for-serious-agentbased-simulation/</guid>
      <description>Shachi shows why serious LLM-based agent simulation needs modular cognitive architecture, not just better personas and larger crowds.</description>
    </item>
    <item>
      <title>Failures, Taxonomized: How Multi‑Level Reflection Turns Agents Into Self‑Learners</title>
      <link>https://cognaptus.com/blog/2025-10-02-failures-taxonomized-how-multilevel-reflection-turns-agents-into-selflearners/</link>
      <pubDate>Thu, 02 Oct 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-10-02-failures-taxonomized-how-multilevel-reflection-turns-agents-into-selflearners/</guid>
      <description>How SaMuLe turns failed agent traces into a reusable diagnostic layer—and what that means for enterprise automation.</description>
    </item>
    <item>
      <title>Memory That Fights Back: How SEDM Turns Agent Logs into Verified Knowledge</title>
      <link>https://cognaptus.com/blog/2025-09-17-memory-that-fights-back-how-sedm-turns-agent-logs-into-verified-knowledge/</link>
      <pubDate>Wed, 17 Sep 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-09-17-memory-that-fights-back-how-sedm-turns-agent-logs-into-verified-knowledge/</guid>
      <description>SEDM reframes agent memory as an auditable lifecycle: verify before storing, schedule before retrieving, consolidate before scaling, and revalidate before transfer.</description>
    </item>
    <item>
      <title>Search Party in a Notebook: JUPITER Turns Data Analysis into a Tree Game</title>
      <link>https://cognaptus.com/blog/2025-09-17-search-party-in-a-notebook-jupiter-turns-data-analysis-into-a-tree-game/</link>
      <pubDate>Wed, 17 Sep 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-09-17-search-party-in-a-notebook-jupiter-turns-data-analysis-into-a-tree-game/</guid>
      <description>JUPITER shows how real notebook traces and value-guided search can make smaller open models more reliable at multi-step data analysis.</description>
    </item>
    <item>
      <title>Small Gains, Long Games: Why Tiny Accuracy Bumps Explode into Big Execution Wins</title>
      <link>https://cognaptus.com/blog/2025-09-17-small-gains-long-games-why-tiny-accuracy-bumps-explode-into-big-execution-wins/</link>
      <pubDate>Wed, 17 Sep 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-09-17-small-gains-long-games-why-tiny-accuracy-bumps-explode-into-big-execution-wins/</guid>
      <description>A controlled long-horizon execution study shows why small per-step reliability gains can create large business value—and why agents need execution architecture, not just clever prompts.</description>
    </item>
    <item>
      <title>Guardrails Before Gas: Secure Plan‑Then‑Execute Agents for Real Work</title>
      <link>https://cognaptus.com/blog/2025-09-14-guardrails-before-gas-secure-planthenexecute-agents-for-real-work/</link>
      <pubDate>Sun, 14 Sep 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-09-14-guardrails-before-gas-secure-planthenexecute-agents-for-real-work/</guid>
      <description>A mechanism-first guide to why Plan-then-Execute agents improve control-flow security, where they still fail, and how enterprises should harden them before production.</description>
    </item>
    <item>
      <title>Tool Time, Any Time: Inside RLFactory’s Plug‑and‑Play RL for Multi‑Turn Tool Use</title>
      <link>https://cognaptus.com/blog/2025-09-13-tool-time-any-time-inside-rlfactorys-plugandplay-rl-for-multiturn-tool-use/</link>
      <pubDate>Sat, 13 Sep 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-09-13-tool-time-any-time-inside-rlfactorys-plugandplay-rl-for-multiturn-tool-use/</guid>
      <description>RLFactory shows how agent RL can be rebuilt around tool feedback, async invocation, and modular rewards—useful plumbing, not magic autonomy.</description>
    </item>
    <item>
      <title>Plan, Act, Replan: When LLM Agents Run the Aisles</title>
      <link>https://cognaptus.com/blog/2025-09-08-plan-act-replan-when-llm-agents-run-the-aisles/</link>
      <pubDate>Mon, 08 Sep 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-09-08-plan-act-replan-when-llm-agents-run-the-aisles/</guid>
      <description>JD.com’s supply-chain agent framework shows that the business value of GenAI planning is not prettier forecasts, but a shorter loop between intent, execution, diagnosis, and correction.</description>
    </item>
    <item>
      <title>Plan, Don&#39;t Spam: The Goldilocks Rule for Test‑Time Compute</title>
      <link>https://cognaptus.com/blog/2025-09-08-plan-dont-spam-the-goldilocks-rule-for-testtime-compute/</link>
      <pubDate>Mon, 08 Sep 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-09-08-plan-dont-spam-the-goldilocks-rule-for-testtime-compute/</guid>
      <description>A new agent-planning paper shows why the best LLM agents should treat explicit reasoning as a scarce operational resource, not a reflex.</description>
    </item>
    <item>
      <title>Rules of Engagement: How Meta‑Policy Reflexion Turns Agent Memory into Guardrails</title>
      <link>https://cognaptus.com/blog/2025-09-08-rules-of-engagement-how-metapolicy-reflexion-turns-agent-memory-into-guardrails/</link>
      <pubDate>Mon, 08 Sep 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-09-08-rules-of-engagement-how-metapolicy-reflexion-turns-agent-memory-into-guardrails/</guid>
      <description>A mechanism-first reading of Meta-Policy Reflexion, a training-free approach that turns failed agent trajectories into reusable rules and test-time guardrails.</description>
    </item>
    <item>
      <title>From Prompts to Policies: The Agentic RL Playbook</title>
      <link>https://cognaptus.com/blog/2025-09-04-from-prompts-to-policies-the-agentic-rl-playbook/</link>
      <pubDate>Thu, 04 Sep 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-09-04-from-prompts-to-policies-the-agentic-rl-playbook/</guid>
      <description>A mechanism-first guide to why Agentic RL reframes LLMs as trainable policies operating across tools, memory, environments, and long-horizon business workflows.</description>
    </item>
    <item>
      <title>Patience Is Profit: Can LLM Agents Stabilize DePIN’s Token Rails?</title>
      <link>https://cognaptus.com/blog/2025-09-01-patience-is-profit-can-llm-agents-stabilize-depins-token-rails/</link>
      <pubDate>Mon, 01 Sep 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-09-01-patience-is-profit-can-llm-agents-stabilize-depins-token-rails/</guid>
      <description>A mechanism-first reading of EconAgentic, a DePIN market simulation showing how token incentives, node-provider patience, and LLM-style decision rules may affect network inclusion, stability, and market value.</description>
    </item>
    <item>
      <title>Mirror, Signal, Maneuver: How &#39;Self&#39; Labels Nudge LLM Cooperation</title>
      <link>https://cognaptus.com/blog/2025-08-27-mirror-signal-maneuver-how-self-labels-nudge-llm-cooperation/</link>
      <pubDate>Wed, 27 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-08-27-mirror-signal-maneuver-how-self-labels-nudge-llm-cooperation/</guid>
      <description>A behavioural-economics reading of how simple self-labels can shift LLM cooperation, and what that means for multi-agent AI operations.</description>
    </item>
    <item>
      <title>Agents on the Clock: Turning a 3‑Layer Taxonomy into a Build‑Ready Playbook</title>
      <link>https://cognaptus.com/blog/2025-08-26-agents-on-the-clock-turning-a-3layer-taxonomy-into-a-buildready-playbook/</link>
      <pubDate>Tue, 26 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-08-26-agents-on-the-clock-turning-a-3layer-taxonomy-into-a-buildready-playbook/</guid>
      <description>A mechanism-first reading of agentic reasoning frameworks as operational control loops, not just model upgrades.</description>
    </item>
    <item>
      <title>Preference Chains of Command: Making LLM Agents Pick Like People</title>
      <link>https://cognaptus.com/blog/2025-08-25-preference-chains-of-command-making-llm-agents-pick-like-people/</link>
      <pubDate>Mon, 25 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-08-25-preference-chains-of-command-making-llm-agents-pick-like-people/</guid>
      <description>A mechanism-first reading of Preference Chain, a Graph RAG method that uses small behavioural samples to make LLM mobility agents less generic and more locally plausible.</description>
    </item>
    <item>
      <title>Enemy at the Gates, Friends at the Table: Why Competition Makes LLM Agents More Cooperative</title>
      <link>https://cognaptus.com/blog/2025-08-24-enemy-at-the-gates-friends-at-the-table-why-competition-makes-llm-agents-more-cooperative/</link>
      <pubDate>Sun, 24 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-08-24-enemy-at-the-gates-friends-at-the-table-why-competition-makes-llm-agents-more-cooperative/</guid>
      <description>A mechanism-first reading of why external rivalry can make LLM agent teams cooperate more internally, and why that lesson is useful but not yet operational proof.</description>
    </item>
    <item>
      <title>Stackelbergs &amp; Stakeholders: Turning Bits into Boardroom Moves</title>
      <link>https://cognaptus.com/blog/2025-08-24-stackelbergs-stakeholders-turning-bits-into-boardroom-moves/</link>
      <pubDate>Sun, 24 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-08-24-stackelbergs-stakeholders-turning-bits-into-boardroom-moves/</guid>
      <description>BusiAgent shows how multi-agent LLMs can turn broad business requests into governed workflows, but its real value is orchestration discipline, not artificial executive genius.</description>
    </item>
    <item>
      <title>USB‑C for Agents, Stress‑Tested: What MCP‑Universe Really Reveals</title>
      <link>https://cognaptus.com/blog/2025-08-23-usbc-for-agents-stresstested-what-mcpuniverse-really-reveals/</link>
      <pubDate>Sat, 23 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-08-23-usbc-for-agents-stresstested-what-mcpuniverse-really-reveals/</guid>
      <description>MCP-Universe shows that connecting agents to real tools is easy; making them reliable across messy, live workflows is still the hard part.</description>
    </item>
    <item>
      <title>Prefix, Not Pretext: A One‑Line Fix for Agent Misalignment</title>
      <link>https://cognaptus.com/blog/2025-08-20-prefix-not-pretext-a-oneline-fix-for-agent-misalignment/</link>
      <pubDate>Wed, 20 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-08-20-prefix-not-pretext-a-oneline-fix-for-agent-misalignment/</guid>
      <description>A mechanism-first reading of why benign agent fine-tuning can erode refusals, and why PING works by steering the first response tokens rather than rewriting the model.</description>
    </item>
    <item>
      <title>Quants With a Plan: Agentic Workflows That Outtrade AutoML</title>
      <link>https://cognaptus.com/blog/2025-08-20-quants-with-a-plan-agentic-workflows-that-outtrade-automl/</link>
      <pubDate>Wed, 20 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-08-20-quants-with-a-plan-agentic-workflows-that-outtrade-automl/</guid>
      <description>TS-Agent shows that financial modelling agents improve when they are constrained by curated model banks, refinement knowledge, feedback loops, and auditable code-edit trails.</description>
    </item>
    <item>
      <title>Crystal Ball, Meet Cron Job: What FutureX Reveals About ‘Live’ Forecasting Agents</title>
      <link>https://cognaptus.com/blog/2025-08-19-crystal-ball-meet-cron-job-what-futurex-reveals-about-live-forecasting-agents/</link>
      <pubDate>Tue, 19 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-08-19-crystal-ball-meet-cron-job-what-futurex-reveals-about-live-forecasting-agents/</guid>
      <description>FutureX shows that forecasting agents should be judged as live operating systems, not static models with better vibes.</description>
    </item>
    <item>
      <title>Consent, Coaxing, and Countermoves: Simulating Privacy Attacks on LLM Agents</title>
      <link>https://cognaptus.com/blog/2025-08-18-consent-coaxing-and-countermoves-simulating-privacy-attacks-on-llm-agents/</link>
      <pubDate>Mon, 18 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-08-18-consent-coaxing-and-countermoves-simulating-privacy-attacks-on-llm-agents/</guid>
      <description>A search-based privacy red-teaming framework shows how agent-agent attacks evolve from blunt requests into forged consent and impersonation—and why static privacy prompts are not enough.</description>
    </item>
    <item>
      <title>Skip or Split? How LLMs Can Make Old-School Planners Run Circles Around Complexity</title>
      <link>https://cognaptus.com/blog/2025-08-18-skip-or-split-how-llms-can-make-oldschool-planners-run-circles-around-complexity/</link>
      <pubDate>Mon, 18 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-08-18-skip-or-split-how-llms-can-make-oldschool-planners-run-circles-around-complexity/</guid>
      <description>A comparison-based reading of why LLM-assisted planning works best when the model predicts constrained intermediate states rather than pretending to be the planner.</description>
    </item>
    <item>
      <title>Therapy, Explained: How Multi‑Agent LLMs Turn DSM‑5 Screens into Auditable Logic</title>
      <link>https://cognaptus.com/blog/2025-08-18-therapy-explained-how-multiagent-llms-turn-dsm5-screens-into-auditable-logic/</link>
      <pubDate>Mon, 18 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-08-18-therapy-explained-how-multiagent-llms-turn-dsm5-screens-into-auditable-logic/</guid>
      <description>DSM5AgentFlow shows that the business value of clinical LLM agents is not autonomous diagnosis, but auditable intake infrastructure.</description>
    </item>
    <item>
      <title>Confounder Hunters: How LLM Agents are Rewriting the Rules of Causal Inference</title>
      <link>https://cognaptus.com/blog/2025-08-12-confounder-hunters-how-llm-agents-are-rewriting-the-rules-of-causal-inference/</link>
      <pubDate>Tue, 12 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-08-12-confounder-hunters-how-llm-agents-are-rewriting-the-rules-of-causal-inference/</guid>
      <description>A mechanism-first look at how LLM agents can help causal ML systems discover confounders, refine unstable subgroups, and reduce expert review burden without pretending to automate causal truth.</description>
    </item>
    <item>
      <title>Meta-Game Theory: What a Pokémon League Taught Us About LLM Strategy</title>
      <link>https://cognaptus.com/blog/2025-08-09-metagame-theory-what-a-pokmon-league-taught-us-about-llm-strategy/</link>
      <pubDate>Sat, 09 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-08-09-metagame-theory-what-a-pokmon-league-taught-us-about-llm-strategy/</guid>
      <description>A small Pokémon tournament shows why LLM evaluation should measure strategy, rationale, and constraint exploitation—not just polished reasoning.</description>
    </item>
    <item>
      <title>Forecast First, Ask Later: How DCATS Makes Time Series Smarter with LLMs</title>
      <link>https://cognaptus.com/blog/2025-08-07-forecast-first-ask-later-how-dcats-makes-time-series-smarter-with-llms/</link>
      <pubDate>Thu, 07 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-08-07-forecast-first-ask-later-how-dcats-makes-time-series-smarter-with-llms/</guid>
      <description>DCATS shows how LLM agents can improve time series forecasting by curating auxiliary data, not by inventing a cleverer forecasting model.</description>
    </item>
    <item>
      <title>The Forest Within: How Galaxy Reinvents LLM Agents with Self-Evolving Cognition</title>
      <link>https://cognaptus.com/blog/2025-08-07-the-forest-within-how-galaxy-reinvents-llm-agents-with-selfevolving-cognition/</link>
      <pubDate>Thu, 07 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-08-07-the-forest-within-how-galaxy-reinvents-llm-agents-with-selfevolving-cognition/</guid>
      <description>Galaxy shows how proactive personal agents may need cognition, system design, privacy handling, and self-repair to evolve as one mechanism rather than four separate features.</description>
    </item>
    <item>
      <title>Forkcast: How Pro2Guard Predicts and Prevents LLM Agent Failures</title>
      <link>https://cognaptus.com/blog/2025-08-04-forkcast-how-pro2guard-predicts-and-prevents-llm-agent-failures/</link>
      <pubDate>Mon, 04 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-08-04-forkcast-how-pro2guard-predicts-and-prevents-llm-agent-failures/</guid>
      <description>ProbGuard shows how runtime monitoring can move from catching unsafe agent actions to forecasting risky trajectories before they fail.</description>
    </item>
    <item>
      <title>From Autocomplete to Autonomy: How LLM Code Agents are Rewriting the SDLC</title>
      <link>https://cognaptus.com/blog/2025-08-04-from-autocomplete-to-autonomy-how-llm-code-agents-are-rewriting-the-sdlc/</link>
      <pubDate>Mon, 04 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-08-04-from-autocomplete-to-autonomy-how-llm-code-agents-are-rewriting-the-sdlc/</guid>
      <description>A practical map of how LLM-based code agents are moving from code completion into planning, debugging, testing, refactoring, and autonomous software workflows.</description>
    </item>
    <item>
      <title>The Lion Roars in Crypto: How Multi-Agent LLMs Are Taming Market Chaos</title>
      <link>https://cognaptus.com/blog/2025-08-03-the-lion-roars-in-crypto-how-multiagent-llms-are-taming-market-chaos/</link>
      <pubDate>Sun, 03 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-08-03-the-lion-roars-in-crypto-how-multiagent-llms-are-taming-market-chaos/</guid>
      <description>MountainLion shows how multi-agent LLM systems can compress crypto research workflows, but its strongest value is interpretability and decision support rather than proven autonomous trading alpha.</description>
    </item>
    <item>
      <title>Mind&#39;s Eye for Machines: How SimuRA Teaches AI to Think Before Acting</title>
      <link>https://cognaptus.com/blog/2025-08-02-minds-eye-for-machines-how-simura-teaches-ai-to-think-before-acting/</link>
      <pubDate>Sat, 02 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-08-02-minds-eye-for-machines-how-simura-teaches-ai-to-think-before-acting/</guid>
      <description>SimuRA shows why useful AI agents may need less blind action and more internal rehearsal before they touch the browser.</description>
    </item>
    <item>
      <title>Layers of Thought: How Hierarchical Memory Supercharges LLM Agent Reasoning</title>
      <link>https://cognaptus.com/blog/2025-08-01-layers-of-thought-how-hierarchical-memory-supercharges-llm-agent-reasoning/</link>
      <pubDate>Fri, 01 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-08-01-layers-of-thought-how-hierarchical-memory-supercharges-llm-agent-reasoning/</guid>
      <description>A practical reading of H-MEM, a hierarchical memory architecture that makes long-term LLM agents faster and more coherent by changing the shape of memory, not merely its size.</description>
    </item>
    <item>
      <title>SIMURA Says: Don’t Guess, Simulate</title>
      <link>https://cognaptus.com/blog/2025-08-01-simura-says-dont-guess-simulate/</link>
      <pubDate>Fri, 01 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-08-01-simura-says-dont-guess-simulate/</guid>
      <description>A mechanism-first look at SiRA, the agent architecture that uses explicit world-model simulation to plan before acting.</description>
    </item>
    <item>
      <title>Echo Chambers or Stubborn Minds? Simulating Social Influence with LLM Agents</title>
      <link>https://cognaptus.com/blog/2025-07-31-echo-chambers-or-stubborn-minds-simulating-social-influence-with-llm-agents/</link>
      <pubDate>Thu, 31 Jul 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-07-31-echo-chambers-or-stubborn-minds-simulating-social-influence-with-llm-agents/</guid>
      <description>A comparison-based reading of how different LLM agent types conform, polarize, or preserve dissent in synthetic forum discussions.</description>
    </item>
    <item>
      <title>Mirage Agents: When LLMs Act on Illusions</title>
      <link>https://cognaptus.com/blog/2025-07-29-mirage-agents-when-llms-act-on-illusions/</link>
      <pubDate>Tue, 29 Jul 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-07-29-mirage-agents-when-llms-act-on-illusions/</guid>
      <description>MIRAGE-Bench reframes agent hallucination as context-unfaithful action and gives operators a sharper way to test LLM agents before they touch real systems.</description>
    </item>
    <item>
      <title>From Graph to Grit: Diagnosing Warehouse Bottlenecks with LLMs and Knowledge Graphs</title>
      <link>https://cognaptus.com/blog/2025-07-26-from-graph-to-grit-diagnosing-warehouse-bottlenecks-with-llms-and-knowledge-graphs/</link>
      <pubDate>Sat, 26 Jul 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-07-26-from-graph-to-grit-diagnosing-warehouse-bottlenecks-with-llms-and-knowledge-graphs/</guid>
      <description>A mechanism-first reading of how knowledge graphs and step-wise LLM reasoning can turn warehouse simulation logs into evidence-linked bottleneck diagnosis.</description>
    </item>
    <item>
      <title>Planners, Meet Your Smart Sidekick</title>
      <link>https://cognaptus.com/blog/2025-07-26-planners-meet-your-smart-sidekick/</link>
      <pubDate>Sat, 26 Jul 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-07-26-planners-meet-your-smart-sidekick/</guid>
      <description>SMARTAPS shows how LLMs can make advanced planning systems more usable by routing planner questions to expert-built optimisation tools rather than pretending to optimise from vibes.</description>
    </item>
    <item>
      <title>The Most Dangerous Query Is the One You Don&#39;t Question</title>
      <link>https://cognaptus.com/blog/2025-07-25-the-most-dangerous-query-is-the-one-you-dont-question/</link>
      <pubDate>Fri, 25 Jul 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-07-25-the-most-dangerous-query-is-the-one-you-dont-question/</guid>
      <description>VeriMinder shows why enterprise NL2SQL systems need a pre-query reasoning layer, not just better SQL generation.</description>
    </item>
    <item>
      <title>Tools of Thought: Why Reasoning Isn’t an Illusion After All</title>
      <link>https://cognaptus.com/blog/2025-07-24-tools-of-thought-why-reasoning-isnt-an-illusion-after-all/</link>
      <pubDate>Thu, 24 Jul 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-07-24-tools-of-thought-why-reasoning-isnt-an-illusion-after-all/</guid>
      <description>A closer look at why tool-augmented reasoning models beat ordinary prompting only when the model, task, and tool interface actually fit.</description>
    </item>
    <item>
      <title>The Watchdog at the Gates: How HalMit Hunts Hallucinations in LLM Agents</title>
      <link>https://cognaptus.com/blog/2025-07-23-the-watchdog-at-the-gates-how-halmit-hunts-hallucinations-in-llm-agents/</link>
      <pubDate>Wed, 23 Jul 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-07-23-the-watchdog-at-the-gates-how-halmit-hunts-hallucinations-in-llm-agents/</guid>
      <description>HalMit reframes hallucination monitoring as boundary mapping: probe where an agent tends to fail, store those risk zones, and flag nearby queries before trust becomes expensive.</description>
    </item>
    <item>
      <title>The Butterfly Defect: Diagnosing LLM Failures in Tool-Agent Chains</title>
      <link>https://cognaptus.com/blog/2025-07-22-the-butterfly-defect-diagnosing-llm-failures-in-toolagent-chains/</link>
      <pubDate>Tue, 22 Jul 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-07-22-the-butterfly-defect-diagnosing-llm-failures-in-toolagent-chains/</guid>
      <description>A mechanism-first reading of how small parameter errors in LLM tool agents propagate into failed automation chains, and what operators should govern before they scale agents.</description>
    </item>
    <item>
      <title>Agents of Disruption: How LLMs Became Adversarial Testers for Autonomous Driving</title>
      <link>https://cognaptus.com/blog/2025-07-21-agents-of-disruption-how-llms-became-adversarial-testers-for-autonomous-driving/</link>
      <pubDate>Mon, 21 Jul 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-07-21-agents-of-disruption-how-llms-became-adversarial-testers-for-autonomous-driving/</guid>
      <description>AGENTS-LLM shows how agentic LLM loops can turn real driving logs into harder autonomous-vehicle test scenarios without pretending to generate reality from scratch.</description>
    </item>
    <item>
      <title>Game of Prompts: How Game Theory and Agentic LLMs Are Rewriting Cybersecurity</title>
      <link>https://cognaptus.com/blog/2025-07-16-game-of-prompts-how-game-theory-and-agentic-llms-are-rewriting-cybersecurity/</link>
      <pubDate>Wed, 16 Jul 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-07-16-game-of-prompts-how-game-theory-and-agentic-llms-are-rewriting-cybersecurity/</guid>
      <description>A practical reading of how game theory and agentic LLMs can reshape cybersecurity architecture, from strategic threat modeling to multi-agent SOC workflows.</description>
    </item>
    <item>
      <title>Tables Turned: Why LLM-Based Table Agents Are the Next Big Leap in Business AI</title>
      <link>https://cognaptus.com/blog/2025-07-15-tables-turned-why-llmbased-table-agents-are-the-next-big-leap-in-business-ai/</link>
      <pubDate>Tue, 15 Jul 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-07-15-tables-turned-why-llmbased-table-agents-are-the-next-big-leap-in-business-ai/</guid>
      <description>A practical reading of why table agents need operational workflows, not just better prompts or larger context windows.</description>
    </item>
    <item>
      <title>Threading the Needle: How GRAFT Reinvents Document Translation with DAGs and LLM Agents</title>
      <link>https://cognaptus.com/blog/2025-07-12-threading-the-needle-how-graft-reinvents-document-translation-with-dags-and-llm-agents/</link>
      <pubDate>Sat, 12 Jul 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-07-12-threading-the-needle-how-graft-reinvents-document-translation-with-dags-and-llm-agents/</guid>
      <description>GRAFT shows that better document translation may come less from longer context windows than from explicit discourse structure, dependency graphs, and selective memory.</description>
    </item>
    <item>
      <title>Secret Handshakes at Scale: How LLM Agents Learn to Collude</title>
      <link>https://cognaptus.com/blog/2025-07-07-secret-handshakes-at-scale-how-llm-agents-learn-to-collude/</link>
      <pubDate>Mon, 07 Jul 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-07-07-secret-handshakes-at-scale-how-llm-agents-learn-to-collude/</guid>
      <description>A mechanism-first reading of how LLM market agents coordinate prices, why communication and pressure matter, and what operators should test before deploying autonomous pricing agents.</description>
    </item>
    <item>
      <title>From ETL to Orchestral Intelligence: The Rise of the Data Agent</title>
      <link>https://cognaptus.com/blog/2025-07-03-from-etl-to-orchestral-intelligence-the-rise-of-the-data-agent/</link>
      <pubDate>Thu, 03 Jul 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-07-03-from-etl-to-orchestral-intelligence-the-rise-of-the-data-agent/</guid>
      <description>A mechanism-first reading of Data Agents as the orchestration layer that could sit above enterprise data tools, agents, benchmarks, and execution engines.</description>
    </item>
    <item>
      <title>Hive Minds and Hallucinations: A Smarter Way to Trust LLMs</title>
      <link>https://cognaptus.com/blog/2025-07-03-hive-minds-and-hallucinations-a-smarter-way-to-trust-llms/</link>
      <pubDate>Thu, 03 Jul 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-07-03-hive-minds-and-hallucinations-a-smarter-way-to-trust-llms/</guid>
      <description>A mechanism-first look at how a pharmacy SMS prototype uses deterministic parsing, fuzzy logic, and LLM cross-checking to make hallucination risk operationally manageable.</description>
    </item>
    <item>
      <title>Chains of Causality, Not Just Thought</title>
      <link>https://cognaptus.com/blog/2025-07-02-chains-of-causality-not-just-thought/</link>
      <pubDate>Wed, 02 Jul 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-07-02-chains-of-causality-not-just-thought/</guid>
      <description>How Causal Influence Prompting turns agent safety from vague warning text into an explicit causal control layer for risky actions.</description>
    </item>
    <item>
      <title>Chatbot at the Table: Rethinking Group Recommendations with GenAI</title>
      <link>https://cognaptus.com/blog/2025-07-02-chatbot-at-the-table-rethinking-group-recommendations-with-genai/</link>
      <pubDate>Wed, 02 Jul 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-07-02-chatbot-at-the-table-rethinking-group-recommendations-with-genai/</guid>
      <description>A mechanism-first reading of why generative AI may turn group recommenders from ranking engines into decision facilitators.</description>
    </item>
    <item>
      <title>Grounded and Confused: Why RAG Systems Still Fail in the Enterprise</title>
      <link>https://cognaptus.com/blog/2025-07-01-grounded-and-confused-why-rag-systems-still-fail-in-the-enterprise/</link>
      <pubDate>Tue, 01 Jul 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-07-01-grounded-and-confused-why-rag-systems-still-fail-in-the-enterprise/</guid>
      <description>HERB shows why enterprise RAG failures are less about chatbot polish and more about evidence retrieval across messy, heterogeneous work.</description>
    </item>
    <item>
      <title>Catalysts of Thought: How LLM Agents are Reinventing Chemical Process Optimization</title>
      <link>https://cognaptus.com/blog/2025-06-27-catalysts-of-thought-how-llm-agents-are-reinventing-chemical-process-optimization/</link>
      <pubDate>Fri, 27 Jun 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-06-27-catalysts-of-thought-how-llm-agents-are-reinventing-chemical-process-optimization/</guid>
      <description>A mechanism-first reading of how multi-agent LLM systems can infer missing process constraints, guide simulations, and reduce the setup bottleneck in chemical optimisation.</description>
    </item>
    <item>
      <title>Mind Games for Machines: How Decrypto Reveals the Hidden Gaps in AI Reasoning</title>
      <link>https://cognaptus.com/blog/2025-06-26-mind-games-for-machines-how-decrypto-reveals-the-hidden-gaps-in-ai-reasoning/</link>
      <pubDate>Thu, 26 Jun 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-06-26-mind-games-for-machines-how-decrypto-reveals-the-hidden-gaps-in-ai-reasoning/</guid>
      <description>A mechanism-first look at Decrypto, a benchmark showing why strong single-agent reasoning does not automatically translate into social reasoning, coordination, or theory of mind.</description>
    </item>
    <item>
      <title>The Joy of Many Minds: How JoyAgents-R1 Unleashes the Power of Multi-LLM Reinforcement Learning</title>
      <link>https://cognaptus.com/blog/2025-06-25-the-joy-of-many-minds-how-joyagentsr1-unleashes-the-power-of-multillm-reinforcement-learning/</link>
      <pubDate>Wed, 25 Jun 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-06-25-the-joy-of-many-minds-how-joyagentsr1-unleashes-the-power-of-multillm-reinforcement-learning/</guid>
      <description>HiMA-Ecom and HiMA-R1 show how vertical-domain agent teams can be trained jointly, remembered selectively, and evaluated more honestly than ordinary chatbot benchmarks allow.</description>
    </item>
    <item>
      <title>From Sparse to Smart: How PROGRM Elevates GUI Agent Training</title>
      <link>https://cognaptus.com/blog/2025-05-26-from-sparse-to-smart-how-progrm-elevates-gui-agent-training/</link>
      <pubDate>Mon, 26 May 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-05-26-from-sparse-to-smart-how-progrm-elevates-gui-agent-training/</guid>
      <description>ProgRM shows that GUI agents may learn more from measuring partial progress than from waiting for a final pass/fail signal.</description>
    </item>
    <item>
      <title>Divide and Model: How Multi-Agent LLMs Are Rethinking Real-World Problem Solving</title>
      <link>https://cognaptus.com/blog/2025-05-23-divide-and-model-how-multiagent-llms-are-rethinking-realworld-problem-solving/</link>
      <pubDate>Fri, 23 May 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-05-23-divide-and-model-how-multiagent-llms-are-rethinking-realworld-problem-solving/</guid>
      <description>ModelingAgent shows that real-world AI problem solving improves less from raw tool access than from structured agent roles, shared memory, and critic-driven refinement.</description>
    </item>
    <item>
      <title>Mind the Context: How ContextAgent Listens, Sees, and Acts Before You Ask</title>
      <link>https://cognaptus.com/blog/2025-05-21-mind-the-context-how-contextagent-listens-sees-and-acts-before-you-ask/</link>
      <pubDate>Wed, 21 May 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-05-21-mind-the-context-how-contextagent-listens-sees-and-acts-before-you-ask/</guid>
      <description>ContextAgent shows that proactive AI assistants are less about speaking first and more about sensing, scoring, and acting only when context makes interruption worthwhile.</description>
    </item>
    <item>
      <title>Reflections in the Mirror Maze: Why LLM Reasoning Isn&#39;t Quite There Yet</title>
      <link>https://cognaptus.com/blog/2025-05-17-reflections-in-the-mirror-maze-why-llm-reasoning-isnt-quite-there-yet/</link>
      <pubDate>Sat, 17 May 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-05-17-reflections-in-the-mirror-maze-why-llm-reasoning-isnt-quite-there-yet/</guid>
      <description>A practical reading of why reflection, planning, and heuristic prompts help LLM agents only when the task, model, and feedback loop are aligned.</description>
    </item>
    <item>
      <title>Flashcards for Giants: How RAL Lets Large Models Learn Without Fine-Tuning</title>
      <link>https://cognaptus.com/blog/2025-05-06-flashcards-for-giants-how-ral-lets-large-models-learn-without-finetuning/</link>
      <pubDate>Tue, 06 May 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-05-06-flashcards-for-giants-how-ral-lets-large-models-learn-without-finetuning/</guid>
      <description>A practical reading of Retrieval Augmented Learning, a train-free framework that lets LLM agents build validated experience memories through retrial rather than parameter updates.</description>
    </item>
    <item>
      <title>When Smart AI Gets It Wrong: Diagnosing the Knowing-Doing Gap in Language Model Agents</title>
      <link>https://cognaptus.com/blog/2025-04-23-when-smart-ai-gets-it-wrong-diagnosing-the-knowingdoing-gap-in-language-model-agents/</link>
      <pubDate>Wed, 23 Apr 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-04-23-when-smart-ai-gets-it-wrong-diagnosing-the-knowingdoing-gap-in-language-model-agents/</guid>
      <description>A mechanism-first reading of why LLM agents can explain good decisions yet still act greedily, and what that means for enterprise automation.</description>
    </item>
    <item>
      <title>Overqualified, Underprepared: Why FinLLMs Matter More Than Reasoning</title>
      <link>https://cognaptus.com/blog/2025-04-20-overqualified-underprepared-why-finllms-matter-more-than-reasoning/</link>
      <pubDate>Sun, 20 Apr 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-04-20-overqualified-underprepared-why-finllms-matter-more-than-reasoning/</guid>
      <description>A practical reading of three finance-AI papers: useful FinLLMs are not market oracles, but components in a layered decision stack.</description>
    </item>
    <item>
      <title>Agents in Formation: Fine-Tune Meets Fine-Structure in Quant AI</title>
      <link>https://cognaptus.com/blog/2025-04-17-agents-in-formation-finetune-meets-finestructure-in-quant-ai/</link>
      <pubDate>Thu, 17 Apr 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-04-17-agents-in-formation-finetune-meets-finestructure-in-quant-ai/</guid>
      <description>A business-oriented reading of how adaptive workflows and verifier-trained reasoning models point toward more reliable vertical AI systems.</description>
    </item>
    <item>
      <title>Case Closed: How CBR-LLMs Unlock Smarter Business Automation</title>
      <link>https://cognaptus.com/blog/2025-04-10-case-closed-how-cbrllms-unlock-smarter-business-automation/</link>
      <pubDate>Thu, 10 Apr 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-04-10-case-closed-how-cbrllms-unlock-smarter-business-automation/</guid>
      <description>A comparison-based reading of why case-based reasoning may give enterprise LLM agents something plain RAG often lacks: structured precedent, adaptation, and accountable memory.</description>
    </item>
    <item>
      <title>Passing as Human: How AI Personas Are Rewriting the Marketing Playbook</title>
      <link>https://cognaptus.com/blog/2025-04-07-passing-as-human-how-ai-personas-are-rewriting-the-marketing-playbook/</link>
      <pubDate>Mon, 07 Apr 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-04-07-passing-as-human-how-ai-personas-are-rewriting-the-marketing-playbook/</guid>
      <description>A practical guide to using AI personas for marketing simulation without mistaking synthetic shoppers for real customers.</description>
    </item>
    <item>
      <title>From Gomoku AI to Boardroom Breakthroughs: How Generative AI Can Transform Corporate Strategy</title>
      <link>https://cognaptus.com/blog/2025-03-28-from-gomoku-ai-to-boardroom-breakthroughs/</link>
      <pubDate>Fri, 28 Mar 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-03-28-from-gomoku-ai-to-boardroom-breakthroughs/</guid>
      <description>A practical interpretation of LLM-Gomoku as a design pattern for AI-assisted corporate strategy: structured options, constraint checks, feedback loops, and disciplined human oversight.</description>
    </item>
  </channel>
</rss>
