<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Software Engineering on Cognaptus</title>
    <link>https://cognaptus.com/tags/software-engineering/</link>
    <description>Recent content in Software Engineering on Cognaptus</description>
    <generator>Hugo -- 0.145.0</generator>
    <language>en-us</language>
    <lastBuildDate>Sat, 27 Jun 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://cognaptus.com/tags/software-engineering/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Memory Has to Earn Its Keep</title>
      <link>https://cognaptus.com/blog/2026-06-27-memory-has-to-earn-its-keep/</link>
      <pubDate>Sat, 27 Jun 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-06-27-memory-has-to-earn-its-keep/</guid>
      <description>A mechanism-first reading of MemOp, a closed-loop framework that treats coding-agent memory as an evaluated and optimized operational asset rather than a bigger scrapbook.</description>
    </item>
    <item>
      <title>The Code Agent Wasn’t Self-Correcting. The Test Harness Was.</title>
      <link>https://cognaptus.com/blog/2026-06-22-the-code-agent-wasnt-selfcorrecting-the-test-harness-was/</link>
      <pubDate>Mon, 22 Jun 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-06-22-the-code-agent-wasnt-selfcorrecting-the-test-harness-was/</guid>
      <description>A mechanism-first reading of why execution-feedback loops make LLM coding assistants more useful, but only for the failures that feedback can actually localize.</description>
    </item>
    <item>
      <title>Compile Once, Train Later: Offline RL Moves Code-Model Verification Upstream</title>
      <link>https://cognaptus.com/blog/2026-06-03-compile-once-train-later-offline-rl-moves-codemodel-verification-upstream/</link>
      <pubDate>Wed, 03 Jun 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-06-03-compile-once-train-later-offline-rl-moves-codemodel-verification-upstream/</guid>
      <description>A mechanism-first reading of how offline reinforcement learning can post-train code models by turning pre-verified code datasets into cheaper, harder-task learning signals.</description>
    </item>
    <item>
      <title>Memory Lane Meets Mainframe: Why Coding Agents Need Better Memories, Not Bigger Egos</title>
      <link>https://cognaptus.com/blog/2026-04-16-memory-lane-meets-mainframe-why-coding-agents-need-better-memories-not-bigger-egos/</link>
      <pubDate>Thu, 16 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-16-memory-lane-meets-mainframe-why-coding-agents-need-better-memories-not-bigger-egos/</guid>
      <description>A mechanism-first reading of Memory Transfer Learning, showing why coding-agent memory works best when it transfers abstract operational discipline rather than brittle code traces.</description>
    </item>
    <item>
      <title>Thinking Fast, Remembering Slow: Why SWE-AGILE Fixes the Memory Crisis of AI Agents</title>
      <link>https://cognaptus.com/blog/2026-04-14-thinking-fast-remembering-slow-why-sweagile-fixes-the-memory-crisis-of-ai-agents/</link>
      <pubDate>Tue, 14 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-14-thinking-fast-remembering-slow-why-sweagile-fixes-the-memory-crisis-of-ai-agents/</guid>
      <description>A mechanism-first reading of SWE-AGILE: why the next bottleneck for AI agents is not only reasoning depth, but remembering the right layer of reasoning at the right cost.</description>
    </item>
    <item>
      <title>Proofs at Scale: When 30,000 Agents Replace the Referee</title>
      <link>https://cognaptus.com/blog/2026-04-06-proofs-at-scale-when-30000-agents-replace-the-referee/</link>
      <pubDate>Mon, 06 Apr 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-04-06-proofs-at-scale-when-30000-agents-replace-the-referee/</guid>
      <description>A mechanism-first reading of automatic textbook formalization: why the breakthrough is not just stronger theorem proving, but disciplined agent orchestration at repository scale.</description>
    </item>
    <item>
      <title>Double Helix, Double Checks: Why Agentic AI Needs Governance Before It Writes Your Code</title>
      <link>https://cognaptus.com/blog/2026-03-05-double-helix-double-checks-why-agentic-ai-needs-governance-before-it-writes-your-code/</link>
      <pubDate>Thu, 05 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-03-05-double-helix-double-checks-why-agentic-ai-needs-governance-before-it-writes-your-code/</guid>
      <description>A WebGIS case study shows why reliable agentic AI depends less on bigger prompts and more on persistent memory, enforceable rules, and auditable workflow structure.</description>
    </item>
    <item>
      <title>Agents That Hire Themselves: Why OpenSage Signals the End of Hand-Crafted AI Workflows</title>
      <link>https://cognaptus.com/blog/2026-02-21-agents-that-hire-themselves-why-opensage-signals-the-end-of-handcrafted-ai-workflows/</link>
      <pubDate>Sat, 21 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-21-agents-that-hire-themselves-why-opensage-signals-the-end-of-handcrafted-ai-workflows/</guid>
      <description>OpenSage shows why the next bottleneck in business automation may be agent infrastructure: systems that let models create sub-agents, tools, and structured memory at runtime.</description>
    </item>
    <item>
      <title>Breaking Things on Purpose: How CLI-Gym Teaches AI to Fix the Real World</title>
      <link>https://cognaptus.com/blog/2026-02-13-breaking-things-on-purpose-how-cligym-teaches-ai-to-fix-the-real-world/</link>
      <pubDate>Fri, 13 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-13-breaking-things-on-purpose-how-cligym-teaches-ai-to-fix-the-real-world/</guid>
      <description>A mechanism-first reading of CLI-Gym, a pipeline that turns working Dockerized repositories into scalable environment-repair tasks for stronger coding agents.</description>
    </item>
    <item>
      <title>When Agents Believe Their Own Hype: The Hidden Cost of Agentic Overconfidence</title>
      <link>https://cognaptus.com/blog/2026-02-09-when-agents-believe-their-own-hype-the-hidden-cost-of-agentic-overconfidence/</link>
      <pubDate>Mon, 09 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-09-when-agents-believe-their-own-hype-the-hidden-cost-of-agentic-overconfidence/</guid>
      <description>A comparison-based reading of agentic uncertainty research, showing why AI agents’ confidence scores are useful for routing work but dangerous as acceptance signals.</description>
    </item>
    <item>
      <title>Tokens, Watts, and Waste: The Hidden Energy Bill of LLM Inference</title>
      <link>https://cognaptus.com/blog/2026-02-08-tokens-watts-and-waste-the-hidden-energy-bill-of-llm-inference/</link>
      <pubDate>Sun, 08 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-02-08-tokens-watts-and-waste-the-hidden-energy-bill-of-llm-inference/</guid>
      <description>A mechanism-first reading of why LLM inference energy is shaped by prefill, decoding, prompt length, and unnecessary generation—not merely model size.</description>
    </item>
    <item>
      <title>Vibe Coding a Theorem Prover: When LLMs Prove (and Break) Themselves</title>
      <link>https://cognaptus.com/blog/2026-01-11-vibe-coding-a-theorem-prover-when-llms-prove-and-break-themselves/</link>
      <pubDate>Sun, 11 Jan 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-01-11-vibe-coding-a-theorem-prover-when-llms-prove-and-break-themselves/</guid>
      <description>Why Isabellm’s real lesson is not autonomous AI reasoning, but verifier-gated system design for domains where being plausibly right is still wrong.</description>
    </item>
    <item>
      <title>When Solvers Guess Smarter: Teaching SMT to Think in Functions</title>
      <link>https://cognaptus.com/blog/2026-01-11-when-solvers-guess-smarter-teaching-smt-to-think-in-functions/</link>
      <pubDate>Sun, 11 Jan 2026 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2026-01-11-when-solvers-guess-smarter-teaching-smt-to-think-in-functions/</guid>
      <description>AquaForte shows how LLMs can guide quantified SMT solving by proposing mathematical function instantiations while traditional solvers keep the formal guarantees.</description>
    </item>
    <item>
      <title>Many Arms, Fewer Bugs: Why Coding Agents Need to Stop Working Alone</title>
      <link>https://cognaptus.com/blog/2025-12-31-many-arms-fewer-bugs-why-coding-agents-need-to-stop-working-alone/</link>
      <pubDate>Wed, 31 Dec 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-12-31-many-arms-fewer-bugs-why-coding-agents-need-to-stop-working-alone/</guid>
      <description>BOAD shows that coding-agent performance depends less on assembling more agents than on discovering a small team, assigning individual credit, and controlling what each agent needs to remember.</description>
    </item>
    <item>
      <title>Guardrails Over Gigabytes: Making LLM Coding Agents Behave</title>
      <link>https://cognaptus.com/blog/2025-12-27-guardrails-over-gigabytes-making-llm-coding-agents-behave/</link>
      <pubDate>Sat, 27 Dec 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-12-27-guardrails-over-gigabytes-making-llm-coding-agents-behave/</guid>
      <description>A mechanism-first reading of why deterministic post-condition guards can make LLM coding agents more reliable—while still failing to solve autonomous software repair.</description>
    </item>
    <item>
      <title>When Agents Compare Notes: How Shared Memory Quietly Rewires Software Development</title>
      <link>https://cognaptus.com/blog/2025-11-15-when-agents-compare-notes-how-shared-memory-quietly-rewires-software-development/</link>
      <pubDate>Sat, 15 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-15-when-agents-compare-notes-how-shared-memory-quietly-rewires-software-development/</guid>
      <description>Spark shows why the next leap in coding agents may come less from bigger models than from shared, curated experience.</description>
    </item>
    <item>
      <title>The Esperanto of AI Agents: How the Agent Data Protocol Unifies a Fragmented Ecosystem</title>
      <link>https://cognaptus.com/blog/2025-11-02-the-esperanto-of-ai-agents-how-the-agent-data-protocol-unifies-a-fragmented-ecosystem/</link>
      <pubDate>Sun, 02 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-02-the-esperanto-of-ai-agents-how-the-agent-data-protocol-unifies-a-fragmented-ecosystem/</guid>
      <description>ADP reframes agent training as a data interoperability problem, showing how typed trajectories can turn scattered agent datasets into reusable fine-tuning infrastructure.</description>
    </item>
    <item>
      <title>When Agents Learn to Test Themselves: TDFlow and the Future of Software Engineering</title>
      <link>https://cognaptus.com/blog/2025-11-02-when-agents-learn-to-test-themselves-tdflow-and-the-future-of-software-engineering/</link>
      <pubDate>Sun, 02 Nov 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-11-02-when-agents-learn-to-test-themselves-tdflow-and-the-future-of-software-engineering/</guid>
      <description>TDFlow shows that coding agents become far more useful when humans define correctness as executable tests and agents are constrained to solve them.</description>
    </item>
    <item>
      <title>Pipes by Prompt, DAGs by Design: Why Hybrid Beats Hero Prompts</title>
      <link>https://cognaptus.com/blog/2025-10-01-pipes-by-prompt-dags-by-design-why-hybrid-beats-hero-prompts/</link>
      <pubDate>Wed, 01 Oct 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-10-01-pipes-by-prompt-dags-by-design-why-hybrid-beats-hero-prompts/</guid>
      <description>Prompt2DAG shows why reliable AI-generated data pipelines need structured intermediate representations, validation gates, and template scaffolding—not heroic one-shot prompts.</description>
    </item>
    <item>
      <title>Guard Rails &gt; Horsepower: Why Environment Scaffolding Beats Bigger Models</title>
      <link>https://cognaptus.com/blog/2025-09-06-guard-rails-horsepower-why-environment-scaffolding-beats-bigger-models/</link>
      <pubDate>Sat, 06 Sep 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-09-06-guard-rails-horsepower-why-environment-scaffolding-beats-bigger-models/</guid>
      <description>A production study of app.build shows why reliable agentic software generation depends less on model size than on structured environments, targeted validation, and repair loops.</description>
    </item>
    <item>
      <title>Mask, Don’t Muse: When Simple Memory Beats Fancy Summaries</title>
      <link>https://cognaptus.com/blog/2025-09-01-mask-dont-muse-when-simple-memory-beats-fancy-summaries/</link>
      <pubDate>Mon, 01 Sep 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-09-01-mask-dont-muse-when-simple-memory-beats-fancy-summaries/</guid>
      <description>A practical reading of why simple observation masking can beat expensive LLM summaries for software-engineering agents.</description>
    </item>
    <item>
      <title>Wheel Smarts &gt; Wheel Reinvention: What GitTaskBench Really Measures</title>
      <link>https://cognaptus.com/blog/2025-08-27-wheel-smarts-wheel-reinvention-what-gittaskbench-really-measures/</link>
      <pubDate>Wed, 27 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-08-27-wheel-smarts-wheel-reinvention-what-gittaskbench-really-measures/</guid>
      <description>GitTaskBench shows that code-agent value depends less on writing fresh code and more on surviving the messy chain from repository comprehension to execution, quality, and cost.</description>
    </item>
    <item>
      <title>Longer Yet Dumber: Why LLMs Fail at Catching Their Own Coding Mistakes</title>
      <link>https://cognaptus.com/blog/2025-08-06-longer-yet-dumber-why-llms-fail-at-catching-their-own-coding-mistakes/</link>
      <pubDate>Wed, 06 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-08-06-longer-yet-dumber-why-llms-fail-at-catching-their-own-coding-mistakes/</guid>
      <description>FPBench shows that many LLM code assistants can detect faulty requirements when prompted, but often fail to question bad premises on their own.</description>
    </item>
    <item>
      <title>From Autocomplete to Autonomy: How LLM Code Agents are Rewriting the SDLC</title>
      <link>https://cognaptus.com/blog/2025-08-04-from-autocomplete-to-autonomy-how-llm-code-agents-are-rewriting-the-sdlc/</link>
      <pubDate>Mon, 04 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-08-04-from-autocomplete-to-autonomy-how-llm-code-agents-are-rewriting-the-sdlc/</guid>
      <description>A practical map of how LLM-based code agents are moving from code completion into planning, debugging, testing, refactoring, and autonomous software workflows.</description>
    </item>
    <item>
      <title>The Butterfly Defect: Diagnosing LLM Failures in Tool-Agent Chains</title>
      <link>https://cognaptus.com/blog/2025-07-22-the-butterfly-defect-diagnosing-llm-failures-in-toolagent-chains/</link>
      <pubDate>Tue, 22 Jul 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-07-22-the-butterfly-defect-diagnosing-llm-failures-in-toolagent-chains/</guid>
      <description>A mechanism-first reading of how small parameter errors in LLM tool agents propagate into failed automation chains, and what operators should govern before they scale agents.</description>
    </item>
    <item>
      <title>The Debugger Awakens: Why Kodezi Chronos Leaves GPT-4 in the Dust</title>
      <link>https://cognaptus.com/blog/2025-07-19-the-debugger-awakens-why-kodezi-chronos-leaves-gpt4-in-the-dust/</link>
      <pubDate>Sat, 19 Jul 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-07-19-the-debugger-awakens-why-kodezi-chronos-leaves-gpt4-in-the-dust/</guid>
      <description>A mechanism-first look at why Kodezi Chronos treats debugging as a repository-scale maintenance workflow rather than a longer code-completion prompt.</description>
    </item>
    <item>
      <title>Beyond Stack Overflow: CodeAssistBench Exposes the Real Gaps in LLM Coding Help</title>
      <link>https://cognaptus.com/blog/2025-07-16-beyond-stack-overflow-codeassistbench-exposes-the-real-gaps-in-llm-coding-help/</link>
      <pubDate>Wed, 16 Jul 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-07-16-beyond-stack-overflow-codeassistbench-exposes-the-real-gaps-in-llm-coding-help/</guid>
      <description>CodeAssistBench shows why coding assistants that shine on Q&amp;amp;A benchmarks still struggle inside real, recent, multi-turn software support workflows.</description>
    </item>
    <item>
      <title>The First Hurdle: Why Coding Agents Struggle with Setup</title>
      <link>https://cognaptus.com/blog/2025-07-15-the-first-hurdle-why-coding-agents-struggle-with-setup/</link>
      <pubDate>Tue, 15 Jul 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-07-15-the-first-hurdle-why-coding-agents-struggle-with-setup/</guid>
      <description>SetupBench shows that coding agents still struggle with the unglamorous first mile of software work: making the project actually run.</description>
    </item>
    <item>
      <title>Guardians of the Chain: How Smart-LLaMA-DPO Turns Code into Clarity</title>
      <link>https://cognaptus.com/blog/2025-06-24-guardians-of-the-chain-how-smartllamadpo-turns-code-into-clarity/</link>
      <pubDate>Tue, 24 Jun 2025 00:00:00 +0000</pubDate>
      <guid>https://cognaptus.com/blog/2025-06-24-guardians-of-the-chain-how-smartllamadpo-turns-code-into-clarity/</guid>
      <description>Smart-LLaMA-DPO shows that the next useful leap in AI security tooling may come from expert preference training, not larger generic models.</description>
    </item>
  </channel>
</rss>
