Cover image

Agreeable to a Fault: Why LLM ‘People’ Can’t Hold Their Ground

A mechanism-first look at why synthetic LLM personas can sound socially plausible while failing stricter tests of behavioural coherence.

September 8, 2025 · 14 min · Zelina
Cover image

Pieces, Not Puzzles: How ArcMemo Turns LLM Reasoning into Reusable Skills

ArcMemo shows why useful agent memory is less about storing past prompts and more about distilling verified reasoning into reusable, selectable concepts.

September 8, 2025 · 15 min · Zelina
Cover image

Plan, Act, Replan: When LLM Agents Run the Aisles

JD.com’s supply-chain agent framework shows that the business value of GenAI planning is not prettier forecasts, but a shorter loop between intent, execution, diagnosis, and correction.

September 8, 2025 · 13 min · Zelina
Cover image

Plan, Don't Spam: The Goldilocks Rule for Test‑Time Compute

A new agent-planning paper shows why the best LLM agents should treat explicit reasoning as a scarce operational resource, not a reflex.

September 8, 2025 · 15 min · Zelina
Cover image

Rules of Engagement: How Meta‑Policy Reflexion Turns Agent Memory into Guardrails

A mechanism-first reading of Meta-Policy Reflexion, a training-free approach that turns failed agent trajectories into reusable rules and test-time guardrails.

September 8, 2025 · 14 min · Zelina
Cover image

Cheap Thrills, Hard Guarantees: BARGAINing with LLM Cascades

BARGAIN shows how LLM cascades can reduce expensive model calls while preserving finite-sample guarantees on accuracy, precision, or recall.

September 6, 2025 · 17 min · Zelina
Cover image

Deep Queries, Fast Answers: Why ‘Deep Research’ Wants to Be Your New Analytics Runtime

MIT’s Palimpzest prototype shows why enterprise Deep Research needs query planning, semantic operators, and reusable context—not just a clever agent with a Python console.

September 6, 2025 · 16 min · Zelina
Cover image

Fusion Cuisine for RAG: Z‑Scores, Rankers, and the Two‑Source Diet

HF-RAG shows why enterprise retrieval systems should stop choosing between labelled exemplars and open corpora, and start making their scores comparable.

September 6, 2025 · 15 min · Zelina
Cover image

Guard Rails > Horsepower: Why Environment Scaffolding Beats Bigger Models

A production study of app.build shows why reliable agentic software generation depends less on model size than on structured environments, targeted validation, and repair loops.

September 6, 2025 · 14 min · Zelina
Cover image

Razor Burn: Why LLMs Nick Themselves on Induction and Abduction

A mechanism-first reading of InAbHyD, a benchmark showing why LLMs can explain observations without finding the simplest useful hypothesis.

September 6, 2025 · 15 min · Zelina