Agreeable to a Fault: Why LLM ‘People’ Can’t Hold Their Ground
A mechanism-first look at why synthetic LLM personas can sound socially plausible while failing stricter tests of behavioural coherence.
A mechanism-first look at why synthetic LLM personas can sound socially plausible while failing stricter tests of behavioural coherence.
ArcMemo shows why useful agent memory is less about storing past prompts and more about distilling verified reasoning into reusable, selectable concepts.
JD.com’s supply-chain agent framework shows that the business value of GenAI planning is not prettier forecasts, but a shorter loop between intent, execution, diagnosis, and correction.
A new agent-planning paper shows why the best LLM agents should treat explicit reasoning as a scarce operational resource, not a reflex.
A mechanism-first reading of Meta-Policy Reflexion, a training-free approach that turns failed agent trajectories into reusable rules and test-time guardrails.
BARGAIN shows how LLM cascades can reduce expensive model calls while preserving finite-sample guarantees on accuracy, precision, or recall.
MIT’s Palimpzest prototype shows why enterprise Deep Research needs query planning, semantic operators, and reusable context—not just a clever agent with a Python console.
HF-RAG shows why enterprise retrieval systems should stop choosing between labelled exemplars and open corpora, and start making their scores comparable.
A production study of app.build shows why reliable agentic software generation depends less on model size than on structured environments, targeted validation, and repair loops.
A mechanism-first reading of InAbHyD, a benchmark showing why LLMs can explain observations without finding the simplest useful hypothesis.