Cover image

Prompted and Confused: When LLMs Forget the Assignment

A close reading of why LLM-generated optimisation models can look correct, compile occasionally, and still misunderstand the problem hiding in plain sight.

November 20, 2025 · 14 min · Zelina
Cover image

Skills to Pay the Agent Bills: Why LLMs Need Better Moves, Not Bigger Models

SkillGen shows why the next gain in LLM agents may come from reusable procedural skills, not longer prompts or larger models.

November 20, 2025 · 18 min · Zelina
Cover image

Thresholds, Trade-offs, and the Art of Not Overthinking Your Robot

How calibrated symbolic uncertainty helps robots decide when to act, when to look again, and when confidence becomes expensive.

November 20, 2025 · 14 min · Zelina
Cover image

Tools of Habit: Why LLM Agents Benefit from a Little Inertia

AutoTool shows how agent systems can cut repeated tool-selection costs by learning when workflow habits are reliable enough to bypass another LLM call.

November 20, 2025 · 14 min · Zelina
Cover image

Value Collision Course: When LLM Alignment Plays Favorites

A mechanism-first reading of how human-feedback design choices quietly decide whose values an aligned model learns.

November 20, 2025 · 14 min · Zelina
Cover image

Ask, Navigate, Repeat: Why Socially Aware Agents Are the Next Frontier

FreeAskWorld shows why embodied AI needs interaction as an operational information channel, not just prettier simulation scenery.

November 18, 2025 · 15 min · Zelina
Cover image

Benchmarked Brilliance: How CreBench Rewrites the Rules of Machine Creativity

CreBench shows why evaluating AI creativity requires rubrics for ideas, process, and products—not another beauty contest for generated images.

November 18, 2025 · 14 min · Zelina
Cover image

Ghostwriters in the Machine: How Multi‑Agent LLMs Turn Raw Transport Data Into Decisions

A new multi-agent LLM framework shows how transport analytics can become stakeholder-ready reports, provided we remember it is automating interpretation, not operational judgement.

November 18, 2025 · 14 min · Zelina
Cover image

Graph Medicine: When RAG Stops Guessing and Starts Diagnosing

A mechanism-first look at how retrieval-augmented LLMs can turn clinical guidelines into structured medical knowledge graphs—and why the hard part is still clinical reliability.

November 18, 2025 · 15 min · Zelina
Cover image

LLMs, Trade-Offs, and the Illusion of Choice: When AI Preferences Fall Apart

A new preference-coherence test shows that many frontier LLMs can produce trade-off behaviour, but very few show stable preference structures across AI-specific scenarios.

November 18, 2025 · 12 min · Zelina