Prompted and Confused: When LLMs Forget the Assignment
A close reading of why LLM-generated optimisation models can look correct, compile occasionally, and still misunderstand the problem hiding in plain sight.
A close reading of why LLM-generated optimisation models can look correct, compile occasionally, and still misunderstand the problem hiding in plain sight.
SkillGen shows why the next gain in LLM agents may come from reusable procedural skills, not longer prompts or larger models.
How calibrated symbolic uncertainty helps robots decide when to act, when to look again, and when confidence becomes expensive.
AutoTool shows how agent systems can cut repeated tool-selection costs by learning when workflow habits are reliable enough to bypass another LLM call.
A mechanism-first reading of how human-feedback design choices quietly decide whose values an aligned model learns.
FreeAskWorld shows why embodied AI needs interaction as an operational information channel, not just prettier simulation scenery.
CreBench shows why evaluating AI creativity requires rubrics for ideas, process, and products—not another beauty contest for generated images.
A new multi-agent LLM framework shows how transport analytics can become stakeholder-ready reports, provided we remember it is automating interpretation, not operational judgement.
A mechanism-first look at how retrieval-augmented LLMs can turn clinical guidelines into structured medical knowledge graphs—and why the hard part is still clinical reliability.
A new preference-coherence test shows that many frontier LLMs can produce trade-off behaviour, but very few show stable preference structures across AI-specific scenarios.