Agents, Not Tasks: Rethinking Business Processes in the Age of AI
A mechanism-first reading of how agentic AI changes business process design from fixed task flows to goal-driven coordination.
A mechanism-first reading of how agentic AI changes business process design from fixed task flows to goal-driven coordination.
ChartM3 shows why chart editing needs more than language: models must connect what users point at to the code that actually controls the visual object.
A clearer operator-focused reading of how circuit-level interpretability turns transformer explanations from attractive diagrams into testable engineering claims.
A practical reading of DGP, a graph-enhanced LLM framework that improves fraud detection by preserving target-node detail while compressing noisy neighbourhood context.
IBM’s OneShield paper shows why enterprise LLM safety is becoming a configurable control plane, not merely a better moderation classifier.
A close reading of UserBench and what it reveals about the gap between tool-using agents and genuinely user-centred assistants.
Warmth in AI assistants may improve the user experience, but this paper shows it can also make models less reliable and more sycophantic.
A practical look at FRED, a finance-focused framework for detecting and editing hallucinations in retrieval-grounded LLM outputs.
A practical reading of programmable virtual humans as a proposed bridge between molecular AI, digital twins, and physiology-first drug discovery.
MIRAGE-Bench reframes agent hallucination as context-unfaithful action and gives operators a sharper way to test LLM agents before they touch real systems.