Cover image

Correct to Select: Choosing OCR Without Ground Truth

DocOCR-Eval turns MLLM correction into a proxy signal for choosing an OCR engine before a document collection has been manually transcribed.

August 11, 2026 · 8 min · Zelina
Cover image

Reasoning Tokens Are Compute, Not an Audit Trail

A mechanism-centered survey explains when chain-of-thought adds useful computation, why fluent rationales can still be unfaithful, and how teams should test reasoning systems before deployment.

August 11, 2026 · 7 min · Zelina
Cover image

One Demonstration Is Not One Deployment: Regrind’s Real Robotics Lesson

Regrind shows how one human demonstration can accelerate dexterous robot training—and why simulation success still requires strict hardware validation.

August 10, 2026 · 8 min · Zelina
Cover image

Shared Blind Spots: Why Narrow AI Needs an Independent Detector

A formal model shows when narrow systems with escalation create value—and when a monitor sharing the same blind spot makes the safety case collapse.

August 10, 2026 · 6 min · Zelina
Cover image

Stable Enough to Be Wrong: Why Neuron Selectors Need Causal Audits

A direct intervention test shows why reproducible neuron rankings can still misidentify the components that actually carry model capability or refusal behavior.

August 10, 2026 · 9 min · Zelina
Cover image

Busy Compute, Idle Memory: ServerlessT2I Rebuilds Image Serving Around Model DAGs

ServerlessT2I shows how model-granular workflow serving can turn unused GPU memory into higher image-generation capacity, tighter SLOs, and more defensible tenant accounting.

August 9, 2026 · 7 min · Zelina
Cover image

Clean Less, Route Better: DataOrchestra Reframes Pretraining Data Curation

DataOrchestra shows how per-example routing can improve pretraining-data utility while avoiding the cost and damage of unnecessary rewriting.

August 9, 2026 · 8 min · Zelina
Cover image

Ten Clusters, One Dominant Signal: Rethinking Learner Archetypes in Edtech

A national-scale GCSE study finds that detailed learner profiles add only modest predictive value beyond overall performance, narrowing where complex personalisation is worth deploying.

August 9, 2026 · 6 min · Zelina
Cover image

Common Is Not Defining: Testing Whether Language Models Understand Category Relations

A prevalence-controlled test shows why semantic similarity can overstate conceptual understanding—and how model-review teams can evaluate the difference.

August 8, 2026 · 6 min · Zelina
Cover image

Reasoning on Demand: AdaHome’s Case for Tiered Local Assistants

AdaHome shows how selective reasoning and dedicated preference memory can improve the accuracy, efficiency, and adaptability of a locally deployed smart-home assistant.

August 8, 2026 · 7 min · Zelina