Correct to Select: Choosing OCR Without Ground Truth
DocOCR-Eval turns MLLM correction into a proxy signal for choosing an OCR engine before a document collection has been manually transcribed.
DocOCR-Eval turns MLLM correction into a proxy signal for choosing an OCR engine before a document collection has been manually transcribed.
A mechanism-centered survey explains when chain-of-thought adds useful computation, why fluent rationales can still be unfaithful, and how teams should test reasoning systems before deployment.
Regrind shows how one human demonstration can accelerate dexterous robot training—and why simulation success still requires strict hardware validation.
A formal model shows when narrow systems with escalation create value—and when a monitor sharing the same blind spot makes the safety case collapse.
A direct intervention test shows why reproducible neuron rankings can still misidentify the components that actually carry model capability or refusal behavior.
ServerlessT2I shows how model-granular workflow serving can turn unused GPU memory into higher image-generation capacity, tighter SLOs, and more defensible tenant accounting.
DataOrchestra shows how per-example routing can improve pretraining-data utility while avoiding the cost and damage of unnecessary rewriting.
A national-scale GCSE study finds that detailed learner profiles add only modest predictive value beyond overall performance, narrowing where complex personalisation is worth deploying.
A prevalence-controlled test shows why semantic similarity can overstate conceptual understanding—and how model-review teams can evaluate the difference.
AdaHome shows how selective reasoning and dedicated preference memory can improve the accuracy, efficiency, and adaptability of a locally deployed smart-home assistant.