Where to Go Deeper Beyond This Academy

A curated guide to textbooks, authors, websites, and papers for readers who want to study transformer internals, attention math, fine-tuning, GPU optimization, and benchmarking in more depth.

April 23, 2026 · 9 min · Michelle
Cover image

Left of Whom? Spatial Agents Need More Than an Explicit Viewpoint

TL;DR for operators An embodied assistant may already have seen a room and still fail when asked to place something “to the left of the chair” from a person’s point of view. Supplying more information about where that person is helps less than many teams might expect, because identifying the observer is only one part of the reference-frame problem. ...

September 27, 2026 · 8 min · Zelina
Cover image

New Signer, Old Sentence: What Sign-Language Translation Benchmarks Are Actually Testing

TL;DR for operators A held-out test set does not necessarily tell you how a sign-language translation system will behave with a new user. In Rethinking Sign Language Translation: The Impact of Signer Dependence on Model Evaluation, Artiaga and colleagues re-evaluate three gloss-free systems after excluding the test signer from training and development.1 ...

September 24, 2026 · 7 min · Zelina
Cover image

A Proof Can Pass and Still Mean the Wrong Thing: What AxQM Changes About Formal AI Evaluation

TL;DR for operators When an AI system produces a formal proof, an evaluation team can make one part of the verdict unusually objective: either the proof is accepted under the allowed rules, or it is not. That removes much of the grader variance found in rubric scoring or LLM judging. AxQM provides 1,019 Lean 4 proof-synthesis tasks over 479 finite-dimensional quantum-mechanics textbook items. It keeps a private reference solution for every task and grades submissions through successful compilation, absence of sorry in the proof or its dependencies, and absence of newly introduced axioms.1 ...

September 22, 2026 · 6 min · Zelina
Cover image

Reasoning Labels Don’t Travel: What UrduBench Changes About Model Selection

TL;DR for operators UrduBench1 tests 23 open and open-weight models on 2,390 held-out Urdu questions spanning arithmetic reasoning, formal mathematics, commonsense, and knowledge tasks. The ranking gives little support to selecting an Urdu model from parameter count or a “reasoning” label alone: Gemma-3-12B-it leads the reported aggregate at 59.4%, while the larger reasoning-oriented DeepSeek-R1-Distill-Qwen-14B scores 44.9%. ...

September 18, 2026 · 7 min · Zelina
Cover image

Think Harder, See Less? What Visual Illusions Reveal About Multimodal Reasoning

TL;DR for operators Giving a vision-language model more time to reason sounds like a reliability upgrade. On deceptive visual tasks, that assumption is too broad. Across seven model configurations tested with and without additional deliberation, longer reasoning improved free-form explanations of why an illusion occurs, yet detection accuracy fell for several Qwen models and multiple-choice reasoning fell for several open-source systems. ...

September 14, 2026 · 7 min · Zelina
Cover image

Correct on the Frame, Wrong on the Timeline

TL;DR for operators A video system can produce the right answer without reliably understanding what happened over time. That becomes a procurement and QA problem when the decision depends on whether one event happened before another, how long something lasted, which direction it moved, whether an action repeated, or whether one event depended on another. ...

September 13, 2026 · 7 min · Zelina
Cover image

Four Inputs In, One Modality Out: Testing Whether Omnimodal Models Actually Arbitrate Evidence

TL;DR for operators A model can receive a camera feed, spoken report, reference image, and text record without meaningfully reasoning across all four. In C$^3$PO1, 86–95% of observed failures across ten models in the paper’s failure analysis were classified as dominance-driven: one modality or prior drove the answer while other evidence was effectively ignored. ...

September 13, 2026 · 7 min · Zelina
Cover image

When Worse Inputs Score Better: Audit the Credibility Behind the Benchmark

TL;DR for operators A benchmark score can be high without being equally trustworthy as a measure of generalization. In this study, the researchers deliberately degraded benchmark questions before they reached the answering model. At a noisy-router count of eight, 10 of the 12 evaluated models nevertheless scored above their own clean baseline. At nine routers, eight models still did so, and the mean positive excess among above-baseline cases reached 0.086. ...

September 10, 2026 · 7 min · Zelina
Cover image

Reasoning Under a Running Clock: Why Agent Rankings Reverse in Real Time

TL;DR for operators A better plan can become a worse agent when the environment keeps moving while the system thinks. STAR makes that reversal unusually clear: Kimi-K2-Thinking leads the unlimited-deliberation evaluation with a rating of 1206.1 and a 1.00 win rate, then falls to 842.6 and a 0.210 win rate in real-time play, where GLM-4.6 leads at 1180.8.1 ...

September 7, 2026 · 7 min · Zelina