When the Right Answer Is No Answer: Teaching AI to Refuse Messy Math
MathDoc shows why document AI needs calibrated refusal, not just better transcription, when real exam papers are noisy, occluded, and incomplete.
MathDoc shows why document AI needs calibrated refusal, not just better transcription, when real exam papers are noisy, occluded, and incomplete.
A mechanism-first reading of LIBERTy, a structural-counterfactual benchmark that tests whether concept-based explanations actually track causal model behavior rather than merely producing plausible edits.
GUI-Eyes shows why GUI agents need learned active perception, not just bigger models staring harder at screenshots.
MatchTIR shows why multi-turn tool agents need fine-grained credit assignment, not just bigger models or louder final-answer rewards.
A mechanism-first look at PCN-Rec, a proof-carrying architecture that turns LLM recommenders from trusted decision-makers into auditable proposers.
A mechanism-first reading of why transformer scaling laws can survive even when the data itself has no power-law structure.
A business-facing reading of AI existential risk as a portfolio of survival assumptions, not one melodramatic prediction.
STITCH shows why long-horizon agents need memory indexed by task intent, not just larger context windows or better embeddings.
A practical reading of Context Bubble construction: why enterprise RAG needs constrained, auditable context assembly rather than larger top-k piles.
A mechanism-first reading of experimental evidence showing why GenAI helps novice architectural designers, fails to broadly lift performance, and can quietly weaken creative agency.