Cover image

LeanCat-astrophe: Why Category Theory Is Where LLM Provers Go to Struggle

LeanCat reveals why verified AI reasoning still fails when agents must navigate large libraries, preserve abstraction, and construct missing conceptual bridges.

January 2, 2026 · 17 min · Zelina
Cover image

MI-ZO: Teaching Vision-Language Models Where to Look

MI-ZO shows how a lightweight inference-time controller can improve 2D-trained vision-language models on 3D scenes by learning which views contain useful, non-redundant evidence.

January 2, 2026 · 16 min · Zelina
Cover image

Planning Before Picking: When Slate Recommendation Learns to Think

HiGR shows that generative recommendation becomes practical only when item representation, slate planning, and preference alignment are designed as one coordinated system.

January 2, 2026 · 18 min · Zelina
Cover image

Question Banks Are Dead. Long Live Encyclo-K.

Encyclo-K replaces fixed benchmark questions with dynamically composed knowledge statements, creating a reusable evaluation engine that exposes the gap between knowing facts and reliably combining them.

January 2, 2026 · 14 min · Zelina
Cover image

Secrets, Context, and the RAG Illusion

PrivacyBench reveals why personalized RAG assistants can recognize secrets yet still expose them—and why reliable privacy controls must begin before retrieval.

January 2, 2026 · 14 min · Zelina
Cover image

Deployed, Retrained, Repeated: When LLMs Learn From Being Used

How selective reuse of validated deployment traces can quietly turn ordinary supervised fine-tuning into an implicit reinforcement-learning loop.

January 1, 2026 · 18 min · Zelina
Cover image

Gen Z, But Make It Statistical: Teaching LLMs to Listen to Data

GenZ reverses the usual LLM feature-discovery workflow by letting proprietary data identify useful distinctions before asking a foundation model to explain them.

January 1, 2026 · 17 min · Zelina
Cover image

Label Now, Drive Later: Why Autonomous Driving Needs Fewer Clicks, Not Smarter Annotators

A practical reading of the Correction Acceleration Ratio, which exposes why the most accurate 3D detector is not always the cheapest annotation assistant.

January 1, 2026 · 14 min · Zelina
Cover image

Learning the Rules by Breaking Them: Exception-Aware Constraint Mining for Care Scheduling

Historical schedules contain both operating rules and emergency compromises; this paper shows how to extract the former without institutionalizing the latter.

January 1, 2026 · 15 min · Zelina
Cover image

Let It Flow: ROME and the Economics of Agentic Craft

ROME shows that competitive agent performance depends less on possessing the largest model than on operating a disciplined learning loop around execution, verification, training, and control.

January 1, 2026 · 19 min · Zelina