Cover image

Talking to Yourself, but Make It Useful: Intrinsic Self‑Critique in LLM Planning

A procedural self-critique loop can make LLM planners markedly more reliable—but only when reflection is converted into explicit rule checking, state tracking, and conservative approval.

January 3, 2026 · 17 min · Zelina
Cover image

Think First, Grasp Later: Why Robots Need Reasoning Benchmarks

ERIQ and GenieReasoner reveal why understanding the right action and physically executing it are separate engineering problems that robotics teams must diagnose separately.

January 3, 2026 · 17 min · Zelina
Cover image

When Models Start to Forget: The Hidden Cost of Training LLMs Too Well

A practical reading of why LLM memorization becomes hard to remove once training entangles recall with general capability.

January 3, 2026 · 16 min · Zelina
Cover image

When Three Examples Beat a Thousand GPUs

A controlled study of LLM-generated neural networks shows why moderate prompt context can improve architecture synthesis—and why more examples eventually break the pipeline.

January 3, 2026 · 15 min · Zelina
Cover image

Big AI and the Metacrisis: When Scaling Becomes a Liability

A systems-level reading of how AI scale can amplify ecological, social, linguistic, and institutional risks—and what organizations can do about it.

January 2, 2026 · 18 min · Zelina
Cover image

Ethics Isn’t a Footnote: Teaching NLP Responsibility the Hard Way

A four-year experiment in hands-on NLP ethics shows why responsibility is learned through difficult choices, public explanation, and repeated practice—not compliance slides.

January 2, 2026 · 16 min · Zelina
Cover image

LeanCat-astrophe: Why Category Theory Is Where LLM Provers Go to Struggle

LeanCat reveals why verified AI reasoning still fails when agents must navigate large libraries, preserve abstraction, and construct missing conceptual bridges.

January 2, 2026 · 17 min · Zelina
Cover image

MI-ZO: Teaching Vision-Language Models Where to Look

MI-ZO shows how a lightweight inference-time controller can improve 2D-trained vision-language models on 3D scenes by learning which views contain useful, non-redundant evidence.

January 2, 2026 · 16 min · Zelina
Cover image

Planning Before Picking: When Slate Recommendation Learns to Think

HiGR shows that generative recommendation becomes practical only when item representation, slate planning, and preference alignment are designed as one coordinated system.

January 2, 2026 · 18 min · Zelina
Cover image

Question Banks Are Dead. Long Live Encyclo-K.

Encyclo-K replaces fixed benchmark questions with dynamically composed knowledge statements, creating a reusable evaluation engine that exposes the gap between knowing facts and reliably combining them.

January 2, 2026 · 14 min · Zelina