Cover image

Beyond the Pareto Frontier: Pricing LLM Mistakes in the Real World

A practical reading of an economic LLM evaluation framework that turns accuracy, latency, abstention, and inference cost into dollar-valued deployment decisions.

July 8, 2025 · 19 min · Zelina
Cover image

Collapse to Forget: Turning Model Collapse into a Privacy Feature for LLMs

A mechanism-first look at Partial Model Collapse, a machine unlearning method that turns self-generation drift into targeted output removal for LLMs.

July 8, 2025 · 16 min · Zelina
Cover image

Mind Games: How LLMs Subtly Rewire Human Judgment

A mechanism-first reading of how LLM summaries can reframe evidence, overweight early context, hallucinate authority, and measurably shift user decisions.

July 8, 2025 · 19 min · Zelina
Cover image

Passing Humanity's Last Exam: X-Master and the Emergence of Scientific AI Agents

X-Master shows that scientific AI progress can come from inference-time orchestration, tool access, critique, rewriting, and selection—not only from larger model training.

July 8, 2025 · 16 min · Zelina
Cover image

The Phantom Menace in Your Knowledge Base

A practical reading of PhantomText, a study showing how invisible document manipulations can poison RAG systems before the model ever sees a prompt.

July 8, 2025 · 19 min · Zelina
Cover image

Backtrack to the Future: How ASTRO Teaches LLMs to Think Like Search Algorithms

ASTRO shows that reasoning gains can come from training models to recover from wrong turns, not merely from scaling models or wrapping them in external agent scaffolds.

July 7, 2025 · 18 min · Zelina
Cover image

Secret Handshakes at Scale: How LLM Agents Learn to Collude

A mechanism-first reading of how LLM market agents coordinate prices, why communication and pressure matter, and what operators should test before deploying autonomous pricing agents.

July 7, 2025 · 17 min · Zelina
Cover image

Talk is Flight: How RALLY Bridges Language and Learning in UAV Swarms

RALLY shows how UAV swarms may coordinate better when language handles semantic intent and reinforcement learning handles role credit.

July 7, 2025 · 16 min · Zelina
Cover image

From Trendlines to Transformers: DeepSupp Redefines Support Level Detection

DeepSupp turns support-level detection from static charting into a dynamic representation-learning problem, but its real value is consistency rather than trading alpha.

July 6, 2025 · 18 min · Zelina
Cover image

Ping, Probe, Prompt: Teaching AI to Troubleshoot Networks Like a Pro

A practical reading of an early AI-agent troubleshooting playground: useful less as proof of autonomy, more as a template for repeatable network-operations evaluation.

July 6, 2025 · 16 min · Zelina