Cover image

When Images Pretend to Be Interfaces: Stress‑Testing Generative Models as GUI Environments

GEBench shows why beautiful generated interfaces are not yet reliable environments for training or testing GUI agents.

February 9, 2026 · 14 min · Zelina
Cover image

When Privacy Meets Chaos: Making Federated Learning Behave

A careful reading of FedCompDP shows why privacy, client heterogeneity, and aggregation stability must be designed together—not bolted together after the model starts shaking.

February 9, 2026 · 15 min · Zelina
Cover image

CompactRAG: When Multi-Hop Reasoning Stops Burning Tokens

CompactRAG shows how multi-hop RAG can shift cost from repeated online LLM calls to reusable offline knowledge compaction.

February 8, 2026 · 16 min · Zelina
Cover image

Freeze Now, Learn Faster: When Parameter Freezing Meets Pipeline Reality

TimelyFreeze shows that parameter freezing only becomes a real training-speed lever when it is aligned with the pipeline schedule’s wall-clock bottlenecks.

February 8, 2026 · 19 min · Zelina
Cover image

Learning to Inject: When Prompt Injection Becomes an Optimization Problem

AutoInject shows why prompt injection should be tested as an adaptive optimization problem, not merely as a list of hand-written attack templates.

February 8, 2026 · 17 min · Zelina
Cover image

Speculation, But With Standards: Training Draft Models That Actually Get Accepted

VSD shows why speculative decoding improves when draft models are trained for accepted paths, not merely probable tokens.

February 8, 2026 · 13 min · Zelina
Cover image

Tokens, Watts, and Waste: The Hidden Energy Bill of LLM Inference

A mechanism-first reading of why LLM inference energy is shaped by prefill, decoding, prompt length, and unnecessary generation—not merely model size.

February 8, 2026 · 14 min · Zelina
Cover image

Ultra‑Sparse Embeddings Without Apology

CSRv2 shows that ultra-sparse embeddings fail less because sparsity is impossible, and more because we have been training them badly.

February 8, 2026 · 19 min · Zelina
Cover image

When Words Start Walking: Rethinking Semantic Search Beyond Averages

A comparison-based reading of why Word Mover’s Distance with GloVe outperforms centroid-style semantic search in statement-level retrieval, and where that lesson actually applies in business systems.

February 8, 2026 · 15 min · Zelina
Cover image

Benchmarks Lie, Rooms Don’t: Why Embodied AI Fails the Moment It Enters Your House

A mechanism-first reading of TEA, an in-situ task-generation framework showing why embodied AI needs environment-specific evaluation before deployment.

February 7, 2026 · 17 min · Zelina