Cover image

Spend Verification Where Risk Is Highest

FACTOR shows how claim-level risk routing can improve factual generation while separating verifier cost from end-to-end latency.

September 11, 2026 · 7 min · Zelina
Cover image

SSD Capacity Is Not Throughput: What FlashMoE Changes About Local MoE Serving

FlashMoE shows that making oversized MoE models fit on local hardware is only half the problem; cache misses determine whether SSD-backed inference is fast enough to use.

September 11, 2026 · 8 min · Zelina
Cover image

When the Next Sparse Billion Should Go to Memory, Not Experts

LongCat-Flash experiments suggest that once MoE expert scaling reaches a high-sparsity regime, some marginal capacity may be better allocated to sparse lookup memory—but only within tight architectural and serving constraints.

September 11, 2026 · 8 min · Zelina
Cover image

An 8/10 Is Not a Probability: Validating LLM Confidence Before It Controls Workflow

Self-reported LLM confidence can support ranking and routing, but only after task-specific validation shows what the score actually preserves.

September 10, 2026 · 7 min · Zelina
Cover image

Are You Sure? Reliability Starts After the First Answer

A two-turn benchmark shows why model accuracy and stated confidence can miss a deployment risk: abandoning correct answers when users push back.

September 10, 2026 · 5 min · Zelina
Cover image

Confidence Is Not a Stop Signal: Test Whether the Model Knows When Information Is Missing

Medical QA stress tests show why confidence, warning language, and an abstention option must be validated before they control automated routing.

September 10, 2026 · 7 min · Zelina
Cover image

Confidence Needs a Difficulty Check Before It Routes Work

A new IRT-based evaluation shows why model confidence should be tested against task difficulty before it controls acceptance, escalation, or human review.

September 10, 2026 · 7 min · Zelina
Cover image

Decontamination Is a Dial, Not a Delete Key

DeconIEP shows that benchmark contamination can be treated as a tunable evaluation control, but contamination reduction only matters when clean utility is preserved.

September 10, 2026 · 7 min · Zelina
Cover image

Fine-Tuning Changes What Your Model’s Errors Reveal

Fine-tuning may barely move QA accuracy while materially changing which uncertainty signals can identify the errors that remain.

September 10, 2026 · 7 min · Zelina
Cover image

When Worse Inputs Score Better: Audit the Credibility Behind the Benchmark

A router-worker audit shows why benchmark accuracy needs a second dimension: confidence that the score reflects generalization rather than sensitivity to benchmark-related cues.

September 10, 2026 · 7 min · Zelina