Cover image

Outside the Radius: Reject Unsupported Requests Before You Route Them

A multi-cluster MiniLM gate improves out-of-scope rejection, while exposing why rejection and intent classification should be evaluated separately.

July 26, 2026 · 8 min · Zelina
Cover image

Same Answer, Different Risk: Visual Semantic Entropy for VLM Review Routing

Visual Semantic Entropy detects visual instability that repeated VLM answers and joint image-text perturbations can conceal.

July 26, 2026 · 10 min · Zelina
Cover image

Control the Crowd Before It Arrives: What PedNStream Can—and Cannot—Test

PedNStream offers fast, network-scale testing of crowd interventions, but its operational value depends on disciplined calibration and validation.

July 25, 2026 · 8 min · Zelina
Cover image

Coverage Is Not a Range: RPS for Ordinal Conformal Prediction

RPS conformal prediction turns class probabilities into contiguous ordinal ranges that balance interval width against the severity of uncovered outcomes.

July 25, 2026 · 9 min · Zelina
Cover image

Slow Policy, Fast Power: Where Agentic Control Belongs in Wireless Networks

Agentic-LTPO shows how agents can adapt wireless objectives without placing probabilistic language models inside the real-time beamforming loop.

July 25, 2026 · 9 min · Zelina
Cover image

Agree Once, Remember Later: The Commit Boundary in Personal Agents

A benchmark of stateful personal agents shows that the largest sycophancy risk emerges when user claims are written into durable state and reused later.

July 24, 2026 · 8 min · Zelina
Cover image

Approved One by One, Risky in Combination: Testing Agent Skills Before They Execute

SkillFuzz shows how marketplace operators can screen composition-induced agent risks before committing scarce sandbox and review capacity.

July 24, 2026 · 9 min · Zelina
Cover image

Rotate Before You Update: Deeper Test-Time Adaptation for Graph Models

T3R shows how graph models can adapt deeper layers from unlabeled test data, provided teams qualify the auxiliary signal, bound the update budget, and accept the latency cost.

July 24, 2026 · 8 min · Zelina
Cover image

Average at Your Own Risk: The Metric Setting That Can Reverse the Winner

A practical guide to choosing micro, macro, weighted, and exemplar aggregation according to the operational unit a classifier must serve.

July 23, 2026 · 8 min · Zelina
Cover image

Look Again Before You Answer: Visual RAG Needs a Search Policy

ProMSA shows that visual retrieval improves when models learn to switch search modality, retry failed matches, and stop under explicit budgets.

July 23, 2026 · 8 min · Zelina