Cover image

The Generalization Stack: Why HAR Robustness Is a Pipeline Property

A 410,400-experiment HAR benchmark shows why deployment robustness depends on the full training-and-selection pipeline, not a single domain-generalization method.

September 28, 2026 · 8 min · Zelina
Cover image

The Guard Is Already in the Draft: Reusing Speculative Decoding for LLM Monitoring

Speculative Probing shows how systems already paying for speculative decoding may reuse that infrastructure for sequence-aware monitoring at much lower incremental cost.

September 28, 2026 · 8 min · Zelina
Cover image

The KL You Weren’t Watching: Moving RLVR Stability to the Query Side

ERPO shifts explicit drift control from model responses to the queries themselves, aiming to preserve exploration while improving RLVR stability.

September 28, 2026 · 8 min · Zelina
Cover image

Turn Down the Work, Not Just the Clock: Governing On-Device LLM Power Across Hardware and Decoding

PELM shows why efficient on-device LLM inference is a joint control problem spanning processor frequency, model computation, latency, thermal headroom, and output quality.

September 28, 2026 · 7 min · Zelina
Cover image

Watching a Signal the Model Can Move

Activation probes are not automatically tamper-resistant: a new benchmark shows that current LLMs can deliberately reshape some of the internal signals those monitors read.

September 28, 2026 · 9 min · Zelina
Cover image

Keep the Rollout Honest: Miles Treats Throughput and Fidelity as One System

Miles v0.1 shows why frontier post-training teams should optimize utilization, data freshness, numerical fidelity, memory placement, and weight movement as one coupled system.

September 27, 2026 · 6 min · Zelina
Cover image

Left of Whom? Spatial Agents Need More Than an Explicit Viewpoint

POVBench shows that explicitly supplying a human viewpoint does not remove the hardest parts of observer-relative spatial reasoning, changing how embodied AI teams should diagnose and remediate failures.

September 27, 2026 · 8 min · Zelina
Cover image

Local Evidence, Global Rule: When Knowledge Graph Embeddings Generalize Too Far

PogRE shows how knowledge-graph models can turn sparsely supported relational patterns into overly broad rules—and how generalization scope can instead grow with accumulated evidence.

September 27, 2026 · 7 min · Zelina
Cover image

Safety Without the Data Lake: Federating the Guard, Not the Traces

FGLGuard shows why multi-agent safety may require local adaptation even when organizations cannot centralize the sensitive traces needed to train it.

September 27, 2026 · 7 min · Zelina
Cover image

Same Count, Different Bugs: Build LLM Security Scans as an Evidence Funnel

Bugstone-E2E shows how CVE patch history can become reusable detection knowledge—and why LLM security findings need staged verification rather than one-shot trust.

September 27, 2026 · 7 min · Zelina