Cover image

Refusal Is Not a Result: Vera Tests What Agents Actually Changed

Vera reframes agent safety evaluation as reproducible software testing built around observable effects, adaptive attacks, and deterministic verification.

July 29, 2026 · 8 min · Zelina
Cover image

Name the Speaker, Then Ask the Plot: Selective Reasoning for Drama Transcripts

DramaSR-532K shows that multimodal reasoning is most valuable when acoustic speaker attribution becomes uncertain, not as a replacement for the acoustic pipeline.

July 28, 2026 · 9 min · Zelina
Cover image

Refactor, Then Run It: SemaDiff Tests Whether Behavior Actually Changed

SemaDiff shows how generated cross-version callers can turn refactoring review from a structural guess into high-precision behavioral evidence.

July 28, 2026 · 8 min · Zelina
Cover image

The Mask Is Not the Model: MMIR-TCM Makes Clinical Memory Inspectable

MMIR-TCM shows how staged perception, reporting, and retrieval can support auditable TCM assistance—without establishing autonomous prescribing.

July 28, 2026 · 8 min · Zelina
Cover image

No Runtime, No Signal: Java Energy Prediction Beyond Static Metrics

Static code metrics barely predict Java method energy use; lightweight execution timing offers a better signal, but not a production-grade energy meter.

July 27, 2026 · 9 min · Zelina
Cover image

Search the Graph, Not the Model: RSF-GLLM Separates Traversal from Generation

RSF-GLLM shows how a dedicated graph reasoner can traverse weakly related bridge entities, expose auditable evidence paths, and reduce reliance on repeated LLM calls.

July 27, 2026 · 10 min · Zelina
Cover image

The Missing Present Is a Distribution: DUPO for Delayed Control

DUPO shows why delayed controllers should evaluate actions across several plausible current states instead of trusting one reconstructed present.

July 27, 2026 · 8 min · Zelina
Cover image

Merge Before You Stop: Checkpoint Choice Depends on the Operator

A controlled model-merging study shows why expert checkpoints should be selected jointly with the merge operator rather than fixed by standalone validation loss.

July 26, 2026 · 9 min · Zelina
Cover image

Outside the Radius: Reject Unsupported Requests Before You Route Them

A multi-cluster MiniLM gate improves out-of-scope rejection, while exposing why rejection and intent classification should be evaluated separately.

July 26, 2026 · 8 min · Zelina
Cover image

Same Answer, Different Risk: Visual Semantic Entropy for VLM Review Routing

Visual Semantic Entropy detects visual instability that repeated VLM answers and joint image-text perturbations can conceal.

July 26, 2026 · 10 min · Zelina