The Generalization Stack: Why HAR Robustness Is a Pipeline Property
A 410,400-experiment HAR benchmark shows why deployment robustness depends on the full training-and-selection pipeline, not a single domain-generalization method.
A 410,400-experiment HAR benchmark shows why deployment robustness depends on the full training-and-selection pipeline, not a single domain-generalization method.
Speculative Probing shows how systems already paying for speculative decoding may reuse that infrastructure for sequence-aware monitoring at much lower incremental cost.
ERPO shifts explicit drift control from model responses to the queries themselves, aiming to preserve exploration while improving RLVR stability.
PELM shows why efficient on-device LLM inference is a joint control problem spanning processor frequency, model computation, latency, thermal headroom, and output quality.
Activation probes are not automatically tamper-resistant: a new benchmark shows that current LLMs can deliberately reshape some of the internal signals those monitors read.
Miles v0.1 shows why frontier post-training teams should optimize utilization, data freshness, numerical fidelity, memory placement, and weight movement as one coupled system.
POVBench shows that explicitly supplying a human viewpoint does not remove the hardest parts of observer-relative spatial reasoning, changing how embodied AI teams should diagnose and remediate failures.
PogRE shows how knowledge-graph models can turn sparsely supported relational patterns into overly broad rules—and how generalization scope can instead grow with accumulated evidence.
FGLGuard shows why multi-agent safety may require local adaptation even when organizations cannot centralize the sensitive traces needed to train it.
Bugstone-E2E shows how CVE patch history can become reusable detection knowledge—and why LLM security findings need staged verification rather than one-shot trust.