Spend Verification Where Risk Is Highest
FACTOR shows how claim-level risk routing can improve factual generation while separating verifier cost from end-to-end latency.
FACTOR shows how claim-level risk routing can improve factual generation while separating verifier cost from end-to-end latency.
FlashMoE shows that making oversized MoE models fit on local hardware is only half the problem; cache misses determine whether SSD-backed inference is fast enough to use.
LongCat-Flash experiments suggest that once MoE expert scaling reaches a high-sparsity regime, some marginal capacity may be better allocated to sparse lookup memory—but only within tight architectural and serving constraints.
Self-reported LLM confidence can support ranking and routing, but only after task-specific validation shows what the score actually preserves.
A two-turn benchmark shows why model accuracy and stated confidence can miss a deployment risk: abandoning correct answers when users push back.
Medical QA stress tests show why confidence, warning language, and an abstention option must be validated before they control automated routing.
A new IRT-based evaluation shows why model confidence should be tested against task difficulty before it controls acceptance, escalation, or human review.
DeconIEP shows that benchmark contamination can be treated as a tunable evaluation control, but contamination reduction only matters when clean utility is preserved.
Fine-tuning may barely move QA accuracy while materially changing which uncertainty signals can identify the errors that remain.
A router-worker audit shows why benchmark accuracy needs a second dimension: confidence that the score reflects generalization rather than sensitivity to benchmark-related cues.