Sparse Per Token, Dense Per Batch: XShare Reprices MoE Routing at Inference Time
XShare shows that MoE serving efficiency depends on the expert footprint of the whole batch, turning expert budgets into a controllable inference-time throughput and quality decision.