MoEBlaze Cuts the Routing Buffers Behind the MoE Memory Wall
MoEBlaze shows that sparse MoE training can remain memory-bound unless routing buffers, activation lifetimes, and GPU data movement are redesigned together.
MoEBlaze shows that sparse MoE training can remain memory-bound unless routing buffers, activation lifetimes, and GPU data movement are redesigned together.
A new MoE scaling law suggests that teams should redesign the split between attention and expert computation as training budgets and sparsity change.
XShare shows that MoE serving efficiency depends on the expert footprint of the whole batch, turning expert budgets into a controllable inference-time throughput and quality decision.
CFM-Bench shows why wireless teams should select channel models by matched domain-task performance rather than a single leaderboard score or compute budget.
OpenEuroLLM scaling experiments show why learning rate and batch size must be recalibrated across model scale, token budget, hardware constraints, and annealing.
CoM-PT shows how sequential knowledge reuse can reduce the total cost of training a vision-model family, while turning model-size selection into a chain-design problem.
A new MoE quantization result suggests that precision should follow expert fragility, not simply expert usage.
MOSAIC shows why sparse-MoE architecture decisions should price the useful computation a real cluster can deliver, not model FLOPs alone.
A sparse MoE model may activate far fewer parameters than it owns, but routing, capacity, communication, and deployment topology determine whether that sparsity becomes real system efficiency.
New evidence suggests sparse MoE routing may make model behavior easier to inspect at the expert level, not merely cheaper to execute.