Cover image

Before You Spend 10 Trillion Tokens: Separate Width From Horizon

TL;DR for operators A training team preparing a multi-trillion-token MoE run usually cannot afford to test several full-scale learning rates. Kim et al. show a way to reduce that search before the expensive run begins: transfer the learning-rate optimum across model width, then estimate separately how that optimum moves as the token budget grows.1 ...

September 23, 2026 · 7 min · Zelina
Cover image

MoEBlaze Cuts the Routing Buffers Behind the MoE Memory Wall

TL;DR for operators Sparse expert activation does not guarantee a small GPU-memory footprint. Conventional token routing can physically duplicate token representations into expert-specific storage, while expert FFNs preserve additional intermediates for backward computation. In the paper’s illustrative DeepSeek-like case, one materialized routing buffer alone is estimated at roughly 94 GB. MoEBlaze1 attacks that data movement rather than changing the router or MoE architecture. It keeps compact token-to-expert location information, gathers original activations when expert computation needs them, reduces results directly back into token order, builds the required routing metadata without global sorting, and combines fused SwiGLU kernels with selective recomputation. ...

September 12, 2026 · 8 min · Zelina
Cover image

Scale the Split: MoE Compute Allocation Should Move With the Budget

TL;DR for operators When a sparse Mixture-of-Experts model receives a larger training budget, the additional computation has at least two architectural destinations: attention capacity and expert feed-forward capacity. Treating their existing balance as fixed can leave the larger model internally misallocated. Li et al. show this experimentally in GPT-style sparse MoE models.1 At fixed compute and sparsity, varying the expert-versus-attention allocation produces a clear loss minimum. As total compute increases, the loss-minimizing allocation shifts toward more expert computation, but the rate of that shift depends on sparsity. ...

September 12, 2026 · 7 min · Zelina
Cover image

Sparse Per Token, Dense Per Batch: XShare Reprices MoE Routing at Inference Time

TL;DR for operators Sparse MoE routing does not guarantee a sparse serving batch. As more tokens are decoded together, especially under speculative decoding, their individual top-k choices can collectively touch a large fraction of the model’s experts. That increases expert-weight movement during a memory-bound stage of inference. XShare treats this as a serving-policy problem rather than a model-retraining problem. It ranks experts using router scores from the current batch, limits the shared expert pool, and then lets each token perform its normal top-k refinement within that pool. In the reported GPT-OSS 120B experiments, moderate restrictions improve output-token throughput by roughly 6-13% in standard decoding and around 13-14% in several speculative settings while keeping accuracy comparatively close to baseline. More aggressive restriction can go much faster, but the quality loss can be large. ...

September 12, 2026 · 7 min · Zelina
Cover image

Give the Quiet Experts More Bits

TL;DR for operators A fixed MoE memory budget does not tell you which experts can safely absorb the lowest precision. Efficient Quantization of Mixture-of-Experts with Theoretical Generalization Guarantees1 argues that the experts most frequently or strongly used are not necessarily the ones that need the most bits. In its theory, experts associated with less-prevalent but still task-relevant features develop weaker activations and smaller router-norm changes, leaving less margin for quantization error. ...

September 11, 2026 · 6 min · Zelina
Cover image

More FLOPs, Worse Choice: Sparse MoE Scaling Has to Price the Cluster

TL;DR for operators Suppose a team has already bought the cluster time: 32 B200 nodes for 20 days. It might seem sensible to prefer the sparse-MoE design that manages to execute the most model FLOPs during that window. In the paper’s reported search, that rule would pick the wrong configuration. The loss-optimal candidate uses about $1.23\times10^{23}$ model FLOPs—roughly 36% fewer than the candidate that realizes about $1.94\times10^{23}$—yet reaches the lower predicted loss. ...

September 11, 2026 · 7 min · Zelina
Cover image

Sparse Is Not Cheap by Default: What MoE Efficiency Actually Depends On

TL;DR for operators Sparse mixture-of-experts models can own hundreds of billions—or even a trillion—parameters while activating only a much smaller subset for each token. That makes activated parameters a more relevant starting point than total parameters when comparing computational burden. It does not settle the infrastructure question. The surveyed literature shows that tokens still have to be assigned to experts, overloaded experts need capacity controls, unused capacity can create padding, overflow may result in dropped tokens, and experts distributed across devices require substantial network traffic. Dong Pan and colleagues’ survey brings these model-level and systems-level constraints into one view.1 ...

September 11, 2026 · 8 min · Zelina
Cover image

Sparse Routing May Buy You Inspectability, Not Just Efficiency

TL;DR for operators Herbst, Wermter, and Lee find that the analyzed Mixture-of-Experts models often represent tested concepts in far fewer neurons than comparable dense transformers.1 The difference is largest under the hardest probe constraint: when only one neuron is available, MoE experts often approach their own best probe performance while dense feed-forward layers need more dimensions. Models with sparser routing also tend to show cleaner representations. ...

September 11, 2026 · 8 min · Zelina
Cover image

SSD Capacity Is Not Throughput: What FlashMoE Changes About Local MoE Serving

TL;DR for operators Moving inactive MoE experts to SSD can make an oversized sparse model fit when the full checkpoint does not fit comfortably in DRAM. But every expert missing from fast memory must be fetched before decoding can continue. In FlashMoE’s evaluated system, an expert load takes roughly 3 milliseconds and expert loading accounts for more than 70% of decoding time. SSD therefore solves capacity only if the cache avoids enough expensive misses. ...

September 11, 2026 · 8 min · Zelina
Cover image

When the Next Sparse Billion Should Go to Memory, Not Experts

TL;DR for operators A sparse-model team with a fixed marginal parameter budget should not assume that the next increment belongs in more experts. In the LongCat-Flash experiments reported by Liu et al.,1 parameter-equivalent expert scaling performs better earlier, but N-gram embedding scaling takes the lead after the MoE reaches a sufficiently high-sparsity regime. ...

September 11, 2026 · 8 min · Zelina