Give the Quiet Experts More Bits
TL;DR for operators A fixed MoE memory budget does not tell you which experts can safely absorb the lowest precision. Efficient Quantization of Mixture-of-Experts with Theoretical Generalization Guarantees1 argues that the experts most frequently or strongly used are not necessarily the ones that need the most bits. In its theory, experts associated with less-prevalent but still task-relevant features develop weaker activations and smaller router-norm changes, leaving less margin for quantization error. ...