Smaller Is Not a Latency Strategy
TL;DR for operators A team trying to move a transformer onto a phone, camera, robot, or embedded device has several levers: reduce the model, lower numerical precision, change the runtime, or move to a better accelerator. Hema Hariharan Samson’s survey of lightweight transformers1 suggests these choices cannot be evaluated independently. In the paper’s analyzed batch-size-1 setting, smaller workloads can become limited by how quickly the device moves model data rather than by raw arithmetic capacity; the paper reports roughly 60–75% hardware utilization in a favorable 15–40M-parameter range. Its broader comparisons similarly show that INT8 execution, operator fusion, optimized runtimes, and specialized accelerators change realized latency by very different amounts across devices. For an operator, the practical target is not the smallest model. It is the smallest accuracy loss that satisfies the product’s measured latency and energy budget on the actual deployment hardware. ...