Cover image

Smaller Is Not a Latency Strategy

TL;DR for operators A team trying to move a transformer onto a phone, camera, robot, or embedded device has several levers: reduce the model, lower numerical precision, change the runtime, or move to a better accelerator. Hema Hariharan Samson’s survey of lightweight transformers1 suggests these choices cannot be evaluated independently. In the paper’s analyzed batch-size-1 setting, smaller workloads can become limited by how quickly the device moves model data rather than by raw arithmetic capacity; the paper reports roughly 60–75% hardware utilization in a favorable 15–40M-parameter range. Its broader comparisons similarly show that INT8 execution, operator fusion, optimized runtimes, and specialized accelerators change realized latency by very different amounts across devices. For an operator, the practical target is not the smallest model. It is the smallest accuracy loss that satisfies the product’s measured latency and energy budget on the actual deployment hardware. ...

September 5, 2026 · 7 min · Zelina
Cover image

Not Every Spike Is Positive: The MTJ Neuron Built for Signed Signals

TL;DR for operators The paper proposes a magnetic tunnel junction, or MTJ, neuron that can implement signed leaky integrate-and-fire dynamics: positive and negative spikes, not merely ordinary one-direction spiking dressed up in new device terminology.1 The important move is geometric. The authors align the pinned-layer easy axis with the short axis of an elliptical free layer, while the free layer’s own easy axis points along the height direction. That orthogonal-easy-axis arrangement changes how the free-layer magnetization accumulates, relaxes, and crosses thresholds. In business language, the paper is not saying “spintronics is cool.” It is saying “a particular magnetic geometry may give a compact physical substrate for richer spiking representations.” Subtle difference. Useful difference. ...

June 20, 2026 · 15 min · Zelina