Cover image

Resident, Not Running: Separating LLM Memory from Compute at the Edge

TL;DR for operators An edge device can have enough compute to run an LLM yet still fail the deployment because too many weights must remain resident in memory. Shrinking the representation through quantization is one response, but it couples several engineering choices together. SelectInfer1 proposes another control surface. In its primary configuration, it loads 70% of FFN neurons but computes 40%, keeping the loaded weights at full precision. On Llama3.2-3B, reported peak memory falls from 6.88 GB to 5.97 GB. On the same Jetson Orin Nano evaluation, decoding throughput is reported at 9.85 tokens/s, compared with 6.41 tokens/s for 4-bit bitsandbytes quantization and 0.75 tokens/s for disk offloading. ...

September 6, 2026 · 8 min · Zelina
Cover image

Approximate the Ranking, Not the Answer: Prox’s Two-Stage Bet on Sparse LLM Inference

TL;DR for operators Prox1 addresses a deployment problem that appears whenever sparsity itself requires computation: how much work should an inference system spend deciding which work to skip? Its answer is unusually specific. Use a cheap, input-sparse INT4 calculation to rank likely-important feed-forward channels, not to approximate their final values. Then recompute only the selected channels with the original model weights. The distinction matters empirically: at 70% effective FFN sparsity, removing the exact recomputation stage drops the aggregate downstream score from 68.6 to 44.3 on Qwen3-8B and from 74.8 to 56.7 on Qwen3-14B. ...

August 22, 2026 · 7 min · Zelina
Cover image

The Experts Are Sparse Inside: Why MoE Cost Cuts Stop at 1.2x

The Experts Are Sparse Inside: Why MoE Cost Cuts Stop at 1.2x Cost has a way of making architecture fashionable. Mixture-of-Experts models became attractive because they promise a pleasant bargain: keep a large total parameter count, but activate only a small part of the model for each token. In business language, that sounds like capacity without the full compute bill. In engineering language, it means routing each token to a few expert feed-forward networks instead of running every expert all the time. ...

May 27, 2026 · 16 min · Zelina