Cover image

Turn Down the Work, Not Just the Clock: Governing On-Device LLM Power Across Hardware and Decoding

TL;DR for operators Lowering an edge processor’s clock can reduce power, but it also slows token generation. That leaves a local LLM runtime with a constrained control problem: keep responses fast enough, stay below thermal limits, limit energy use, and avoid degrading output quality more than the application can tolerate. PELM1 expands the control surface instead of optimizing processor frequency alone. Its runtime governor adjusts CPU and GPU frequencies together with how the model drafts and verifies tokens. On the tested LLaMA-13B workload on Jetson AGX Orin, the paper reports 29.0% to 52.4% lower energy than the evaluated comparison methods across cooling conditions, while PELM delivered the highest average generation speed in three of four conditions. ...

September 28, 2026 · 7 min · Zelina
Cover image

Before the First Token: Put the Jailbreak Gate Inside the Model

TL;DR for operators A product team normally has several places to stop a dangerous request: filter the prompt, ask another model to inspect it, regenerate under stricter instructions, or moderate the output after generation. All of those controls act around the language model. GUARD-SLM asks whether the model’s own internal representation can provide the stop signal earlier. ...

September 21, 2026 · 7 min · Zelina
Cover image

SSD Capacity Is Not Throughput: What FlashMoE Changes About Local MoE Serving

TL;DR for operators Moving inactive MoE experts to SSD can make an oversized sparse model fit when the full checkpoint does not fit comfortably in DRAM. But every expert missing from fast memory must be fetched before decoding can continue. In FlashMoE’s evaluated system, an expert load takes roughly 3 milliseconds and expert loading accounts for more than 70% of decoding time. SSD therefore solves capacity only if the cache avoids enough expensive misses. ...

September 11, 2026 · 8 min · Zelina
Cover image

A Richer Map Can Make the Planner Slower

TL;DR for operators A robot can perceive more of its environment than its planner should necessarily receive. In the experiments summarized here, adding task-irrelevant objects to structured scene representations increases the burden on classical planners and can leave harder problems unsolved. The proposed response is not a new end-to-end planner. It is a learned relevance layer that decides which objects and relations should survive into the planning problem. ...

September 7, 2026 · 7 min · Zelina
Cover image

Resident, Not Running: Separating LLM Memory from Compute at the Edge

TL;DR for operators An edge device can have enough compute to run an LLM yet still fail the deployment because too many weights must remain resident in memory. Shrinking the representation through quantization is one response, but it couples several engineering choices together. SelectInfer1 proposes another control surface. In its primary configuration, it loads 70% of FFN neurons but computes 40%, keeping the loaded weights at full precision. On Llama3.2-3B, reported peak memory falls from 6.88 GB to 5.97 GB. On the same Jetson Orin Nano evaluation, decoding throughput is reported at 9.85 tokens/s, compared with 6.41 tokens/s for 4-bit bitsandbytes quantization and 0.75 tokens/s for disk offloading. ...

September 6, 2026 · 8 min · Zelina
Cover image

The Optimizer Comes Before the Quantizer

TL;DR for operators A small model can perform acceptably after task specialization and still lose more accuracy than expected when compressed. The paper examined here suggests that part of that loss may be shaped earlier, during fine-tuning, rather than being determined by the quantizer alone. Paper evidence: Muon-optimized models lost less accuracy after quantization than Adam-optimized models on six of eight benchmarks. On ARC-e, for example, accuracy fell by 0.55 percentage points after quantization with Muon versus 3.16 points with Adam. The integrated pipeline also produced a 2.86 GB quantized model from a 6.01 GB pre-quantized model while improving reported serving throughput and roughly halving per-token latency. ...

September 6, 2026 · 8 min · Zelina
Cover image

Smaller Is Not a Latency Strategy

TL;DR for operators A team trying to move a transformer onto a phone, camera, robot, or embedded device has several levers: reduce the model, lower numerical precision, change the runtime, or move to a better accelerator. Hema Hariharan Samson’s survey of lightweight transformers1 suggests these choices cannot be evaluated independently. In the paper’s analyzed batch-size-1 setting, smaller workloads can become limited by how quickly the device moves model data rather than by raw arithmetic capacity; the paper reports roughly 60–75% hardware utilization in a favorable 15–40M-parameter range. Its broader comparisons similarly show that INT8 execution, operator fusion, optimized runtimes, and specialized accelerators change realized latency by very different amounts across devices. For an operator, the practical target is not the smallest model. It is the smallest accuracy loss that satisfies the product’s measured latency and energy budget on the actual deployment hardware. ...

September 5, 2026 · 7 min · Zelina
Cover image

The Front End Holds the Line: What Heart-Sound CNNs Lose When Models Shrink

TL;DR for operators A team shrinking an audio classifier for a low-cost screening device has more than one place to spend scarce compute. The network can stay larger, or the input representation can do more work before the signal reaches the network. That trade-off became visible when this heart-sound CNN was reduced from three convolutional blocks to two. Using a plain log-mel spectrogram, modified accuracy—a score that gives equal weight to abnormal-case sensitivity and normal-case specificity—fell from 0.910 to 0.826. With a front-end that normalizes each frequency band against its recent energy so locally unusual sounds stand out more clearly, called PCEN, it fell only from 0.915 to 0.894. A front-end representing the same sound at several time-frequency resolutions, called multi-resolution log-mel, similarly fell from 0.916 to 0.894. ...

August 13, 2026 · 7 min · Zelina
Cover image

Reasoning on Demand: AdaHome’s Case for Tiered Local Assistants

TL;DR for operators A household assistant should not spend the same computational effort on “turn on the light” as on “make the room comfortable.” It also should not treat one unusual request as a permanent preference change. AdaHome applies that principle through tiered local processing. Explicit commands take a short planning path, while requests that need interpretation or personal context receive additional reasoning, validation, and—where appropriate—user confirmation. Under a common Llama 3.2-3B setup, it achieved 86.7% success on both direct and indirect commands while recording the lowest latency and token use in every command category among the compared systems. ...

August 8, 2026 · 7 min · Zelina
Cover image

The Training Rhythm Survives Encryption

TL;DR for operators A federated-learning client repeatedly downloads a model, computes locally, and uploads an update while the mobile network assigns radio resources to each step. Encryption hides the transmitted contents, but not the timing, direction, allocation size, and recurring cadence created by this training cycle. FLINT1 reconstructs those scheduling traces and uses them to classify CNN, RNN, and Transformer families. With a 300-second observation window, it reaches a macro F1 of 0.930 ± 0.021 in the evaluated closed-world testbed. ...

August 5, 2026 · 8 min · Zelina