Turn Down the Work, Not Just the Clock: Governing On-Device LLM Power Across Hardware and Decoding
TL;DR for operators Lowering an edge processor’s clock can reduce power, but it also slows token generation. That leaves a local LLM runtime with a constrained control problem: keep responses fast enough, stay below thermal limits, limit energy use, and avoid degrading output quality more than the application can tolerate. PELM1 expands the control surface instead of optimizing processor frequency alone. Its runtime governor adjusts CPU and GPU frequencies together with how the model drafts and verifies tokens. On the tested LLaMA-13B workload on Jetson AGX Orin, the paper reports 29.0% to 52.4% lower energy than the evaluated comparison methods across cooling conditions, while PELM delivered the highest average generation speed in three of four conditions. ...