TL;DR for operators

Lowering an edge processor’s clock can reduce power, but it also slows token generation. That leaves a local LLM runtime with a constrained control problem: keep responses fast enough, stay below thermal limits, limit energy use, and avoid degrading output quality more than the application can tolerate.

PELM1 expands the control surface instead of optimizing processor frequency alone. Its runtime governor adjusts CPU and GPU frequencies together with how the model drafts and verifies tokens. On the tested LLaMA-13B workload on Jetson AGX Orin, the paper reports 29.0% to 52.4% lower energy than the evaluated comparison methods across cooling conditions, while PELM delivered the highest average generation speed in three of four conditions.

The result is most relevant to teams building sustained local inference rather than chasing minimum joules in isolation. Under the lighter 1B AGX workload, PELM sometimes consumed 0.4% to 7.3% more energy than the lowest-energy baseline because it continued to meet the 25-token/s target that several frequency-focused approaches missed. The runtime is optimizing within latency and thermal constraints, not simply selecting the lowest-power state.

Frequency reduction solves only half the problem

An on-device LLM has a straightforward power control: reduce processor voltage and frequency. Under the approximation used in the paper, dynamic power rises steeply with frequency,

$$ P \propto V^2f,\qquad V \propto f \Rightarrow P \propto f^3. $$

That makes frequency reduction attractive, but it does not remove computational work. The same model layers still have to execute, only more slowly. A governor that reacts to heat by repeatedly lowering frequency can preserve the device while degrading generation speed below the product’s acceptable response rate.

PELM changes the second side of that equation. Instead of treating the model workload as fixed, it can also change how much computation is performed for accepted tokens.

Self-speculative decoding lets a shallower portion of the same model draft several tokens and then verifies those candidates. When the drafts are accepted, the system can reduce effective computation per generated token. PELM adds another control: verification itself can stop at a selected intermediate depth rather than always traversing the full model.

That last mechanism changes the trade-off. Shallower verification saves work but may diverge from full-depth output, so verification depth becomes both a compute control and a quality-sensitive parameter. PELM is consequently not a lossless decoding accelerator in its full variable-depth configuration.

The governor controls hardware and model execution together

Every 100 ms, PELM observes a state that includes CPU and GPU power and utilization, device temperature and temperature change, decoding speed, processor frequencies, current verification depth, and speculative-decoding behavior.

A branch-DQN controller then selects across five action dimensions: CPU frequency, GPU frequency, draft exit depth, number of speculative tokens, and verification depth. The chosen configuration remains fixed for the tokens generated during that control interval.

This architecture matters because the controls interact. Lower frequency reduces instantaneous processor power but may violate the decoding-speed target. Shallower execution reduces layer computation but can affect output fidelity. Longer speculation can amortize verification when drafts are accepted, but its benefit depends on the current query and model behavior.

PELM’s controller is attempting to find a feasible combination of these variables under explicit speed and thermal constraints rather than tuning each one independently.

The implementation also has to handle a systems problem created by variable depth. If execution later resumes at a deeper layer, some deeper KV-cache state may not exist because an earlier token exited sooner. The paper’s Hidden Queue Cache stores intermediate hidden states so deeper processing can resume without recomputing every preceding layer. Reported reload latency ranges from 0.090 to 0.320 ms for tested token lengths, versus 4.195 to 20.416 ms for one decoder-layer forward pass on AGX Orin. That measurement supports HQCache as an implementation mechanism rather than a second source of the headline energy result.

The gains concentrate in moderate-to-heavy workloads

The main evaluation uses 200 prompts: 40 each from GSM8K, NQ-open, HumanEval, WMT14-DE-EN, and CNN/Daily Mail. The authors test LayerSkip-pretrained LLaMA models at 1B, 8B, and 13B sizes on Jetson AGX Orin, with the 1B model also tested on Orin Nano. Four cooling regimes vary the thermal environment.

For LLaMA-13B on AGX Orin, PELM reports energy reductions of 29.0% to 52.4% relative to the evaluated baselines. For 8B, the reported reduction is 13.4% to 32.5% while maintaining comparable generation speed. Across moderate-to-heavy workloads, performance per joule improves by as much as 45.4%.

Thermal completion gives another view of the result. PELM and zTT complete all queries in the tested cases except the most demanding 13B minimum-cooling condition. There, PELM completes 45% of queries versus 34% for zTT. That result matters because instantaneous throughput is of limited value if sustained inference drives the system into its thermal failure condition.

The light 1B workload provides the useful counterexample. PELM is not consistently the minimum-energy option: it consumes about 0.4% to 7.3% more than the lowest-energy baseline in the reported AGX comparisons. Its advantage is that it continues to satisfy the 25-token/s target where several alternatives do not.

This is evidence for constrained efficiency, not universal energy dominance.

The ablations test whether the cross-layer design is doing real work

The leave-one-component-out experiments are mechanism tests rather than separate headline claims.

Removed control Observed effect What the test supports
Variable depth Generally increases computation per accepted token and slows inference Changing model execution depth contributes efficiency beyond frequency control
Self-speculative decoding Produces substantial speed and energy degradation Speculation is a material source of reduced effective computation
PELM frequency control Can favor speed on a light workload but performs poorly on moderate-to-heavy cases Hardware power control remains necessary when workload pressure rises

Together, these results support the paper’s central architectural claim: the gains are not attributable to one unusually effective DVFS policy. The three control dimensions play complementary roles.

The controller itself is small relative to the LLM: 5,687 parameters, about 22 KB in FP32. Measured action-selection latency ranges from 0.087 to 0.922 ms, while training updates take 9 to 18 ms within the 100 ms control period. Those measurements make controller overhead a bounded implementation cost in the tested systems.

For deployment teams, the control boundary moves upward

What the paper directly shows is that, within its tested environment, allowing the runtime to manipulate both processor operating points and model execution can produce large energy gains under heavier workloads while respecting explicit throughput and thermal requirements.

Cognaptus infers a broader design consequence for local-AI products: power governance may need an interface into the model runtime rather than ending at the hardware scheduler. If intermediate exits, speculation behavior, and verification depth are exposed as controllable runtime variables, the platform can trade hardware effort against model computation instead of treating inference as an opaque fixed load.

That also makes early-exit-capable model design relevant to platform engineering. Intermediate-layer outputs are no longer only an acceleration technique; they become part of the resource-control surface. Product teams adopting such a design would need quality thresholds alongside latency, energy, and thermal objectives because variable verification depth explicitly changes the fidelity trade-off.

The evidence stops at a specific deployment envelope

PELM’s full design requires useful intermediate-layer outputs produced through early-exit supervision. The authors evaluate LayerSkip-pretrained LLaMA checkpoints; arbitrary off-the-shelf models do not automatically expose the same control mechanism. The authors describe a reduced configuration using DVFS and dynamic self-speculative decoding with full-depth verification when variable-depth operation is unavailable.

The checkpoints are also unquantized. Interactions among quantization, early exits, variable verification depth, and cached hidden states remain untested. Hardware evidence comes from Jetson AGX Orin and Orin Nano, so deployment on smartphone SoCs, NPUs, or other heterogeneous accelerators would require new mappings between the controller and device-specific power interfaces.

Finally, the evaluation does not include concurrent background workloads or interaction with a broader platform governor. Those conditions could change both available thermal headroom and the meaning of utilization signals used by the policy.

The paper nevertheless makes a concrete systems point within its tested boundary. Once the runtime can change both the processor’s operating state and the amount of model computation being requested from it, edge-LLM efficiency stops being a frequency-setting problem. It becomes a closed-loop allocation problem across speed, energy, heat, and output fidelity—and the 1B counterexample shows why those objectives cannot be collapsed into minimum energy alone.

Cognaptus: Automate the Present, Incubate the Future.


  1. Weisi Yang and Stephen Xia (2026). PELM: Power Efficient On-Device LLM Inference with Speculative Decoding and Dynamic Voltage Frequency Scaling. arXiv:2609.09662. https://arxiv.org/abs/2609.09662 ↩︎