Cover image

Probe Before You Prune: Measuring Which MoE Experts a Workload Can Lose

Lightweight adaptation can serve as a diagnostic probe for MoE expert pruning, turning a short task-relevant run into a practical signal for trading model capacity against memory and latency.

September 6, 2026 · 9 min · Zelina
Cover image

Resident, Not Running: Separating LLM Memory from Compute at the Edge

SelectInfer shows how edge LLM deployment can treat memory residency and runtime computation as separate budgets rather than compressing both through one fixed setting.

September 6, 2026 · 8 min · Zelina
Cover image

The Optimizer Comes Before the Quantizer

A study of Llama 3.2 3B suggests that the fine-tuning optimizer can affect how much task accuracy survives later 4-bit quantization, making compression an end-to-end deployment decision.

September 6, 2026 · 8 min · Zelina
Cover image

The Planner Trusted the Wrong State: A New Security Boundary for Embodied Agents

State-semantic injection shows why embodied-agent security must protect the internal world representation between perception and planning, not only prompts and models.

September 6, 2026 · 7 min · Zelina
Cover image

When Confidence Drives the Workflow: What HypeLoRA Changes About Adapter Selection

HypeLoRA shows why confidence-sensitive classifiers should be selected on calibration and task performance separately, not on accuracy alone.

September 6, 2026 · 7 min · Zelina
Cover image

When Model Output Can Change State: An Architecture Guide to Agent Reliability

A systems view of Agentic AI shows where autonomy creates operational risk—and where architecture can contain it.

September 6, 2026 · 7 min · Zelina
Cover image

Cache Hits Can Create Queues

A new KV-serving design shows why cache eviction and request routing should be controlled together when prefix reuse competes with worker congestion.

September 5, 2026 · 7 min · Zelina
Cover image

Compress the Activation, Spend the Memory on Rank

CARE-LoRA shows how LoRA fine-tuning can reclaim activation memory without freezing its projection-down matrix, then reinvest that memory in adapter capacity.

September 5, 2026 · 7 min · Zelina
Cover image

LoRA’s Missing Budget: Which Matrices Deserve an Adapter?

Condition-number-guided LoRA selection shifts fine-tuning efficiency from shrinking every adapter to deciding which pretrained matrices deserve adaptation at all.

September 5, 2026 · 7 min · Zelina
Cover image

Shrink the KV Cache, Miss the Bottleneck

A serving team should choose KV-cache optimizations by the resource constraint they relieve, not by how much memory they remove.

September 5, 2026 · 8 min · Zelina