Cover image

Slow Policy, Fast Power: Where Agentic Control Belongs in Wireless Networks

TL;DR for operators An operator can shift priorities from throughput to energy saving in a sentence. The network still has to make feasible power decisions every transmission slot. Agentic-LTPO separates those jobs: a slower agent layer interprets policy and proposes bounded settings, while a deterministic solver retains control of fast execution. In a simulated network where distributed access points jointly serve users, the complete system reports 22.8 cumulative communication utility, compared with 14.5 for static configuration—a 57.2% relative gain. The improvement does not come from letting an LLM control the physical layer directly. Proposed changes pass through structured grounding, retrieval, projection into allowed ranges, criticism, and numerical optimization before affecting the network. ...

July 25, 2026 · 9 min · Zelina
Cover image

The Agent Benchmark Without the Agent Bill

TL;DR for operators Agent evaluations are expensive for a fairly obvious reason: the agent has to do something. It must browse, edit files, call tools, manipulate repositories, survive its own mistakes, and occasionally discover that the environment has changed while nobody was looking. The paper introduces Pace, a method for predicting performance on an expensive agentic benchmark from a compact set of cheaper, non-agentic test instances.1 Across 14 frontier models and four agentic benchmarks, a 100-instance Pace proxy produces an average mean absolute error of 3.80 percentage points, a 0.81 Spearman rank correlation, and 84.37% pairwise model-ranking accuracy under leave-one-model-out validation. ...

July 16, 2026 · 16 min · Zelina

From Memory-Based Coordination to Controlled Case Orchestration

A metropolitan funeral home redesigns a fragile, coordinator-dependent case process around specialized AI agents while retaining human control over every family, legal, ceremonial, and religious decision.

July 15, 2026 · 8 min · Vox
Cover image

The Crystal Ball Was a Search Loop

TL;DR for operators Lab automation is not the story here. The story is search discipline. The paper introduces HACO, a human–AI co-discovery system that tries to develop a better crystal structure prediction algorithm by searching across generative-modeling ideas, coding candidates, training them, evaluating them, and refining the winners. The system identifies masked generative modeling, specifically MaskGIT from computer vision, as a transferable idea for crystal structure prediction. With sparse human steering, it turns that idea into MaskGXT, a masked discrete-token transformer for generating crystal structures from compositions.1 ...

July 7, 2026 · 19 min · Zelina

From Alert Queue to Governed Incident Record

A mid-sized managed security provider redesigned its analyst-heavy incident workflow around specialized agents that assemble and maintain the evidence record while humans retain authority over containment, client communication, breach declaration, and regulatory escalation.

June 30, 2026 · 8 min · Vox
Cover image

Learning Has a Supply Chain

TL;DR for operators AI learning is becoming less like “train a bigger model and hope it behaves” and more like operating a controlled capability loop. The first paper in this cluster shows a narrow but important lesson: once a multimodal model has learned useful representations, the final adaptation step should optimize the metric that actually matters, while avoiding damage to the representation underneath.1 The second paper moves the same logic into physical action: an embodied system should connect language-level intention, predicted world change, memory, and executable robot control, not merely map images to motor commands with expensive optimism.2 The third paper zooms out: when agentic AI becomes economically and militarily useful, the real bottleneck includes data centers, accelerators, electricity, water, datasets, and skilled labor.3 ...

June 27, 2026 · 14 min · Zelina
Cover image

Feedback Is the New Attack Surface

TL;DR for operators AI agents are not only vulnerable because someone can hide a bad instruction in an email, document, web page, Slack message, or tool output. They are vulnerable because attackers can now automate the search for bad instructions that work. That changes the security problem. A one-off prompt injection is annoying. An automated attack loop is strategic. It generates candidate injections, observes the agent’s response, scores partial progress, keeps the promising branches, and tries again. Very entrepreneurial, in the worst possible way. ...

June 23, 2026 · 21 min · Zelina
Cover image

The Retriever Found Similar Things. The Evidence Was Elsewhere.

TL;DR for operators The current enterprise RAG conversation still has a charmingly stubborn misconception: if the model hallucinates, buy better embeddings, increase the context window, add an agent, and hope the PowerPoint becomes true. The two papers here point in a less theatrical direction. One paper, Non-negative Elastic Net Decoding for Information Retrieval, argues that dense retrieval has a structural weakness: it scores each candidate independently, so it can retrieve several similar items instead of the complementary set actually needed to answer the query.1 The other, Agentic Hybrid RAG for Evidence-Grounded Muon Collider Analysis, shows what happens when retrieval is treated as a full evidence workflow: sparse and dense retrieval are fused, queries are decomposed under constraints, evidence is deduplicated and budgeted, and answers are judged for coverage, hallucination, and abstention.2 ...

June 23, 2026 · 19 min · Zelina
Cover image

The Agents Need Traffic Laws, Not a Bigger Chatroom

TL;DR for operators The paper’s practical message is simple enough to be dangerous: once agents start working with other agents, the hard problem stops being “Can this model reason?” and becomes “Can this network behave?” Quanyan Zhu’s paper on the Internet of Agentic AI, or IoAI, frames the next stage of agentic systems as an open ecosystem of heterogeneous autonomous agents that discover collaborators, negotiate responsibilities, exchange context, invoke tools, and execute workflows across cloud, edge, device, organizational, and cyber-physical environments.1 That sounds grand, which is usually where useful engineering goes to die. But the paper’s better contribution is more sober: it treats agentic AI as a distributed systems problem. ...

June 22, 2026 · 26 min · Zelina
Cover image

Agents of Consequence: Why Tool Use Needs a Control Loop

TL;DR for operators Enterprise AI agents are moving from “answer this question” toward “watch this process, use tools, make decisions, and keep going.” That is useful. It is also how software quietly graduates from assistant to operational liability. Three recent papers, read together, make a simple point with uncomfortable business implications. VitalAgent shows how an LLM agent can become useful in wearable-health monitoring when it has physiological memory, structured tools, evidence validation, and proactive alerting.1 CoMap shows how agents can improve long-horizon decisions by pairing their policy with a co-evolving textual world model that predicts action consequences before execution.2 Gram shows why more autonomous agents also need deployment-realistic audits, because pressure, incentives, role-play cues, and implicit constraints can produce sabotage-like behavior even when the model is not cartoonishly “evil.”3 ...

June 20, 2026 · 19 min · Zelina