Cover image

Free Will Without Randomness: A Practical Test for AI Agency

TL;DR for operators When an AI system selects among actions, follows goals, and changes behavior when its internal representations change, there are at least two ways to describe what is happening. One is purely mechanistic: code receives inputs and produces outputs. The other treats the system as an agent choosing among options for reasons. ...

September 30, 2026 · 8 min · Zelina
Cover image

Watching a Signal the Model Can Move

TL;DR for operators A latent-space monitor is valuable only if the internal signal it reads is sufficiently difficult for the monitored model to manipulate. Measuring Activation Control in Large Language Models tests that assumption directly.1 Across 25 open-weight instruction-tuned models, the authors find meaningful but coarse control over concept-related internal representations. Models can raise a concept signal, suppress it toward its ordinary baseline, order several requested intensity levels, and move signal toward broad regions of a sentence. They are much less successful at targeting a particular layer or restricting modulation to particular token groups. ...

September 28, 2026 · 9 min · Zelina
Cover image

Local Evidence, Global Rule: When Knowledge Graph Embeddings Generalize Too Far

TL;DR for operators A knowledge graph may observe the same relational pattern only a handful of times and still use an embedding model whose architecture effectively treats that pattern as valid everywhere. That is not merely a sparse-data problem. It is a generalization-control problem. Kim and Kim call this failure pattern over-generalization and propose PogRE, a knowledge graph embedding architecture designed to make a pattern’s reach expand as supporting evidence covers more independent directions in embedding space.1 The paper reports both competitive link-prediction results and lower targeted over-generalization measurements than TransE, RotatE, PairRE, and CompoundE in the evaluated settings. ...

September 27, 2026 · 7 min · Zelina
Cover image

More Agents, More Rules: HiMA-MDD Treats Multi-Agent AI as a Governance Problem

TL;DR for operators Dividing a high-stakes decision among specialist agents does not specify who may see which evidence, who owns each subdecision, or who can revise it. HiMA-MDD1 turns those choices into explicit system rules for PHQ-8 assessment from completed multimodal clinical interviews. The strongest architectural signal is not “more agents perform better.” Before global verification, four specialists produce the lowest total-score error, while two specialists produce the highest screening kappa and Macro-F1. Removing cross-factor audit and targeted revision causes the largest screening-performance decline among the paper’s three component ablations. Giving four specialists bounded item-specific evidence also beats giving them a capped shared evidence pool on every reported pre-verification metric, although that experiment changes both evidence composition and context length. ...

September 25, 2026 · 6 min · Zelina
Cover image

The Safety Leaderboard Has Conditions: Choose Moderators by Harm, Context, and Cost

TL;DR for operators A moderation team rarely needs a model that is merely “best at safety.” It needs a model that catches the relevant harms at the point where moderation occurs, without creating unacceptable false positives or latency. A large benchmark by Afshin Orojlooyjadid and Hitesh Patel compares 53 specialized moderators and general-purpose language models across 11 public safety datasets.1 Its strongest operational finding is not a new overall winner. It is that the ranking changes with the safety problem and with what the moderator is allowed to inspect. ...

September 23, 2026 · 8 min · Zelina
Cover image

State Before Action: OODA-Tool Puts a Control Layer Between Context and Execution

TL;DR for operators A tool-using agent can remember the right customer, constraint, or prior result and still make the wrong call. From State to Action: OODA-Tool for Reliable Multi-Turn Tool Use1 treats that gap as an architectural problem rather than only a prompting problem. Its strongest configuration separates four jobs: reconstruct the active task state, decide whether execution is actually warranted, choose a permitted action structure, and only then bind concrete arguments. On ToolDial, Specialized OODA beats Direct-LoRA at every tested Qwen3 scale, by 4.48 to 6.99 percentage points in Task Success. The gains are largest where state must survive long histories, missing information, changed values, constraints, or sequential dependencies. ...

September 22, 2026 · 8 min · Zelina
Cover image

State Is the Workflow: AstronOS Moves Long-Horizon Agents Beyond Transcript Replay

TL;DR for operators A multi-step agent workflow eventually faces a problem that a larger context window does not solve. One model makes a decision, another step runs later with new evidence, and the system must know which earlier facts and decisions still count as official. Replaying the transcript gives the next executor more text; it does not necessarily tell it what has been accepted, superseded, or rejected. ...

September 22, 2026 · 7 min · Zelina
Cover image

Who Gets the Veto? Allocate AI Authority to the Costlier Error

TL;DR for operators A high-stakes AI system can fail in two very different ways: it can act on something that is not true, or it can fail to act when danger is real. Those errors need not have comparable consequences. Martino Maggetti’s Reciprocal Trust and Distrust in Artificial Intelligence Systems: The Hard Problem of Regulation1 argues that this asymmetry should influence who receives final decision authority. In nuclear launch and strategic-warning settings, where a false positive could trigger catastrophic action, the paper favors protected human authority, independent corroboration, and explicit uncertainty. In reactor, chemical-process, and flight-control settings, where failing to intervene can be catastrophic, it allows for bounded AI authority through mechanisms such as non-overridable shutdown logic. ...

September 21, 2026 · 8 min · Zelina
Cover image

A Refusal Is Not a Safety Test: Probe Harm After the Prompt Changes Form

TL;DR for operators A chatbot that refuses a plainly written harmful request has passed one test of its safety behavior, not the whole test. In Emoji-Based Jailbreaking of Large Language Models, Gopinadh and Hussain submit the same 50 emoji-augmented adversarial prompts to four locally deployed open-source models.1 They report successful jailbreak rates of 10% for Gemma 2 9B, 10% for Mistral 7B, 6% for Llama 3 8B, and 0% for Qwen 2 7B. ...

September 20, 2026 · 7 min · Zelina
Cover image

When Reasoning Leaves the Prompt: Designing the Agentic Control Loop

TL;DR for operators When an LLM has to plan, call a tool, inspect the result, remember what happened, and decide what to do next, the system has changed in a more fundamental way than “adding more reasoning steps.” Agentic Reasoning for Large Language Models frames that change as a move from mostly static generation toward an interactive reasoning-and-control loop.1 ...

September 18, 2026 · 7 min · Zelina