Cover image

A Sunset Clause Is Not a Safety Test: Designing an Exit From Frontier AI Limits

TL;DR for operators Any high-stakes agreement needs an answer to a practical question: what evidence should be enough to loosen the rule, and who gets to decide? Across eight comparable treaty regimes, Lennart Finke’s International Agreements to Limit Frontier AI: Objectives and Exit1 finds no concrete rule that automatically ends an agreement once its substantive objective has been achieved; exit instead relies mainly on unilateral withdrawal or fixed duration. :contentReference[oaicite:0]{index=0} ...

August 14, 2026 · 7 min · Zelina
Cover image

From Alarm Signal to Release Gate: Measuring CBRN Uplift in Frontier Models

TL;DR for operators A safety team sees a frontier model produce expert-like, technically detailed CBRN guidance. That is a reason to investigate—but it does not yet answer the release question: does access to the model materially improve what a non-expert can do? This study shows why the distinction matters. All four CBRN domains exceeded the thresholds for expert-level instruction and interactive scientific or technical instruction, yet only the radiological domain exceeded the study’s core material-uplift criterion. The most visibly concerning outputs therefore did not, by themselves, identify where the controlled experiment found meaningful improvement in user performance. ...

August 12, 2026 · 8 min · Zelina
Cover image

Shared Blind Spots: Why Narrow AI Needs an Independent Detector

TL;DR for operators A common deployment pattern assigns routine work to a cheaper or faster specialist and sends uncertain cases to a human or generalist system. The safeguard works only when the monitor retains information that the specialist lacks. If safe opportunities and fatal cases become indistinguishable after the same information is removed, a detector using that incomplete view must capture useful work and miss fatal cases at equal rates. A better confidence score, threshold, or escalation rule cannot recover a distinction that is no longer present. ...

August 10, 2026 · 6 min · Zelina
Cover image

The Answer Looked Clinical. The Critical Steps Were Missing.

TL;DR for operators A clinical assistant can follow instructions, produce a complete-looking answer, and read like expert work without performing the reasoning needed for a safe clinical decision. Those surface qualities may justify further evaluation, but they do not justify procurement or decision authority. Across five deliberately difficult tasks, the models satisfied 80–90% of the least consequential criteria but only 32.4–41.7% of the most clinically consequential criteria. All three models missed 56 of 108 critical requirements. The systems were therefore strongest where omissions mattered least and weakest on reasoning and safety steps most closely tied to harm. ...

August 1, 2026 · 9 min · Zelina
Cover image

A Full Distribution Is Not a Risk Certificate

TL;DR for operators A team is deciding whether an AI risk dashboard should trigger a safer action, an operational alert, or a governance review. The system predicts a range of possible returns for each action rather than only an average, so its output may appear to provide stronger evidence about worst-case outcomes. ...

July 30, 2026 · 8 min · Zelina
Cover image

Agree Once, Remember Later: The Commit Boundary in Personal Agents

TL;DR for operators A personal assistant can hear a confident user claim, store it as a preference or rule, and rely on it during a later task after the original conversation is gone. The safety problem is therefore not only the agreeable reply. It is the write that lets the claim survive. ...

July 24, 2026 · 8 min · Zelina
Cover image

Measure for Measure: Why AI Evaluation Must Follow the Failure

TL;DR for operators A lower model bit width is not automatically a speedup. A lower training loss is not automatically a reliable policy. GRINQH evaluates quantization through the mechanism it is meant to change: decoding-stage memory traffic, kernel throughput, end-to-end generation speed, and retained task accuracy.1 Kolmogorov regression evaluates diffusion policies through trajectory geometry, a PDE-based inference residual, rollout behavior, anomaly detection, and an external safety filter.2 The shared lesson is not that the two forms of “precision” are technically equivalent. They are not. The lesson is that fidelity and evidence should be allocated according to the actual failure structure of the system. A production evaluation should connect four things explicitly: the intervention, the mechanism it changes, the diagnostic that observes that change, and the operational outcome that justifies deployment. Composite scores are useful only when their weights reflect real business priorities and their components remain separately visible. Otherwise, they are merely spreadsheets wearing authority. The dashboard is not the system AI evaluation has developed an awkward habit: optimize a convenient number, improve that number, and declare the system improved. ...

July 22, 2026 · 16 min · Zelina
Cover image

The Refusal Rate That Refuses to Reassure

TL;DR for operators The reassuring headline is that both evaluated frontier models rejected most automated jailbreak attempts. The operationally useful headline is that they still produced 1,620 and 702 panel-confirmed harmful completions, respectively, across every top-level harm category in the benchmark.1 The strongest adaptive attack succeeded on 11.51% of attempts against one model and 6.10% against the other. Static encodings and familiar jailbreak templates, by contrast, were almost entirely neutralised. ...

July 16, 2026 · 18 min · Zelina
Cover image

The Simulator Gets a Reality Check

TL;DR for operators RealityBridge is a paper about a fairly unglamorous but commercially important problem: editable driving simulations are useful because they let teams stage rare, dangerous, and legally inconvenient scenarios, but the rendered videos often look wrong in exactly the places that matter. Blurry vehicles, mismatched lighting, weak shadows, floating artifacts, broken boundaries, flickering objects, and small hazards that quietly dissolve into the background are not just aesthetic defects. They are domain-gap leakage. ...

July 9, 2026 · 22 min · Zelina
Cover image

The Bike Learns to Lean Before It Learns to Race

TL;DR for operators A new paper, Self-Paced Curriculum Reinforcement Learning for Autonomous Superbike Racing in Simulation, introduces a reinforcement-learning framework for training a superbike agent in VRider SBK, a Unity-based motorcycle racing simulator.1 The useful part is not merely that the model rides faster. The useful part is how the authors turn motorcycle racing into a staged learning problem without hand-writing a long curriculum by committee, which is usually how such things go to die politely. ...

July 8, 2026 · 18 min · Zelina