Cover image

Layers Are Not Independent: Red-Team the Whole AI Safety Stack

Cascade shows why AI safety testing has to include the software and hardware conditions that determine whether model-level defenses actually execute as intended.

September 21, 2026 · 8 min · Zelina
Cover image

Put the LLM Upstream, Not in the Loop

A controlled simulation workflow shows how LLMs can expand behavioural and scenario specification without taking execution authority away from a calibrated model.

September 21, 2026 · 7 min · Zelina
Cover image

Reward the Claim, Not Just the Answer: What V-Rubrics Changes in Multimodal RL

V-Rubrics shows how criterion-level and localized reinforcement-learning credit can make multimodal post-training more targeted when intermediate visual and reasoning errors matter.

September 21, 2026 · 7 min · Zelina
Cover image

Who Gets the Veto? Allocate AI Authority to the Costlier Error

A framework for assigning human vetoes and bounded AI override rights according to which failure—acting wrongly or failing to act—carries the greater cost.

September 21, 2026 · 8 min · Zelina
Cover image

A Refusal Is Not a Safety Test: Probe Harm After the Prompt Changes Form

Emoji-based jailbreak tests show why pre-release safety evaluation must vary how harmful intent is represented, not just whether a model refuses explicit text.

September 20, 2026 · 7 min · Zelina
Cover image

Don’t Give Every Query the Same Context Budget

ATACompressor shows why context compression should decide both what evidence to preserve and how much representation capacity each query deserves.

September 20, 2026 · 7 min · Zelina
Cover image

Don’t Reason Over Every Token: LycheeMemory Turns Long Context Into a Budget

LycheeMemory shows how compressed memory and state-dependent gating can cut long-context inference cost by allocating expensive reasoning only to selected evidence.

September 20, 2026 · 5 min · Zelina
Cover image

Recall Is Not a Retrieval Strategy: Choosing GraphRAG or VectorRAG by Workload

A polymer-literature benchmark shows why retrieval recall alone cannot decide between GraphRAG and VectorRAG for technical knowledge systems.

September 20, 2026 · 7 min · Zelina
Cover image

Safety Has a Memory: Why Multimodal Jailbreak Testing Must Follow the Conversation

A multimodal assistant that refuses one harmful prompt may still fail after benign conversational context has accumulated, changing how safety teams should test interactions and moderate outputs.

September 20, 2026 · 7 min · Zelina
Cover image

Train the Store, Audit the Facts: What KBevo Changes About Retrieval

KBevo shows that an external knowledge store can be optimized for downstream answers and edited later, but better task utility does not make the stored facts automatically trustworthy.

September 20, 2026 · 7 min · Zelina