Cover image

Before the First Token: Put the Jailbreak Gate Inside the Model

TL;DR for operators A product team normally has several places to stop a dangerous request: filter the prompt, ask another model to inspect it, regenerate under stricter instructions, or moderate the output after generation. All of those controls act around the language model. GUARD-SLM asks whether the model’s own internal representation can provide the stop signal earlier. ...

September 21, 2026 · 7 min · Zelina
Cover image

Guide, Don’t Guess: Reallocating Reasoning Compute with Small-Model Hints

TL;DR for operators When a compact model fails on a difficult multi-step problem, the default remedies are expensive: use a larger model, generate many independent attempts, or accept lower accuracy. This paper tests a different allocation of compute—keep the solver small, but give it localized guidance at difficult intermediate steps. HintMR1 separates those jobs. A hinter tells the solver what to consider next without supplying the full solution; the solver then advances its reasoning one step at a time. On AIME-2024, DeepSeek-R1-Distill-Qwen-7B rises from 20.69% accuracy without hints to 68.97% with GPT-5.2-generated hints. The result is not evidence that any second model helps: non-fine-tuned small-model hints are inconsistent and sometimes reduce accuracy below the no-hint baseline. ...

September 18, 2026 · 6 min · Zelina
Cover image

When Prompt Optimization Overloads the Small Model

TL;DR for operators A small model has a finite inference budget. Prompt optimization consumes that budget too: longer instructions occupy context, optimization calls add latency, and extensive rewriting can alter an input the model might already have understood. Shim and colleagues test a different operating rule in POaaS: inspect each query first, leave sufficiently good prompts alone, and apply narrowly targeted repairs only when a specific deficiency is detected.1 On Llama-3.2-3B, this raises average clean task accuracy from 63.7% with no optimization to 66.0%. Under the same fixed-small-model protocol, EvoPrompt, OPRO, and PromptWizard fall to 59.8%, 57.5%, and 48.8%. ...

September 4, 2026 · 7 min · Zelina
Cover image

The 70B Model May Belong Upstream

TL;DR for operators If a team has a small labelled seed set and a large volume of multilingual text to classify, keeping the strongest LLM in every inference request may not be the best allocation of compute. Pecher et al. find that smaller models using examples generated by LLaMA-3 70B can exceed that same 70B model used directly as a zero-shot classifier with roughly 50 synthetic examples in aggregated language groups.1 ...

September 2, 2026 · 8 min · Zelina
Cover image

Reasoning on Demand: AdaHome’s Case for Tiered Local Assistants

TL;DR for operators A household assistant should not spend the same computational effort on “turn on the light” as on “make the room comfortable.” It also should not treat one unusual request as a permanent preference change. AdaHome applies that principle through tiered local processing. Explicit commands take a short planning path, while requests that need interpretation or personal context receive additional reasoning, validation, and—where appropriate—user confirmation. Under a common Llama 3.2-3B setup, it achieved 86.7% success on both direct and indirect commands while recording the lowest latency and token use in every command category among the compared systems. ...

August 8, 2026 · 7 min · Zelina
Cover image

The Yap Trap: Why AI Reasoning Needs a Governor

Long reasoning has become the new luxury trim in AI products. The demo no longer just answers. It pauses, reflects, reconsiders, checks itself, writes a small philosophical memoir, and then hopefully solves the problem. This is not entirely theatrical. Chain-of-thought style reasoning and large reasoning models have improved performance on difficult tasks, especially in mathematics, coding, planning, and multi-step analysis. For business users, that matters. A model that can break down a problem is more useful than one that confidently blurts out the first plausible answer. Nobody wants a legal assistant, financial analyst, or production-support agent whose main cognitive strategy is “vibes, but fast.” ...

June 9, 2026 · 16 min · Zelina
Cover image

Less Chain, More Thought: The Coming Control Layer for LLM Reasoning

Less Chain, More Thought: The Coming Control Layer for LLM Reasoning Enterprise AI has spent the last two years discovering a mildly inconvenient truth: a model that explains itself at length is not necessarily reasoning well. It may be reasoning. It may be narrating. It may also be producing a confident procedural bedtime story with a spreadsheet attached. ...

June 2, 2026 · 15 min · Zelina
Cover image

Prompt and Circumstance: Why One Accuracy Number Is Not a Reliability Audit

Opening — Why this matters now The AI market has learned to worship benchmark tables with the solemnity once reserved for quarterly earnings. One model is up two points on MMLU, another is slightly better at reasoning, a third is cheaper, smaller, faster, and therefore apparently ready to run your compliance workflow by Tuesday. ...

May 7, 2026 · 14 min · Zelina
Cover image

From Perception to Empathy: Why Small Models May Win the Emotional AI Race

Customer support is where emotional AI often goes to embarrass itself. A user says, “Fine, whatever.” The system detects a neutral sentence. A human hears irritation, resignation, and possibly the final five seconds before churn. The difference is not vocabulary. It is context, tone, facial expression, timing, and the reason behind the emotion. Unfortunately, many “emotion AI” systems still behave as if the job is to pick a label from a menu: happy, sad, angry, neutral. Very scientific. Also very convenient, because menus are easier than people. ...

March 3, 2026 · 14 min · Zelina
Cover image

Small Models, Big Skills: When Agent Frameworks Meet Industrial Reality

Compliance has a wonderful way of killing beautiful demos. In a demo, the agent calls a frontier model, loads a tool, reads a document, writes a decision, and everyone nods at the future. In a regulated company, the same workflow meets a less poetic checklist: where did the data go, who pays for the GPU time, can this run inside our perimeter, and why did the model spend twenty seconds “thinking” about a binary classification task? ...

February 19, 2026 · 15 min · Zelina