Cover image

Reward the Right Thing: GUI Agents Need Better Success Criteria, Not Just Better Judges

TL;DR for operators After a GUI agent sends a message, edits a document, moves a file, changes a setting, or searches for information, someone—or something—has to decide whether the instruction was actually completed correctly. It is tempting to treat that decision as mainly a model-capability problem: use a stronger vision-language model, show it more screenshots, or improve the prompt. The evidence in Task-Adaptive Rubrics for GUI Reward Modeling suggests another failure point comes earlier. The verifier first needs an adequately specified definition of success.1 ...

September 29, 2026 · 7 min · Zelina
Cover image

The Right Answer Is Not a Proof: Put Verification Inside the Reasoning Loop

TL;DR for operators A model can produce a correct answer while taking a logically invalid route to get there. That distinction matters whenever downstream execution depends not only on the answer, but on whether the intermediate decisions are trustworthy. Chen, Zhou, and Zhang test a lighter alternative to full theorem-proof generation: PRoSFI, which asks a 7B model to expose small, machine-readable reasoning steps that external formal tools can verify.1 On ProverQA-Hard, outcome-only reinforcement learning reaches 91.31% answer accuracy but only 21.97% GPT Soundness. PRoSFI reaches 92.97% accuracy and 76.07% GPT Soundness. The practical lesson is not that every business workflow should be formalized. It is that when intermediate decisions can be expressed as checkable rules, adding a verification layer can provide a stronger reliability signal than final-answer accuracy alone. ...

September 18, 2026 · 7 min · Zelina
Cover image

Three Loops Toward Self-Improvement

TL;DR for operators Model teams encounter three distinct constraints after a base model exists. New proprietary knowledge may remain accessible through retrieval without becoming reliably encoded in the weights. High-quality unique training text eventually becomes scarce even when additional training compute is available. Improvements to the training recipe still depend heavily on humans proposing, implementing, and testing experiments. ...

September 4, 2026 · 7 min · Zelina
Cover image

Check Your Work: Why Self-Verification Deserves Its Own Training Budget

TL;DR for operators A post-training team deciding where to spend its next training budget should not infer verification ability from task accuracy. In Learning to Self-Verify Makes Language Models Better Reasoners, Chen et al. find that training models to solve mathematical problems better does not reliably make them better at judging whether solutions are correct.1 Training the reverse capability behaves differently: models trained only to judge their own generated solutions subsequently solve problems about as well as models trained directly for generation. ...

September 3, 2026 · 8 min · Zelina
Cover image

Search Wider or Read Deeper: Where Long-Document Agents Should Spend the Next Token

TL;DR for operators A document assistant has already found several relevant pages but still cannot support an answer. The next action should depend on why the evidence is inadequate: perhaps another page is missing, or perhaps the answer is already present but buried in a table, region, or cross-page relationship that needs closer inspection. ...

August 27, 2026 · 8 min · Zelina
Cover image

Squeeze Evolve: When AI Stops Thinking Alone and Starts Allocating Intelligence

Budget is where many impressive AI demos go to become ordinary software. A model can reason longer. It can sample more. It can revise itself, compare candidates, aggregate outputs, and repeat the whole ritual until the invoice starts looking like a small infrastructure project. The obvious response is to ask whether the strongest model should simply do all of this work. Obvious, yes. Economically elegant, not quite. ...

April 11, 2026 · 21 min · Zelina
Cover image

The Latent Cost of Thinking: When LLM Reasoning Becomes a Liability

Thinking is expensive. That sounds obvious when the thinker is a human consultant billing by the hour. It sounds less obvious when the thinker is a large reasoning model producing long chains of thought, checking itself, trying another route, doubting the first answer, then generously spending another few thousand tokens to arrive at the same wrong place with better punctuation. ...

March 29, 2026 · 18 min · Zelina
Cover image

When Agents Hesitate: Smarter Test-Time Scaling for Web AI

Forms are boring. That is exactly why they are dangerous for AI agents. A human filling out an enterprise dashboard does not treat every click as a philosophical crisis. Search here. Scroll there. Submit. Done. A web agent, unfortunately, has no such common sense guarantee. It can overthink a routine step, miss a pivotal one, or spend a small fortune sampling twenty versions of the same obvious action. Very diligent. Also very expensive. ...

February 13, 2026 · 17 min · Zelina
Cover image

Confidence Is Not Truth, But It Can Steer: When LLMs Learn When to Stop

Stop Every production LLM workflow eventually meets the same boring question: should the model answer now, think again, or throw away the current path and try something else? That question sounds less glamorous than “build a bigger model.” It is also closer to where real deployment costs live. Reasoning models can improve by sampling more answers, extending chains of thought, or running repeated critique-and-revision loops. The bill, naturally, arrives in tokens, latency, GPU capacity, and engineering patience. The last item is rarely benchmarked, perhaps because it would make too many papers look expensive. ...

February 10, 2026 · 14 min · Zelina
Cover image

Conformal Thinking: Teaching LLMs When to Stop Thinking

Thinking is not free. That sentence should not need explaining to anyone who has paid an inference bill, waited for a reasoning model to finish its theatrical inner monologue, or watched an AI agent spend half its budget trying to solve a task it was never going to solve. Reasoning models have become better at using more tokens. They have not automatically become better at knowing when more tokens have stopped helping. ...

February 4, 2026 · 17 min · Zelina