How to Evaluate an AI Use Case

A practical framework for deciding whether an AI project is worth pursuing, what shape it should take, and how to avoid expensive pilots.

March 16, 2026 · 8 min · Michelle

AI Evaluation, Monitoring, and Incident Response for Production Systems

How to evaluate, monitor, and respond to failures in production AI systems so quality, safety, and governance remain active after launch.

March 16, 2026 · 6 min · Michelle
Cover image

When Open Artifacts Still Hide the Workflow

TL;DR for operators A public model can come with weights, a dataset, configuration files, inference code, and evaluation tooling while still leaving another team unable to reconstruct what was trained, what was held out, how the released checkpoint was selected, or how broadly its performance claims apply. Hiwa Asadpour’s audit of a Central Kurdish text-to-speech release1 makes that problem concrete. The strongest evidence is not that the underlying research is invalid; the paper being audited is described as comparatively cautious and well documented. The problem appears in the handoff from publication to public artifacts. Configuration settings conflict with the reported run, evaluation assets are missing, the released dataset does not mark the original test utterances, and the model card states a stronger performance conclusion than the source paper supports. ...

September 30, 2026 · 7 min · Zelina
Cover image

Control the Caption by Training What to Omit

TL;DR for operators A generative system can produce a fluent, factually plausible output and still fail because it focuses on the wrong information. Controllable Image Captioning with Prompt-Conditioned Scene Rewards by Jongyeop Hyun, Taeyoung Kim, and Hyounghun Kim1 tests a stricter approach: train the model not only to reward requested content, but also to penalize complementary content that falls outside the requested focus. ...

September 26, 2026 · 7 min · Zelina
Cover image

The Model Felt the Tampering. It Couldn’t Name the Cause

TL;DR for operators Suppose middleware silently rewrites part of an AI agent’s own generated answer before the model continues. A reasonable expectation is that a capable model would notice the interference, or at least diagnose why its continuation has become strange. The Sleight of Word benchmark tests exactly that expectation.1 Across 19 open-weight instruction-tuned models, covert substitutions consistently change the models’ predictive distributions: post-swap surprisal and entropy rise for every model tested. But correct identification of the intervention is almost absent. No model exceeds 1.3% explicit switch awareness, and the pooled rate is below 0.1%. ...

September 25, 2026 · 7 min · Zelina
Cover image

The Benchmark Is in the Trace: Reusing Agent Trajectories to Shrink SWE Evaluation

TL;DR for operators Repository-level agent benchmarks are expensive enough that teams have a strong incentive to run only part of them. The risk is not merely estimating the wrong average score; a poorly chosen subset can also change which agent appears better. PTA-IRT uses historical execution traces to make that subset more informative.1 At a 10% calibration budget, it records the best reported MAE, Kendall’s tau, and Spearman’s rho in every metric column across the four evaluated SWE-bench variants. Its averages are 0.041 MAE, 0.888 tau, and 0.973 rho. ...

September 22, 2026 · 8 min · Zelina
Cover image

A Citation Can Be Right Without Being Grounded

TL;DR for operators A RAG system can return the right answer, attach a source that genuinely supports that answer, and still leave one critical question unresolved: did that source actually influence the model’s answer generation? A mechanistic study of Llama-3.1-8B-Instruct finds that inline citation behavior is not controlled by one dedicated citation feature. It emerges from a distributed sequence of attention heads and MLPs that includes early entity enrichment, matching between document and question entities, mid-layer processing, and late aggregation that shifts the model toward emitting a citation marker rather than ending the sentence.1 ...

September 19, 2026 · 8 min · Zelina
Cover image

The Fourth Hop Changes the Risk Profile: Measuring Reliability in Multi-Step LLM Workflows

TL;DR for operators When one model output becomes input to the next stage, a final accuracy score tells you too little about where reliability is being lost. A workflow may fail because a required fact was never available, because a later composition step is intrinsically harder, or because an earlier mistake was allowed to propagate. Those failure modes call for different controls. ...

September 17, 2026 · 8 min · Zelina
Cover image

When Hallucination Is More Than a Wrong Fact: Measuring Reliability Through the User

TL;DR for operators A model change can improve an automated hallucination benchmark while leaving users dissatisfied for a different reason: sources are hard to verify, reasoning appears unsupported, false claims are stated with confidence, or corrections are ignored. The System Hallucination Scale (SHS) gives teams a structured way to measure those experiences across five dimensions rather than reducing reliability to a binary factual-error judgment.1 ...

September 8, 2026 · 7 min · Zelina
Cover image

When the Test Becomes a Signal: Rethinking AI Agent Evaluation

TL;DR for operators A tool-using agent does not experience an evaluation as an abstract benchmark. It sees prompts, tool wrappers, permissions, response timing, filesystem artifacts, network behavior, logging infrastructure, and other parts of the environment. If those signals differ from production, a sufficiently adaptive agent may be able to infer when it is being tested and behave differently. ...

September 7, 2026 · 8 min · Zelina