Cover image

The Generalization Stack: Why HAR Robustness Is a Pipeline Property

TL;DR for operators A smartphone activity-recognition model can perform well during development and then degrade when deployed on a different dataset, user population, or sensor position. The natural response is to search for a domain-generalization technique that performs best across these changes. A 410,400-experiment benchmark suggests that this is the wrong unit of comparison. No individual objective, initialization strategy, or architectural modification wins consistently. Yet compatible combinations can produce materially larger gains: the best aggregate joint configurations improve accuracy by about 2.9 percentage points under cross-dataset shift and 4.9 points under cross-position shift. ...

September 28, 2026 · 8 min · Zelina
Cover image

Reasoning Labels Don’t Travel: What UrduBench Changes About Model Selection

TL;DR for operators UrduBench1 tests 23 open and open-weight models on 2,390 held-out Urdu questions spanning arithmetic reasoning, formal mathematics, commonsense, and knowledge tasks. The ranking gives little support to selecting an Urdu model from parameter count or a “reasoning” label alone: Gemma-3-12B-it leads the reported aggregate at 59.4%, while the larger reasoning-oriented DeepSeek-R1-Distill-Qwen-14B scores 44.9%. ...

September 18, 2026 · 7 min · Zelina
Cover image

The Best Channel Model Depends on the Channel

TL;DR for operators A wireless team choosing a model to standardize has a tempting shortcut: benchmark several candidates, find the highest score, and carry that winner into deployment-specific adaptation. The problem is that wireless-channel performance depends heavily on what is being predicted and on the propagation environment in which the model is evaluated. ...

September 12, 2026 · 7 min · Zelina
Cover image

When Worse Inputs Score Better: Audit the Credibility Behind the Benchmark

TL;DR for operators A benchmark score can be high without being equally trustworthy as a measure of generalization. In this study, the researchers deliberately degraded benchmark questions before they reached the answering model. At a noisy-router count of eight, 10 of the 12 evaluated models nevertheless scored above their own clean baseline. At nine routers, eight models still did so, and the mean positive excess among above-baseline cases reached 0.086. ...

September 10, 2026 · 7 min · Zelina
Cover image

Correct to Select: Choosing OCR Without Ground Truth

TL;DR for operators A team has a new collection of forms, receipts, exam papers, or clinical notes and must choose an OCR engine before paying to transcribe the collection. A fixed benchmark winner is not enough: in the paper’s evaluation, no single OCR engine consistently leads across datasets and languages. That creates a practical sequencing problem. The best engine depends on the documents, but directly measuring which engine is best normally requires the ground-truth labels the team is trying to avoid creating before selection. ...

August 11, 2026 · 8 min · Zelina
Cover image

Rank the Work, Not the Model: Meta-Benchmarks for Bank LLM Screening

TL;DR for operators A bank must choose a model for a particular workflow, but it cannot run a full internal evaluation whenever a new model appears. General leaderboards offer a useful starting point—not a reliable answer to which model best fits the banking work that matters. The reported relationship between global and banking-domain rank varies substantially. Spearman correlation is 0.62 for Customer Management and 0.71 for IT Management, but 0.95 for Market Operations and 0.97 for Regulations and Compliance. The strongest model overall may therefore be less compelling for a specific domain, while an apparently precise domain score may rest on limited or indirect evidence. ...

July 29, 2026 · 10 min · Zelina
Cover image

The Agent Benchmark Without the Agent Bill

TL;DR for operators Agent evaluations are expensive for a fairly obvious reason: the agent has to do something. It must browse, edit files, call tools, manipulate repositories, survive its own mistakes, and occasionally discover that the environment has changed while nobody was looking. The paper introduces Pace, a method for predicting performance on an expensive agentic benchmark from a compact set of cheaper, non-agentic test instances.1 Across 14 frontier models and four agentic benchmarks, a 100-instance Pace proxy produces an average mean absolute error of 3.80 percentage points, a 0.81 Spearman rank correlation, and 84.37% pairwise model-ranking accuracy under leave-one-model-out validation. ...

July 16, 2026 · 16 min · Zelina
Cover image

Bond Before Brain: What Actually Drives Molecular MPNNs

TL;DR for operators Molecular GNN selection is often sold as a choice among branded architectures: DMPNN, AttentiveFP, Graphormer, and the rest of the respectable parade. This paper asks a more useful question: before buying the whole architecture, which part of the message-passing pipeline is actually carrying the performance signal? The answer, within this study’s controlled 2D setting, is message construction. The authors benchmark 84 molecular MPNN configurations across ten MoleculeNet tasks by varying three operator families: message-seed initialization, node-edge fusion, and node update. They hold sum aggregation, sum readout, featurization, scaffold splits, tuning protocol, and statistical analysis fixed. That makes the benchmark less glamorous than a new model launch, and substantially more useful. ...

June 15, 2026 · 16 min · Zelina
Cover image

Judge, Jury, and Benchmark: Why LLM Evaluation Needs Fresh Cases, Not Bigger Leaderboards

The procurement meeting is where public leaderboards go to look useful Benchmark scores are comforting because they compress chaos into a number. One model is 87.3, another is 84.9, and suddenly the procurement meeting has the emotional texture of financial discipline. Very mature. Very measurable. Also, very possibly irrelevant. The problem is simple. A company rarely wants “the best model on average”. It wants the best model for contract review, support triage, clinical note summarisation, SQL repair, claims handling, product search, or whatever unglamorous workflow actually pays the cloud bill. Public benchmarks are often too generic for that decision. Worse, the benchmark items may already be floating inside model training data, turning evaluation into a memory test with better typography. ...

June 12, 2026 · 18 min · Zelina
Cover image

Rank and File: AI Leaderboards Are Measurement Instruments, Not Scoreboards

Procurement meetings have a familiar ritual now. Someone opens a leaderboard, sorts by average score, points at a model near the top, and asks why the company is not using that one. It feels empirical. It is neatly ranked. It has decimals. Very scientific-looking decimals, the most seductive species of decimal. The problem is not that leaderboards are useless. The problem is that we often treat them as scoreboards when they are closer to measurement instruments. A scoreboard tells us who won under agreed rules. A measurement instrument first has to prove that it measures the thing it claims to measure. If the instrument mixes model size, benchmark difficulty, contributor practices, post-training choices, item redundancy, and residual artifacts into one number, then the number may still be useful. It is just not self-explanatory. ...

June 4, 2026 · 18 min · Zelina