Cover image

Synthetic Data Can Make the Model Worse

TL;DR for operators A team with authoritative domain documents but little labeled training data has an attractive option: ask a capable model to manufacture question-answer pairs, then fine-tune a smaller open model on them. The operational risk is assuming that domain relevance makes those examples safe training material. In this paper, a simple synthetic-data pipeline moved LLaMA 3.1 8B backward on open-ended legal QA: its LegalMC4 score fell from 43.0% to 35.4%. A more structured pipeline raised the same score to 55.4%. Across both LLaMA 3.1 8B and Gemma 3 12B, that structured treatment improved all four tested German legal benchmarks.1 ...

September 3, 2026 · 7 min · Zelina
Cover image

Synthetic Data Needs an Evidence Contract

TL;DR for operators Synthetic data should have a defined job before anyone scales its production. For model training, the relevant test is whether generated examples add nonredundant learning signal and improve held-out performance without unacceptable regressions. For consumer research, the test changes: statistically diverse text is not enough if the business claim concerns what real customers believe. ...

September 3, 2026 · 8 min · Zelina
Cover image

When the Gazetteer Goes Blank, the Records Still Know Where the Place Is

TL;DR for operators When a historical place name is missing from a gazetteer, the usual workflow treats it as a lookup failure. This paper shows another option: use multiple records that mention the same place from already known coordinates, convert descriptions such as “1 km east of X” into constraints on where X must be, and aggregate those clues. ...

August 26, 2026 · 7 min · Zelina
Cover image

Edge Control: Why Synthetic Graphs Need a Repair Pass

TL;DR for operators Synthetic graph data is easy to make look plausible and hard to make structurally right. A graph can have the right number of nodes, a sensible average edge count, and a respectable generative model behind it, while still getting the relational geometry wrong. In graph domains, that is not a cosmetic flaw. The edges are the thing. ...

June 18, 2026 · 19 min · Zelina
Cover image

Mind the Representation Gap: Why Enterprise AI Fails Before It Thinks

Enterprise AI has developed a charming habit: whenever a system fails, someone suggests using a larger model. The chatbot misread a customer complaint? Bigger model. The autonomous system struggled with a new sensor configuration? Bigger model. The video classifier understood the objects but missed the actual message? Bigger model, possibly with a more expensive logo. ...

June 11, 2026 · 14 min · Zelina
Cover image

Synthetic Data, Real Receipts: Why LLM Pipelines Need an Auditor

Opening — Why this matters now Synthetic data has become one of AI’s favorite escape routes. Real data is expensive, legally awkward, slow to collect, unevenly labeled, and sometimes simply unavailable. LLMs offer a tempting alternative: generate the missing examples, fill the long tail, create evaluation suites, simulate edge cases, and keep the training pipeline moving. Convenient. Elegant. Also mildly dangerous, which is usually where the interesting part begins. ...

April 25, 2026 · 12 min · Zelina
Cover image

Seeing Is Judging: Why LLMs Are Better Critics Than Creators in Time-Series Reasoning

A dashboard says revenue demand has “stabilized.” A monitoring agent says a sensor spike is “temporary.” A trading assistant says volatility has “fallen after the regime shift.” The sentence is smooth. The chart is nearby. The user is tired. That is usually enough for a bad explanation to survive. This is the quiet problem behind AI-assisted analytics: not whether a language model can write a plausible story about time-series data, but whether the story is faithful to the numbers. A recent paper, LLM-as-a-Judge for Time Series Explanations, studies exactly this gap by asking models to play two different roles: narrator and critic.1 ...

April 4, 2026 · 16 min · Zelina
Cover image

Redundancy Overload Is Optional: Finding the FDs That Actually Matter

Tables have a talent for pretending to be tidy. A customer table may have fifty columns. A transaction table may have a hundred. A log table may contain derived fields, timestamps, status codes, copied identifiers, normalized labels, and a few columns that nobody remembers creating but everybody is afraid to delete. Then a data profiling tool arrives, dutifully discovers functional dependencies, and returns several hundred thousand “valid” relationships. ...

January 18, 2026 · 19 min · Zelina
Cover image

When Benchmarks Rot: Why Static ‘Gold Labels’ Are a Clinical Liability

Clinical AI has a paperwork problem. Not the usual paperwork problem, where doctors drown in documentation and everyone promises that software will save them. The more interesting problem sits one layer below: the paperwork used to judge the software may itself be wrong. That is the uncomfortable center of Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight, a paper that audits MedCalc-Bench, a benchmark for testing whether language models can compute medical risk scores from patient narratives.1 The paper’s target is not a toy dataset. MedCalc-Bench covers 55 medical calculators and includes 10,053 training instances plus 1,047 test instances. Its labels were produced through an LLM-assisted pipeline: GPT-3.5 matched patient contexts to calculator questions, GPT-4 extracted clinical features, and Python scripts aggregated those features into final scores. ...

December 23, 2025 · 15 min · Zelina
Cover image

Trust Issues: Why Neural Networks Need Their Own Internal Affairs Department

Accuracy is a comforting number. That is precisely the problem. A neural network can score well on a test set and still be operationally suspicious. The labels may be corrupted. The input may be degraded. A small patch may have quietly hijacked part of the model’s learned behavior. The model may be confident, calibrated enough for a dashboard, and still untrustworthy in the one place where the business actually needs it to behave. ...

November 26, 2025 · 16 min · Zelina