Cover image

Synthetic Data Needs an Evidence Contract

TL;DR for operators Synthetic data should have a defined job before anyone scales its production. For model training, the relevant test is whether generated examples add nonredundant learning signal and improve held-out performance without unacceptable regressions. For consumer research, the test changes: statistically diverse text is not enough if the business claim concerns what real customers believe. ...

September 3, 2026 · 8 min · Zelina
Cover image

Synthetic Experience, Real Transfer: Build the Test Before You Scale the Data

TL;DR for operators Synthetic data should not be budgeted as a cheaper substitute for human examples. It should be treated as infrastructure for producing controlled training experience. The operational sequence is: generate tasks that can actually be executed and scored; verify and repair them before spending compute on trajectories; choose a training objective that reinforces the capability you want rather than merely reproducing successful-looking behavior; and test the resulting model outside the environment in which that experience was generated. ...

September 3, 2026 · 8 min · Zelina
Cover image

Train the Agent Before the Sandbox Exists

TL;DR for operators A team can have stable API contracts before it has an executable sandbox, populated user state, and reliable integration infrastructure. The usual assumption is that serious multi-step agent training must wait, because later API responses need to reflect what earlier calls changed. ESAT challenges that sequencing. Lee et al. generate stateful training trajectories from API specifications without the target environment, then validate and filter the interactions before training.1 On AppWorld, training with ESAT-S52 trajectories generated from 52 synthetic applications independent of AppWorld improves Task Goal Completion by 8.4 to 47.0 percentage points across evaluated models and usually outperforms training on trajectories collected inside the real AppWorld environment. ...

September 3, 2026 · 7 min · Zelina
Cover image

When More Problems Stop Helping: RL Data Scaling Becomes an Allocation Problem

TL;DR for operators A post-training team with a fixed compute budget has several ways to spend it: add more verified problems, generate variants of existing problems, increase difficulty, or expose the model to structurally different tasks. The usual accounting metric—number of training problems—does not tell the team which choice produces the most useful learning experience. ...

September 3, 2026 · 8 min · Zelina
Cover image

When the Research Workflow Becomes Training Data

TL;DR for operators A team with an expensive research workflow faces a recurring choice: keep paying for the full workflow on every report, or use it to teach a cheaper system how to reproduce much of its behavior. O-Researcher shows why the second option is plausible.1 With the same GPT-5 model, changing research execution from sequential to parallel raises the reported Overall score from 42.92 to 49.60, while Comprehensiveness rises from 40.59 to 49.61. Workflow structure itself is contributing capability. ...

September 3, 2026 · 7 min · Zelina
Cover image

The 70B Model May Belong Upstream

TL;DR for operators If a team has a small labelled seed set and a large volume of multilingual text to classify, keeping the strongest LLM in every inference request may not be the best allocation of compute. Pecher et al. find that smaller models using examples generated by LLaMA-3 70B can exceed that same 70B model used directly as a zero-shot classifier with roughly 50 synthetic examples in aggregated language groups.1 ...

September 2, 2026 · 8 min · Zelina
Cover image

When the Simulator Becomes the Curriculum

TL;DR for operators A simulator can generate effectively unlimited trajectories. That does not automatically make those trajectories useful training data for a reasoning model. Sim2Reason turns simulated mechanics into questions with answers that can be checked automatically, then uses those questions for reinforcement-learning post-training. The reported transfer is substantial enough to matter operationally: on Qwen2.5-32B, International Physics Olympiad mechanics accuracy rose from 19.8% to 25.2%. By contrast, supervised fine-tuning on 200,000 teacher-generated trajectories lowered the same score to 15.9%. ...

September 2, 2026 · 7 min · Zelina
Cover image

Hide the Worker, Keep the Geometry: What SynthSite Changes About Privacy-Aware Safety Video

TL;DR for operators A safety team may need to conceal workers before video leaves a trusted environment, yet hiding appearance can also remove the geometry needed to detect whether someone is beneath a suspended load. SynthSite tests that conflict directly.1 Across 55 synthetic clips, cartooning produced the highest F2 against human safe/unsafe labels at 0.767, while Canny-edge produced the highest F2 for reproducing the raw-video pipeline at 0.964. ...

August 18, 2026 · 7 min · Zelina
Cover image

Fix the Worst Frame First: Adaptive Anchoring for Synthetic Video Supervision

TL;DR for operators A synthetic training video can look coherent at its beginning and end while losing the target identity somewhere in the middle. The paper proposes changing the synthetic-data factory so that every generated frame is checked against a real target-identity reference, the weakest eligible frame receives an additional identity anchor, and only then is the affected span regenerated. ...

August 17, 2026 · 7 min · Zelina
Cover image

Clean Less, Route Better: DataOrchestra Reframes Pretraining Data Curation

TL;DR for operators A pretraining-data pipeline receives millions of uneven records. Some are unusable, some contain removable noise, some need structural repair or added explanation, and some are already valuable enough that further processing may damage them. The operational problem is therefore not how to apply more cleaning, but how to decide which intervention—if any—each example needs. ...

August 9, 2026 · 8 min · Zelina