Cover image

Teach the Primitive, Not the Picture

TL;DR for operators A synthetic-data budget creates a design choice: generate examples that resemble deployment tasks, or construct simpler exercises that isolate the operations a model will repeatedly need. SpatialBlock1 provides evidence for the second strategy. In a matched Qwen2.5-VL-3B experiment, synthetic training on conventional relative-direction and relative-distance questions improves those same in-domain tasks. But the SpatialBlock curriculum performs better on several external evaluations: 49.1 versus 32.8 on MindCube, 28.1 versus 25.5 on MMSI-Bench, and 46.4 versus 43.6 on MMMU. ...

September 29, 2026 · 7 min · Zelina
Cover image

Make the Structure Survive: Forma Turns Synthetic Clinical Cases Into Auditable Outputs

TL;DR for operators If synthetic cases are going to support training or evaluation, surface plausibility is a weak control objective. The harder requirement is preserving the relationships that make each case meaningful. Forma tests that idea by specifying a person-specific psychological structure before generation and asking whether those directional relationships can be recovered afterward. In the full condition, an external probe reaches MCC +0.41 and AUC 0.70 for directed-edge recovery. When the structural formulation is removed, performance falls close to chance: MCC +0.03/AUC 0.52 with demographics and self-report still present, and +0.01/0.50 under zero-shot generation. ...

September 25, 2026 · 7 min · Zelina
Cover image

Train the Graph Before You Query It: SelfGraphRAG Turns Structure Into Supervision

TL;DR for operators An internal document collection can contain the relationships needed to answer difficult questions while still lacking the labeled examples needed to teach a retriever which relationships matter. That usually leaves teams choosing between manual annotation and retrieval based mostly on embedding similarity. SelfGraphRAG1 tests a third option: build a knowledge graph, turn its structure into generated question-answer examples, and train the retriever on those examples. On MultiHop-RAG, the resulting system reports F1 of 24.62, compared with 2.60 for RAG, 0.98 for LightRAG, and 0.01 for GraphRAG. ...

September 19, 2026 · 7 min · Zelina
Cover image

The Right Answer Is Not Enough: Verify the Reasoning Before You Train on It

TL;DR for operators A synthetic reasoning trace can end with the correct answer and still contain intermediate steps you would not want a model to imitate. ORACLE1 addresses that data-quality problem by checking reasoning one step at a time: it uses a symbolic reasoning engine when a step can be formalized, and LLM-based correctness and feasibility judgments when it cannot. ...

September 17, 2026 · 7 min · Zelina
Cover image

One Model, Two Routes: Why Audio-Omni Unifies Audio by Splitting Its Controls

TL;DR for operators A product that must understand an instruction, preserve a voice or reference sound, and stay synchronized with video is handling several different kinds of control. Audio-Omni’s strongest architectural evidence says those controls should not be forced through the same interface. The system combines a frozen multimodal language model with a trainable audio generator. High-level meaning and transcript information are supplied as flexible context; synchronization and acoustic-reference information are attached directly to the evolving audio representation. In the paper’s conditioning ablation, that allocation performs best across text-to-audio, video-to-audio, text-to-speech, and audio editing. ...

September 14, 2026 · 9 min · Zelina
Cover image

One Stack, Many Crossings: What AMD’s Real2Sim2Real Pipeline Changes for Robotics Infrastructure

TL;DR for operators Robot-learning infrastructure is usually discussed as if the central choice were the model or accelerator. The operational loop is broader: collect or generate experience, simulate behavior, train a policy, validate it, move it onto a robot, observe failures, reconstruct relevant environments, and repeat. Qing Yang and colleagues’ Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline1 is best read as evidence that many of those stages can be kept inside one ROCm + PyTorch-oriented software environment. The authors generate demonstrations in Genesis, fine-tune SmolVLA-450M, validate it in simulation, and deploy it to a physical Franka arm. They also connect real-scene reconstruction, synthetic-data generation, and reinforcement-learning workloads to the broader stack. ...

September 8, 2026 · 7 min · Zelina
Cover image

More Synthetic Data Wasn’t Better

TL;DR for operators Catalog teams often need more labeled examples before an attribute extractor works reliably in a new category or marketplace. Producing additional product-like text is not the difficult part. The training record has to change the intended attribute while leaving unrelated product information coherent. Negri, Martínez Gómez, Balanya, and Rajaram test a controlled generation pipeline built around that requirement.1 In their downstream experiment, synthetic-only training reached 60.48% attribute-extraction accuracy, almost identical to 60.79% for original-only training. But the best tested configuration was 75% original + 25% synthetic, at 68.82%. Accuracy declined as the synthetic share rose to 50% and 75%. ...

September 4, 2026 · 6 min · Zelina
Cover image

Three Loops Toward Self-Improvement

TL;DR for operators Model teams encounter three distinct constraints after a base model exists. New proprietary knowledge may remain accessible through retrieval without becoming reliably encoded in the weights. High-quality unique training text eventually becomes scarce even when additional training compute is available. Improvements to the training recipe still depend heavily on humans proposing, implementing, and testing experiments. ...

September 4, 2026 · 7 min · Zelina
Cover image

Wrong Code, Right Data Budget

TL;DR for operators If verified circuit-training data are scarce, exhaustive validation of every generated design may not be the best use of the next unit of data-engineering budget. In one experiment, a curated 22-circuit synthetic corpus reached 48.49% F1-Micro and 42.82% F1-Macro with a frozen circuit encoder, outperforming the original 22 verified circuits on both metrics and a 110-circuit raw generated corpus on F1-Micro. ...

September 4, 2026 · 7 min · Zelina
Cover image

Synthetic Data Can Make the Model Worse

TL;DR for operators A team with authoritative domain documents but little labeled training data has an attractive option: ask a capable model to manufacture question-answer pairs, then fine-tune a smaller open model on them. The operational risk is assuming that domain relevance makes those examples safe training material. In this paper, a simple synthetic-data pipeline moved LLaMA 3.1 8B backward on open-ended legal QA: its LegalMC4 score fell from 43.0% to 35.4%. A more structured pipeline raised the same score to 55.4%. Across both LLaMA 3.1 8B and Gemma 3 12B, that structured treatment improved all four tested German legal benchmarks.1 ...

September 3, 2026 · 7 min · Zelina