More Synthetic Data Wasn’t Better
TL;DR for operators Catalog teams often need more labeled examples before an attribute extractor works reliably in a new category or marketplace. Producing additional product-like text is not the difficult part. The training record has to change the intended attribute while leaving unrelated product information coherent. Negri, Martínez Gómez, Balanya, and Rajaram test a controlled generation pipeline built around that requirement.1 In their downstream experiment, synthetic-only training reached 60.48% attribute-extraction accuracy, almost identical to 60.79% for original-only training. But the best tested configuration was 75% original + 25% synthetic, at 68.82%. Accuracy declined as the synthetic share rose to 50% and 75%. ...