Skip to content
Latent Space

AI training pipeline is going fully synthetic, from reward signal to environment

[AINews] 10% worse, 100x cheaper, 10000x faster: Why Simulation is taking over

Latent Space traces how every component of the ML pipeline has flipped from human-made to model-made since 2022. The reward signal went synthetic first with InstructGPT's reward model, then Phi's textbook-quality synthetic pretraining data, followed by Alpaca-style distillation where a frontier model acts as teacher. Meta's self-rewarding models automated curriculum design in 2024, and Karpathy's autoresearch loop ran 700 overnight experiments in 2026, cutting GPT-2 training time from 2.02 to 1.80 hours. The latest step is Z.ai's GLM-5.3 synthesizing entire RL environments. The author frames this as '10% worse, but 100x cheaper and 10,000x faster human simulation.'

Why it matters: Latent Space connects 'models generating data instead of humans labeling it' into a traceable arc from 2022 to now, backed by specific papers and product milestones — not just trend talk. The ding is that this is a paid newsletter's Friday roundup, not a scoop or new release; ...

Read the original ↗Export Markdown