Skip to content
Hacker News front page

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

This paper pushes Zero RL—RL with verifiable rewards and no human labels—to 1 trillion parameters. Naive scaling makes reasoning traces bloated and hard to read, so the team adds clipped importance sampling, training-inference ratio correction, and mixed-precision control to stabilize training. Three findings: 1T parameters sharply improve sample efficiency and performance ceilings; training moves from a discovery phase to a sharpening phase; the model spontaneously develops anthropomorphism, self-verification, parallel reasoning, structured formatting, and even 'context anxiety,' making hand-crafted heuristics unnecessary. Ring-2.5-1T-Zero is competitive on seven math benchmarks. They also propose a three-dimensional CoT quality framework—comprehensibility, reproducibility, efficiency—where their model shows clear advantages in structured, concise traces. The post does not disclose specific benchmark scores, training cost, or open-source plans.

Why it matters: Empirical report on trillion-parameter Zero RL, directly addressing community concerns about scaling stability. Score held at 78 because the author team isn't a known major lab and the post doesn't disclose model architecture or training cost.

Read the original ↗Export Markdown