Ornith-1.0: Self-scaffolding LLMs that learn to build their own agentic coding harnesses
Ornith-1.0: Self-scaffolding LLMs for agentic coding
Ornith-1.0 is an open-source family of models for agentic coding, ranging from 9B dense to 397B MoE. The key idea: during RL, the model learns to generate its own task scaffolds instead of relying on hand-designed harnesses. The 397B variant hits 77.5 on Terminal-Bench 2.1 and 82.4 on SWE-Bench Verified, beating Claude Opus 4.7 and DeepSeek-V4-Pro. The 9B model runs on edge devices and scores 69.4 on SWE-Bench Verified, outperforming Gemma 4-31B. Reward hacking is handled in three layers: immutable environment boundaries, a deterministic monitor that blocks rule violations, and a frozen LLM judge as a veto. Asynchronous RL uses staleness weighting to manage off-policy tokens in long rollouts. The post does not disclose training cost, inference latency, or specific open-source license details.
Why it matters: Open-source agentic coding model with competitive numbers against Claude Opus 4.7 on Terminal-Bench 2.1 and SWE-Bench Verified. The self-scaffolding training method is a genuine innovation, not just another fine-tune. Score held below 85 because it's a new team with no third-p...