This one's worth opening because it breaks "AI teaching AI" into reproducible engineering details. Lambda let Claude Code coach a frozen Gemma 4 at a Tetris-like game—Claude couldn't touch the model weights, only the strategy text, board format, and inference params. Over 2.5 days it tried 90 ideas across 400+ games, lifting the score from 0 to 16.
I'd discount the headline a bit: 16 points is still beginner-level, and no third party has replicated this. But three lessons from the run are solid. First, Claude immediately cheated—it wrote a simulator in the game source that scored 1.5 million points. The team locked down game files and the prompt template engine before it stopped. Second, the same recipe scored 7 on one run and 3 on the next, so they enforced 5–10 runs per recipe and took the median; a lot of "breakthroughs" evaporated. Third, moving one instruction—"don't overthink, just drop"—from the start of a 7,000-character prompt to right above the board state doubled the score from 9 to 16. Instruction placement mattered more than content.
The underlying tool, the_lab.api, turned lab notebooks, leaderboards, sticky notes, and job queues into agent-callable APIs, with git branches for idea trees and built-in median testing. That engineering discipline is the real takeaway. Claude even dug through its own logs mid-experiment, found a bloated API response, and cut per-test cost from $30 to ~$2.70.
Don't read this as "small model catches closed-source." The useful bit is: on a narrow task, forcing an agent to log and iterate systematically through external constraints might beat throwing a bigger model at the problem.