A small transformer trained on a 5090 hits 44% on ARC-AGI-1 for 67 cents
44% on ARC-AGI-1 in 67 cents
Mithil Vakde trained a small transformer from scratch on a 5090 in 1.5 hours for 67 cents, scoring 44% on ARC-AGI-1 public eval—matching TRM/HRM—and 7% on ARC-2. The method converts each puzzle into token sequences, uses 3D RoPE and per-task learnable embeddings for cross-task learning, and applies test-time augmentations with voting. Switching to a modern architecture (SwiGLU, RMSNorm) and using fewer augmentations drove the gains and cut costs. Training only on output tokens lifted the score from 40% to 44%, which the author doesn't fully understand yet. Code is open source; the union of solved tasks across runs reaches 55%, and the author sees room in better position embeddings and architecture tweaks.
Why it matters: 44% on ARC-AGI-1 for 67 cents and 1.5 hours on a single 5090 — the numbers carry the story. Architecture details (3D RoPE, per-task embeddings) give a reproducible hook, not just talk. ARC-2 at 7% is the hard gap keeping it below 80.