Skip to content
Hacker News front page

Four frontier models draw the Mona Lisa with colored pencils

"Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok

TryAI built a canvas arena where GPT-5.6 Sol, Claude Fable 5, Grok 4.5, and Gemini 3.6 Flash used colored-pencil tools to reproduce the Mona Lisa and Starry Night, plus five open-ended prompts. GPT-5.6 Sol scored highest on SSIM; Claude Fable 5 took the longest and cost the most while producing worse output. Grok 4.5 struggled, and open-weight models returned blank canvases. The authors argue fuzzy tasks like this separate frontier models from the rest better than benchmarks, and reveal real costs of long-running agent work.

Why it matters: TryAI built a drawing arena where GPT-5.6 Sol, Claude Fable 5, Grok 4.5, and Gemini 3.6 Flash used colored-pencil tools to reproduce famous paintings — 28 drawings total, with cost and structural similarity scores. It's a rare hands-on test of tool use + visual feedback loops,...

Read the original ↗Export Markdown