LEGO-Anything tests agents on 3D reconstruction and self-scoring
What happened
On October 3, The Decoder reported on LEGO-Anything, a study in which coding agents generate runnable Blender scene programs from a single image to rebuild 3D scenes, then judge how accurate the rebuild is. The LEGO-Bench evaluation used 208 images from 104 scenes. The best performer, GPT-6 Astra, reached 53.4% reconstruction accuracy indoors and 39.6% outdoors. The agents' geometric self-assessment landed at or below chance, so they can build a 3D reconstruction but cannot reliably tell how good it is.
Written by AI from the coverage · updated 1 hour ago
Coverage
Follow the reports to see the story from different sides.
- The DecoderAI agents build 3D scenes from photos but have no idea if they got it right
LEGO-Anything 让编码智能体从单张图片生成可执行的 Blender 场景程序,但测试发现智能体的几何自评接近或低于随机水平。LEGO-Bench 包含来自 104 个场景的 208 张图片,表现最好的 GPT-6 Astra 在室内与室外场景的重建准确率分别为 53.4% 和 39.6%。
Heat over time
Not enough continuous observations to draw a trend yet.