Trending storyDeveloping
Epoch AI's InnovationEval: neither model independently matched SDPO
1 report1 sourceupdated 4 hours ago
What happened
AI digest
On October 10, Hacker News surfaced Epoch AI's preliminary InnovationEval results, which test whether AI can carry out AI R&D on its own. The run covered Claude Fable 5 and GPT-5.6 Sol and asked whether each could independently discover a machine learning innovation on par with the human-proposed SDPO. Neither model did so on its own. Only preliminary results are disclosed, with no further progress reported.
Written by AI from the coverage · updated 58 minutes ago
Coverage
Follow the reports to see the story from different sides.
Oct 10
- Hacker News front pageCan AI automate AI R&D yet?
Epoch AI 公布 InnovationEval 初步结果,Claude Fable 5 和 GPT-5.6 Sol 均未独立发现与人类提出的 SDPO 相当的机器学习创新。
Heat over time
Not enough continuous observations to draw a trend yet.