Skip to content
MIT Technology Review · AI

MIT TR: 7 puzzles where AI still flubs—can you beat them?

AI models flub these intelligence tests. Can you fare any better?

MIT Technology Review built an interactive quiz from seven puzzles that have tripped up frontier models. It cites Columbia University data: in late 2024 the best models solved only 18% of NYT Connections puzzles, but by early 2025 some reached near-perfect scores. Visual tasks remain a weak spot—LLMs still fail badly at mental rotation problems even with vision input. A 2024 study by Google and UIUC showed models get tripped by Knights and Knaves variants, defaulting to memorized answers instead of reading the twist; SimpleBench exploits the same pattern. The post does not disclose current model accuracy on these seven puzzles.

Read the original ↗Export Markdown