This is worth opening because Mollick isn't running benchmarks—he's throwing models into messy, real-world tasks to see what they actually do. He had GPT-6 Astra turn the 1977 text adventure Zork into a playable 3D action game, with the AI deciding what the white house and monsters look like and converting 'fight the troll' into an action sequence. In a second experiment, Fable 5.1 reconstructed Umberto Eco's private library from a dozen videos, shelf photos, and two catalogues—no floor plan available—placing roughly 5,000 identified books across 27,000 shelf slots. He also had Astra operate Blender and produce a 3D animated book trailer in 45 minutes.
His core argument is 'capability overhang': these models can already do weeks of human work, but almost nobody taps that potential. He frames four human advantages to close the gap—deep knowledge, wide knowledge, taste, and agency. That framework is more useful than vague 'AI will replace you' talk.
I'd discount this a bit: the post doesn't say when GPT-6 Astra or Fable 5.1 launched, whether they're publicly available, or give any specs. These experiments read more like a power user stress-testing frontier models, not something the average person can replicate today. But the direction is right—dig into what current models can do before obsessing over the next release.