This piece is worth your time because it connects two experiments into one clean argument: AI lifts the floor of student work while present, but once removed, the borrowed floor vanishes, and rubrics still reward conformity over depth.
In the Turkish high school math RCT published in PNAS, students using standard ChatGPT scored 48% higher during practice, then dropped 17% below the control group on a closed-book exam taken minutes later. Server logs showed they were copying answers without reading the reasoning. The Milan business writing study from OpenAI's economics team found ChatGPT boosted scores by nearly a full point on a 5-point scale, pushing novice writing close to expert benchmarks. But here's the twist: students who spelled out failure conditions and causal mechanisms received systematically lower grades. The rubric rewarded smooth, conventional answers and penalized deeper thinking.
I'd discount the Milan study slightly since it comes from OpenAI's own economics team—read with that in mind. The Turkish study is more independent, published in PNAS. Together they point to the same problem: you get what you measure. If rubrics only reward polished conformity, AI just makes mediocre output cheaper to produce.
What's missing: long-term retention data. The article itself notes that no RCT has yet shown that scaffolded AI tutoring produces gains on unassisted exams. The strongest positive evidence so far comes from AI helping human tutors, not replacing them.