TutorMoments: a framework to test if AI tutors know when to help and when to hold back
TutorMoments: Do AI tutors know when to help and when to hold back?
Allen AI released a preview of TutorMoments, a replay-based evaluation that tests whether LLMs over-help when acting as math tutors. It uses real one-on-one tutoring transcripts, with experienced teachers flagging moments where a tutor must choose between scaffolding a problem and pushing the student to reason independently. When told only to 'tutor well,' models tend to give too much support and rarely push for deeper thinking. Prompting the trade-off explicitly improves performance but does not close the gap to human tutors who adapt to the moment. The project includes a de-identified transcript dataset, replay pipeline code, and model tutor replays.
Why it matters: Allen AI's TutorMoments benchmark uses real tutoring transcripts to mark moments where a tutor should step in vs. hold back, then tests models on those decisions. The finding that models over-help is concrete and counterintuitive — H and K are both present. But resonance is na...