Toronto and Wharton study: AI tutor plus three correct answers in a row yields small math gains for over 6,000 Tennessee middle schoolers
Researchers from the University of Toronto and the University of Pennsylvania's Wharton School ran a fractions experiment with more than 6,000 middle schoolers in Tennessee, randomly assigning students to conventional computer-based instruction or the same software with an AI tutor, with half of each group required to answer the same skill correctly three times in a row after a mistake. Students used the software once for 50 minutes during math class and took a 15-minute test a week later. The AI tutor plus mastery-learning combination scored about 3 percentage points higher than conventional computerized instruction, with the advantage concentrated on the easiest fraction questions most similar to what students had practiced.
This is one of the earlier randomized experiments testing an AI tutor combined with mastery learning, offering schools and education product teams verifiable short-term evidence that AI's value may come from making students slow down and review errors rather than showing answers. The advantage was only about 3 percentage points and appeared only on the easiest questions most similar to practice, not on harder problems; the intervention lasted 50 minutes, and the paper is a National Bureau of Economic Research working paper not yet peer-reviewed, so it cannot support claims about long-term learning or overall education quality. Schools and education product teams can treat error review and consecutive-correct thresholds as testable feature combinations rather than adding only an AI chat layer. Parents and teachers should treat the result as short-term, small, and limited to easy problems, not as evidence that AI tutoring broadly raises math ability.