Randomized Trials Show AI Tutoring Effects Vary by Student, Subject and Instructional Design
A University of Maryland working paper published in October 2026 reports a randomized trial of 2,379 undergraduates and 30 instructors at a large public university, where sections offered access to a GPT-4o-based Virtual Study Assistant recorded final grades 0.37 standard deviations lower and learning-management-system participation 0.90 standard deviations lower, though only about 15 percent of students used it at least once. A September preprint, StudentBench, compared one-hour AI tutoring with human tutoring and a no-tutor control among 2,383 adults, finding statistically equivalent immediate gains on Quantitative GRE-style tests but not on Verbal, where human tutors' mean remained higher. A four-week Medly GCSE science study reported an effect of 0.33 standard deviations over self-directed revision, but 644 of 929 students completed post-testing, a 30.7 percent attrition rate.
These randomized trials provide a verifiable deployment warning for schools and product teams: offering access to an AI tool does not by itself guarantee improved learning, and effects depend on instructional design, student uptake and the task being tested. For education product practitioners, this means validating separately by subject and learner group and tracking unaided test performance and longer-term retention rather than relying on satisfaction or in-product activity. Limitations include preprint status for some studies, no long-term retention testing, no school-pupil classroom coverage, and industry funding for StudentBench. Schools should run controlled trials for specific subjects and learner groups before procuring or scaling AI tutoring tools, documenting the tested software version and teacher training. Product teams should distinguish access from sustained use and report separately who benefits and who disengages across language, disability, connectivity and prior attainment.