Randomized Trial of 6,997 Students Finds NUMI AI Tutor Improves Next Answers but Slows Practice and Shows Limited Delayed Gains
A randomized trial involved 6,997 grade 6–8 students in 20 Hamilton County, Tennessee schools, with 6,327 taking a delayed test one week later. All students used NUMI practice, feedback, and explanations; half also received an AI tutor and were assigned to mastery or non-mastery workflows. After an error in the mastery condition, the AI group was 8.5 percentage points more likely to answer the next item correctly and used 0.96 fewer attempts but took 2.88 minutes longer. Delayed accuracy on practiced items was 40.2% for the mastery-plus-AI group and 37.0% for controls; with p=0.065, the paper calls the evidence suggestive and has not yet been peer reviewed.
The trial measures immediate recovery after errors, practice speed, and delayed performance in one design, showing that better next-answer accuracy carries a time cost. Delayed gains were small and statistically limited, and a single session of about 50 minutes cannot establish long-term learning effects. Teachers using this kind of tutor must balance explanation depth with problem coverage in a fixed lesson and watch for students spending too long. Product teams should report next-answer accuracy, practice time, and delayed testing together because immediate accuracy alone omits a key cost.