AI tutor slows math practice but improves post-mistake recovery in 6,997-student trial

Delayed-test gains were modest and concentrated among students using AI inside a three-correct-in-a-row mastery workflow in Tennessee

A student faces a blue digital display in an image illustrating AI-supported learning. The research examined how a guard-railed AI tutor affected math practice and delayed learning.

The study tested an embedded AI math tutor with middle school students in Hamilton County Schools, Tennessee

A randomized trial involving 6,997 US middle school students has found that a guard-railed AI math tutor changed what happened after students made mistakes, increasing next-attempt accuracy and reducing the number of attempts needed to get back to a correct answer. The trade-off was time: students assigned to AI worked more slowly and reached less material during a fixed class period, while the clearest evidence of learning a week later remained modest.

The study, Making AI Tutoring Productive: Evidence from a Mastery-Based Math Practice Experiment, was conducted in Hamilton County Schools, which includes Chattanooga, Tennessee, across 20 schools and just under 100 teachers. Students in grades 6 to 8 used the purpose-built NUMI platform during a regular math class in March 2026, followed by a delayed assessment approximately one week later. Of the 6,997 students in the initial analysis sample, 6,327 took the delayed test.

Every student received instructional videos, math practice, immediate feedback and worked solutions through a computer-assisted learning platform. Half were additionally assigned access to NUMI's AI tutor. That makes the study a test of what an LLM-based tutor adds to an existing digital learning environment, rather than a comparison between AI and conventional classroom teaching.

Students were also independently assigned to either a mastery or non-mastery workflow, creating four main combinations of AI, no AI, mastery and no mastery. In the mastery condition, students had to spend at least one minute on the relevant instructional video and answer three exercise questions correctly in a row before progressing. A mistake reset the streak. Non-mastery students could move on after attempting three questions, whether or not they had reached the same threshold.

The researchers were testing whether the way AI was embedded into practice changed how students used it, particularly when a progression rule forced them to keep working after an error.

Writing on LinkedIn after the paper was released, co-author Alp Süngü said: “Our findings suggest that the potential of AI tutors may depend on how the technology is operationalized in schools, not the technology alone.”

AI made error recovery more accurate, not faster

The strongest result appears immediately after mistakes. Among mastery students who made an observed error in Exercise 1, assignment to AI increased the probability that the next attempt was correct by 8.5 percentage points. It also reduced the number of additional attempts required to reach the next correct answer by 0.96 attempts.

But those students did not recover more quickly in clock time. AI increased the time to the next correct answer by 2.88 minutes.

That combination sits at the heart of the study: the tutor appears to have replaced some rapid retrying with a slower process of working through the problem.

The authors are careful about how far that result can be taken. Because assignment to AI also changed whether and when students made mistakes, the post-error sample itself is affected by the treatment. The paper therefore presents the post-mistake results primarily as evidence about the mechanism through which the tutor operated, rather than treating them as the main causal estimate of learning.

The broader practice data tell a similar story. AI students were less likely to reach later Exercise 1 questions, but they also made fewer observed mistakes. In the non-mastery group, AI reduced the number of Exercise 1 questions completed by 1.04 and increased exercise time by 1.64 minutes. Inside the mastery workflow, progress slowed further.

Süngü put the trade-off more directly in his LinkedIn post: “AI improved post-mistake accuracy! But it also slowed students down, highlighting the trade offs AI-infused learning brings.”

The NUMI activity was designed to fit into roughly 50 minutes, so extra time spent working through an early misconception could mean less exposure to Exercise 2 or the exit ticket.

It is a familiar tutoring problem in a new form. A tutor who asks a student to explain their reasoning will often take longer than simply displaying a solution. The question tested here was whether the additional time translated into better understanding.

NUMI's usage data suggest students were interacting with more than a background chatbot. Among students assigned to both AI and mastery, students used the “Help me get started” feature an average of 1.99 times, experienced 2.28 post-mistake walkthroughs and requested 3.23 step-specific explanations. Around 42% typed at least one substantive math-related message to the tutor.

Mastery changed behavior far more than delayed learning

The mastery condition produced a much larger immediate change in student behavior. Among students using the CAL-only platform, mastery increased the number of Exercise 1 questions completed by 4.62, increased the number answered correctly by 1.23 and raised the probability of reaching three correct answers in a row by 28.7 percentage points.

It also added 6.69 minutes to Exercise 1. On the platform, in other words, mastery appeared to work. Students practiced more and were considerably more likely to clear the threshold the system defined as mastery.

One week later, however, that advantage largely disappeared. The researchers found no clear delayed-test improvement from the mastery condition on its own. The large increase in three-correct-in-a-row attainment did not translate into a detectable increase in retained learning.

The researchers point to several possible explanations, including repeated exposure to similar questions, lucky streaks, guessing or knowledge that was present during practice but not retained. The experiment cannot distinguish between those explanations. What the randomization does show is that increasing the probability of clearing the platform's mastery threshold was not sufficient by itself to raise delayed-test scores.

The more encouraging result came when AI was combined with that structure.

Among mastery students, 37.0% of CAL-only students correctly answered the delayed Exercise 1 question covering material they had practiced. The figure was 40.2% for students assigned to AI, a 3.2 percentage point difference.

The result is only marginally statistically significant, with a p-value of 0.065, so the authors describe it as suggestive rather than definitive. It is also far from a transformational result. A majority of students in both groups still answered the delayed question incorrectly.

The corresponding unpracticed Exercise 1 result was almost unchanged, at 33.9% for CAL-only students and 34.1% for AI students. That distinction strengthens the interpretation that the small difference on the practiced question was connected to the material students had worked on rather than to a general difference between the two groups.

The full experimental results point in the same direction but are uneven. The interaction between mastery and AI was positive for performance on the two practiced delayed-test questions, while the broader practiced-versus-unpracticed measure remained too imprecise to establish a large or consistent learning effect across the assignment.

The appendix also finds a larger practiced Exercise 1 difference for the more difficult topic cells, where AI students scored 5.9 percentage points above the CAL-only group. On the easier topic cells, the difference was effectively zero. Those subgroup results are supportive rather than primary evidence, and the paper does not find statistically significant differences in the AI effect across the demographic and prior-achievement subgroup comparisons it tests.

NUMI was designed to stop AI becoming an answer key

The results also need to be read in the context of the specific AI system students were using. NUMI did not provide unrestricted access to a general-purpose chatbot. The tutor was embedded beside each math problem and designed not to reveal final answers.

Before attempting a problem, students could use “Help me get started,” which introduced the first-step approach and then presented a short reasoning question with two possible answers. The feature was withheld for the next two attempts after use.

Following an incorrect answer, the tutor could identify a likely misconception and move into a step-by-step walkthrough, checking understanding through yes/no questions, two-option prompts, multiple-choice questions or short answers. Students reviewing a worked solution could also select individual steps and ask for a concept-focused explanation.

The platform added another layer behind those interactions. Incoming messages were screened for safety and intent, including crisis content, off-topic messages, frustration, profanity and attempts to obtain an answer directly. Draft tutor responses were then checked for mathematical correctness, formatting and informativeness, with low-scoring responses rewritten before delivery.

The CAL-only version removed the AI chat, hints and walkthroughs but retained the same instructional videos, practice questions, feedback and worked solutions. That distinction limits how broadly the findings can be generalized. The study tests a specific, guard-railed tutor inside a specific learning workflow, not open-ended student use of an LLM.

The study has other significant limits. Students completed one NUMI assignment, not a term or year of AI-supported learning. The delayed assessment contained only four questions, creating relatively noisy measures of retained learning. Treatment also affected whether students reached later material, making Exercise 2 particularly difficult to interpret.

The NBER paper is also a working paper circulated for discussion and comment. It has not been peer reviewed or reviewed by the NBER Board of Directors as an official NBER publication.

The researchers identify repeated use over several weeks or units as a natural next test. In this experiment, the intervention consisted of one approximately 50-minute math practice session between March 23 and 27, 2026, followed by a 15 to 20-minute delayed assessment during the week of March 30 to April 3.

Next
Next

UC Irvine receives $10m federal grant for national AI writing research center