Students offered an AI tutor got lower grades and disengaged from course materials, University of Maryland trial finds
In a randomized study of 2,379 undergraduates, access to a GPT-4o-based tutor was linked to grades about four points lower in matched courses, while participation in the university’s learning platform fell sharply
A University of Maryland randomized trial found students offered access to a course-integrated AI tutor had lower grades in matched classes and used existing course resources less
A large University of Maryland trial has produced a striking result for universities rolling AI into teaching: students offered a course-integrated AI tutor performed worse in matched classes and engaged substantially less with existing course materials.
The randomized study involved 2,379 undergraduate students and 30 instructors across multiple disciplines during fall 2025.
In the clearest comparison, between sections of the same course, students given access to the AI tutor finished around four percentage points lower in their final grades, equivalent to 0.37 standard deviations.
Their use of the university’s learning management system dropped even more sharply. Recorded participation fell by 0.90 standard deviations, while page views and active days also declined.
One detail makes the result particularly interesting: only around 15% of students offered the tutor actually used it.
The researchers say introducing an approved AI tool into a course may also have changed how students approached other learning activities and their wider use of generative AI.
Most students used the AI for answers, not tutoring
The University of Maryland’s Virtual Study Assistant was built into the university learning platform and powered by GPT-4o.
It used retrieval-augmented generation, or RAG, to draw primarily on course materials selected by individual instructors. In practical terms, that meant students could ask questions inside their course and receive responses grounded in materials such as lecture content and readings.
But how students used it is revealing. Nearly 74% of requests were for information, explanations or solutions. Around 11% involved practice or test preparation, while fewer than 1% asked the tutor to give feedback on a student’s own attempt.
So although the tool was presented as a study assistant, students were much more likely to ask it for an answer or explanation than to use it as a back-and-forth tutor.
Jing Liu, Associate Professor and Director of the Center for Educational Data Science and Innovation and lead author of the paper, stressed on LinkedIn that the findings should not be read as evidence against all AI tutoring.
“These findings do not establish that purpose-built AI tutoring tools are not beneficial,” he wrote.
Instructors could choose between a direct instruction mode, which provided more explicit answers and explanations, and a tutoring mode designed to guide students toward an answer. Most retained the default direct instruction setting.
Students also used the course platform less
The most consistent finding across the trial was not grades, but changes in student behavior.
Students with access to the Virtual Study Assistant recorded fewer page views, fewer active days and substantially lower participation in the university learning platform.
That participation measure included activities such as discussion responses, assignment submissions and quizzes.
The researchers do not establish that lower platform use caused the lower grades. Students may have substituted some traditional course activity for AI-assisted work elsewhere. But the scale of the difference is one reason the study stands out.
Liu’s takeaway was less about the technology itself than what it may replace: “when evaluating an AI tool’s impact on learning, we need to consider what it might replace.
“Going to office hours, engaging with course materials, and asking instructors questions take time and may feel inefficient, but they are essential building blocks of learning.”
Survey responses also point toward reduced interaction with instructors, although the researchers caution that the student survey response rate was low.
The 15% uptake problem makes the findings harder to explain
If only around one in seven students used the tutor, why were differences between the treatment and control classrooms so large?
The researchers do not claim to have a definitive answer. Students in both groups already had access to other AI tools, including ChatGPT and Gemini. Introducing a university-approved AI tutor may therefore have changed students’ perception of how acceptable AI use was more generally, even among those who never opened the Virtual Study Assistant itself.
The trial measures the effect of offering access to the tool, not simply the effect on students who actively used it.
The study also identifies a potentially important equity issue. Estimated grade losses were larger among first-generation students than among continuing-generation students, although the researchers caution against overinterpreting subgroup results.
University of Maryland Vice President and Chief Information Officer Jeffrey Hollingsworth highlighted that point on LinkedIn, writing that “the negative impact was larger for first generation students than the overall population.”
There are other limitations. The study uses course grades rather than an independently administered assessment, students already had access to other generative AI tools, and most instructors did not use the more guided tutoring mode.
Taken together, the results do not show that AI tutoring inevitably harms learning.
They do provide evidence that simply placing a course-grounded AI assistant inside a learning platform does not guarantee better outcomes, and that how students use the tool, how instructors integrate it and what learning activities it replaces may be just as important as the technology itself.
Liu has said a second paper from a separate study is due next week.