An AI Tutor Lowered Grades. Blame the Default Setting
The first large randomized trial of a course-integrated AI tutor at a U.S. public university found lower grades and less engagement. The most important detail is a setting almost nobody changed.
Published October 6, 2026 • Jeff Katzman • 5 min read
On September 30, researchers at the University of Maryland released the results of a full-semester randomized controlled trial that every campus rolling out an AI tutor should read before the spring term. Across 2,379 undergraduates and 30 instructors, students in sections randomly given access to a course-integrated AI tutor finished the semester with lower final grades, by 0.27 to 0.37 standard deviations. Among sections of the same course, the drop was 0.37. Their recorded participation in instructor-designed online activities fell by 0.90 standard deviations.
This is not a story about a rogue chatbot. The tool was a GPT-4o-based virtual study assistant built into the university learning management system, grounded in instructor-selected course materials through retrieval-augmented generation. On paper, it was exactly the kind of responsible, course-aligned deployment many institutions are planning right now. That is what makes the result worth taking seriously.
What the Trial Actually Found
The study, published as an Annenberg Institute EdWorkingPaper, randomized instructors within blocks of identical or similar courses spanning STEM, the social sciences, the humanities, and education. That design matters: it compares students taking the same course in the same term, with and without the tutor.
The negative effects were not limited to grades. Students with access visited course pages less, logged fewer active days, and completed fewer discussion responses, assignments, and quizzes. Homework scores showed null or negative effects, not the short-term boost earlier studies have reported. Survey responses suggested tutor users asked their instructors fewer questions about course content.
"These findings suggest that introducing an AI tutor may displace engagement with existing learning activities without improving academic performance, with potentially differential impact across student groups."
That last phrase deserves attention. The estimated final-grade losses among first-generation students were more than twice those of continuing-generation students. The students institutions most hope AI will help appear to be the ones most exposed when it is configured poorly.
The Setting Nobody Changed
Here is the detail that should reshape procurement conversations. The tool offered instructors two modes: a direct-instruction mode, which was the default, and a tutor mode that guides students toward a correct response rather than handing it over. Few instructors switched. According to the authors, only two ended the term with tutor mode in place.
Key insight: In practice, the default setting is the pedagogy. If an AI tutor ships in answer mode and relies on busy faculty to switch it, most students will get an answer engine, not a tutor.
Student behavior followed the setting. Roughly 73.8 percent of requests sought information, explanations, or solutions. About 11.3 percent involved practice or test preparation. Just 0.7 percent asked for feedback on the student's own work. The authors are careful to say these categories do not prove cognitive offloading, but the pattern points toward passive, on-demand help rather than effortful practice.
Fairness to the study requires two caveats. Only about 15 percent of students in treatment sections used the assistant even once, so the authors note that introducing the tool may have shifted how students used other AI, such as ChatGPT and Gemini, which all students could access. And the researchers explicitly limit their conclusions to this assistant as configured, not to more scaffolded AI tutoring. That boundary is the point. Configuration is the variable.
The Research Is Converging
The Maryland result fits a pattern. A 2025 high school math study by Bastani and colleagues found that general-purpose AI access hurt later independent performance, while tutoring guardrails mitigated that harm. A large Tennessee middle school trial found gains only when the AI tutor was paired with a mastery requirement, a finding we examined in an earlier post. Access alone does not teach. Structure does.
Five questions to ask before turning on an AI tutor
- What is the default mode? Guided questioning should be on out of the box, not an opt-in setting.
- Does it require student effort? Look for prompts that ask students to attempt, explain, and revise before help escalates.
- Does it route students back to the course? A good tutor sends learners to readings, activities, and instructors rather than replacing them.
- Who is tracking engagement? Faculty need early signals when participation drops, not end-of-term grades.
- Are results reported by subgroup? First-generation, multilingual, and lower-prior-achievement students must be measured separately.
Socratic by Default
The lesson is not to keep AI out of courses. Students already have unrestricted chatbots in their pockets. The lesson is that an institutional AI tutor has to be designed to keep students in the question, and that design cannot depend on every instructor finding the right toggle.
That is why the Core-LX AI platform uses Socratic questioning as its core behavior rather than an optional mode. It is designed to ask students to reason before it explains, to adapt reading level while holding rigor constant, and to support more than 150 languages so multilingual learners can engage in depth. It deploys inside Canvas, Moodle, D2L, and Blackboard via LTI, and continuous mastery tracking is built to flag students whose engagement is slipping before a grade does.
We do not assume that design works. The Maryland trial shows why no vendor should. Our research pilot measures full-course outcomes, including engagement with existing course activities and results by subgroup, so partner institutions can see whether Socratic AI tutoring complements their courses or quietly displaces them.
One semester at one university does not settle the question of AI tutoring. It does settle something narrower and more useful: the defaults are not neutral. Every institution deploying AI this year is choosing a pedagogy, whether it realizes it or not.
Deploy an AI Tutor That Asks Before It Answers
See how Socratic-by-default AI tutoring works inside your existing LMS, or join our research pilot to measure its impact on your students.
Read the Full Article
Share Your Thoughts
#AIinEducation #AITutoring #HigherEd #EdTech #SocraticMethod #FirstGenStudents #EdResearch
About Core Learning Exchange: We provide turnkey Career and Technical Education (CTE) solutions for grades 6-14, offering 450+ courses from 20+ providers aligned to state standards and industry certifications. Our AI platform uses proven Socratic methodology to develop critical thinking skills through personalized, adaptive learning—deployed in hours via LTI integration.
Related Posts
AI Tutoring Only Worked When Students Had to Prove It
A 6,000-student Tennessee trial found AI tutoring helped only when paired with a mastery requirement.
AI Can Boost Scores and Still Fail to Teach
Why practice-score gains from AI can mask a drop in durable learning, and what purpose-built tutoring does differently.
AI Tutoring Got 918x Cheaper. Measure What Matters
AI tutors now match humans for an hour. The harder question is who keeps learning for a full term.