AI Tutoring Got 918x Cheaper. Measure What Matters

A new preprint says AI tutors now match expert humans on one-hour learning gains. That settles the cost question. It does not settle the question schools actually face.

Published September 29, 2026 • Jeff Katzman • 5 min read

On September 23, a research team funded by Handshake AI Research posted a preprint with a headline number that will end up in a lot of vendor pitch decks. Across 2,383 participants, AI tutoring produced GRE learning gains that were statistically equivalent to expert human tutoring. The cheapest AI tutor that passed the equivalence test cost $0.0052 per percentage point gained. The human tutor cost $4.81. That is a 918-fold difference.

The study, called StudentBench, is careful work. It logged more than 175,000 student-AI messages, randomized participants into AI tutoring, human tutoring, or no tutoring, and published its data and code. In five of seven GRE domains, the best-performing AI tutor beat the human tutor on average. Anyone who still believes AI tutors cannot teach should read it.

But the authors themselves draw a narrow boundary around the result, and school leaders should draw it too. The finding applies to GRE preparation, over a single one-hour session, with motivated college-age volunteers. Human tutors still led on verbal reasoning. That boundary is where the real work in AI in education begins.

The Cost Question Is Mostly Answered

For a decade, the central objection to high-dosage tutoring was price. Human tutoring works, but few districts or colleges can afford it at the scale their students need, especially now that stimulus funding is gone. StudentBench is one more data point, alongside a Harvard and MIT randomized trial published last year, suggesting that a well-designed AI tutor can deliver meaningful gains in a controlled session for pennies.

Key insight: A cost-per-gain number is only meaningful for students who actually use the tutor. If half of them never log in, the real cost per gain doubles and the students who most need help get none of it.

That caveat is not hypothetical. In August, Education Week reported on a Stanford SCALE Initiative research brief that reviewed four models of tutoring, from fully human to fully AI. In the limited research on AI-only tutoring, 40 to 47 percent of students never engaged with the platform at all. The brief summarized it bluntly:

"Many students don't engage with AI tutors. Adding human oversight helps engagement, but not achievement."

StudentBench participants were recruited volunteers who chose to spend an hour studying for a high-stakes exam. That is precisely the population that does not have an engagement problem. A ninth grader in a required algebra course, or a first-generation community college student working thirty hours a week, is a different case.

Hour One Is Not Week Twelve

The other boundary is time. A one-hour session measures whether a tutor can explain well and correct errors quickly. It does not measure whether a student returns on day nine, whether the gains survive a unit test three weeks later, or whether the tutor keeps a struggling learner in the productive discomfort of working a problem rather than extracting an answer. Those are the conditions under which most AI tutoring deployments succeed or fail.

The verbal reasoning gap matters too. Language-heavy reasoning is where students who read below grade level, and students learning in a second language, most need support. It is also where the human tutor kept an edge. Any institution serving multilingual learners should treat that as a design requirement, not a footnote.

Five questions to ask before trusting a cost-per-gain claim

  • Who was studied? Volunteers, or the full enrolled population, including students who opt out?
  • How long did it run? One session, or a full term with repeated use?
  • What share of students engaged? Cost per gain should be calculated on everyone assigned, not only active users.
  • Did gains persist and transfer? Delayed tests and unfamiliar problems reveal whether understanding stuck.
  • Who gained least? Results broken out by reading level, language, and prior achievement show whether gaps narrowed or widened.

Designing for the Students Who Do Not Show Up

If AI tutoring is now cheap enough to give every student, the scarce resource shifts from dollars to sustained engagement and faculty attention. That changes what a good tutor has to do. It needs to live where students already are, inside the LMS course they are required to open, rather than in a separate app they must remember to visit. It needs to adjust reading level so that a verbal reasoning problem is accessible without being simplified. And it needs to tell an instructor, early, which students have gone quiet.

That is the design logic behind the Core-LX AI platform. It uses a Socratic approach that keeps students in the question rather than handing them answers, adapts reading level while holding academic rigor constant, supports more than 150 languages, and deploys inside Canvas, Moodle, D2L, and Blackboard via LTI in hours. Continuous mastery tracking is designed to surface disengaged and struggling students before a traditional grade alert would.

We do not claim those features solve the engagement problem. That is exactly why our research pilot measures outcomes across full courses and full enrollments, including the students who never click, and reports results by subgroup. The next wave of AI tutoring research should look less like a one-hour benchmark and more like a semester.

StudentBench answered whether AI can tutor as well as a human for an hour. The question for every provost, superintendent, and CTE director now is whether it can keep the right students learning for a term. That answer will not come from a price per point.

Help Measure What Actually Matters

Join institutions studying AI tutoring across full courses, full enrollments, and every student subgroup.

Read the Full Article

Read "StudentBench: AI and human tutoring yield equivalent GRE learning gains" on arXiv

Share Your Thoughts

#AITutoring #AIinEducation #EdTech #EdResearch #StudentEngagement #EquityInEducation #HigherEd #K12

About Core Learning Exchange: We provide turnkey Career and Technical Education (CTE) solutions for grades 6-14, offering 450+ courses from 20+ providers aligned to state standards and industry certifications. Our AI platform uses proven Socratic methodology to develop critical thinking skills through personalized, adaptive learning—deployed in hours via LTI integration.