Oral Exams Are Back, and AI Is Giving Them
A voice AI administered 73 personalized oral exams at NYU Stern for under a dollar each. The result is a working answer to the question every institution is now asking: how do you verify that a student actually understands something?
Published July 28, 2026 • Jeff Katzman • 4 min read
The take-home essay is finished as a measurement instrument. The Higher Education Policy Institute found this year that 94 percent of UK undergraduates now use generative AI to help with assessed work, and nearly two thirds say their assessments have already changed significantly in response. Institutions spent two years trying to solve this with detection software. Detection lost.
So a growing number of faculty have reached back for the oldest assessment format in the university: sit the student down and ask them to explain. Cornell, NYU Stern, and Penn all ran oral defenses during spring 2026 midterms. The logic is blunt and hard to argue with. You cannot outsource a conversation.
The problem is that oral exams are the reason we invented the written exam. They do not scale. A professor with 200 students cannot run 200 twenty-minute conversations, which is precisely why the format quietly disappeared from undergraduate education a century ago.
What the NYU Experiment Actually Did
Panos Ipeirotis and Konstantinos Rizakos at NYU Stern built a system that changes that arithmetic. A voice AI conducts a personalized oral exam, drawing on each student's own capstone project. The transcript is then graded by what the authors call a council of three large language models, which score independently and then revise after reading one another's assessments. They ran it across two undergraduate cohorts totaling 73 students, in fall 2025 and spring 2026.
Exams averaged 22 minutes, ranging from 9 to 41 minutes. Exam length had no relationship to performance. And the cost of grading came in under one dollar per exam.
The deliberation step is the finding. When the three models scored independently, agreement was poor: Krippendorff's alpha of 0.52, with one model averaging 16.3 out of 20 while another averaged 13.0. After the models read each other's reasoning and revised, alpha rose to 0.86 — above the 0.80 threshold conventionally treated as good reliability. Dimension-level scores reached 74 percent perfect agreement, with only 1 percent of ratings differing by two points or more.
The Students Were Not Thrilled, and That Matters
Here is where the study earns its credibility. The authors published the uncomfortable numbers alongside the encouraging ones.
Seventy percent of students agreed the exam "tested my actual understanding" — the highest-rated item on the survey. But 83 percent found it more stressful than a written exam, only 33 percent found the questions clear, and when asked to choose a format, 57 percent preferred traditional written exams while just 13 percent preferred the AI-administered oral. It is worth noting that 83 percent had never taken an oral exam in any course before. Unfamiliarity is doing some of the work in those numbers.
The behavioral evidence points a different direction than the sentiment data. Across 73 exams the instructors received two regrading requests and conceded not a single point.
Five Lessons Worth Copying
Any institution considering this should read the engineering section before the results section. The authors are specific about what broke:
What the Researchers Learned the Hard Way
- Constraints must be architectural. Behavioral limits on a language model have to be enforced through system design, not through prompting alone.
- Decompose the exam into phases. Separate agents per phase prevent conversational drift and cascading prompt failures.
- Use multi-model consensus. Independent scoring followed by deliberation measurably reduces single-model bias.
- Randomization requires code. Language models cannot randomize reliably; question selection needs deterministic mapping from a seed.
- Voice design is pedagogy. A cloned professorial voice was perceived as aggressive rather than reassuring.
They are equally direct about limits. Reliability is not validity. As the paper puts it, agreement among raters "does not rule out three models agreeing on a wrong grade." Students with speech-related disabilities or limited verbal fluency face a barrier that written formats avoid, and non-native speakers face compounded difficulty. The self-recorded format deters cheating without preventing it. Voice recordings pass through four third-party vendors, which is a data governance conversation, not a footnote.
Verification Is Not the Same as Teaching
The deeper issue is that a one-shot oral exam is still a snapshot. It confirms at the end of a term what an instructor probably needed to know in week three. If the reason to revive the oral exam is that dialogue reveals understanding in a way a submitted document cannot, then the logical conclusion is not a better final exam. It is more dialogue, distributed across the whole course.
Reliability (do raters agree?) is necessary but not sufficient... it does not rule out three models agreeing on a wrong grade.
— Ipeirotis and Rizakos, NYU Stern
This is the pedagogical bet behind our AI platform. Socrat is built on Socratic methodology, which means it keeps students in question space rather than delivering answers — the same mechanism that makes an oral exam diagnostic, applied continuously instead of once. Because every exchange is a reasoning exchange, mastery tracking accumulates process evidence all term long and flags a struggling student earlier than an end-of-term assessment ever could. Reading level adapts without lowering rigor, and support runs in more than 150 languages, which speaks directly to the accessibility concern the NYU authors raise about non-native speakers in a voice-only format.
That evidence has a second use. Employers do not want a transcript grade; they want to know what a graduate can demonstrably do. Continuous dialogue produces a record of demonstrated competency, which is the same logic behind aligning coursework to the 70 industry certifications in our catalog. Assessment stops being a gate and becomes a signal.
We are not claiming outcomes we have not measured. That is exactly what our research pilot exists to study, with institutional partners and anonymized data. The NYU team modeled the right posture here: publish the reliability figures, publish the 57 percent who preferred a written exam, and call an informal instructor comparison what it is rather than a validation study. Education technology would be in better shape if more of it were reported this way.
Oral exams did not come back because they are novel. They came back because they were always the better instrument, and until recently the labor cost made them impossible. Under a dollar per exam changes that. The question facing institutions now is whether they use the capability to build a harder final, or to make the entire course a conversation.
Make the Whole Course a Conversation
Socratic AI tutoring that surfaces student reasoning continuously, not once at the end. Deploys via LTI in hours.
Read the Full Article
Read "Scalable and Personalized Oral Assessments Using Voice AI" on arXiv
Share Your Thoughts
#AIinEducation #AssessmentRedesign #HigherEd #SocraticMethod #AcademicIntegrity #EdTech #CompetencyBasedEducation
About Core Learning Exchange: We provide turnkey Career and Technical Education (CTE) solutions for grades 6-14, offering 450+ courses from 20+ providers aligned to state standards and industry certifications. Our AI platform uses proven Socratic methodology to develop critical thinking skills through personalized, adaptive learning—deployed in hours via LTI integration.
Related Posts
AI Can Boost Scores and Still Fail to Teach
Practice scores rise while durable learning drops. Why the measurement instrument matters as much as the tool.
AI Did Not Destroy Critical Thinking. We Did.
Assignments that never surfaced student reasoning were vulnerable long before generative AI arrived.
Join Our AI Tutoring Research Project
We are recruiting institutional partners to study AI-enhanced learning outcomes with real data.