← All articles

The AI Tutor That Had to Get Past a Human First

In 3,617 messages, an AI maths tutor made five factual errors. The setup behind that number points to a different way of using AI with students.

Here is a number worth pausing on if you worry about AI making things up: in 3,617 messages to teenagers about maths, an AI tutor made five factual errors. That is about 0.1 percent. It also never sent a message that made its human supervisors uneasy about safety.

Before you read too much into that, look at how the trial was built. This was not an everyday chatbot. The AI had detailed information about each student, and its replies passed through human tutors before students saw them.

What was the experiment?

A randomised controlled trial in Britain involved 165 secondary school students aged 13 to 15, working on maths problems on the Eedi platform, according to reporting by The 74 (The 74). Eedi, an edtech company, put a small group of expert human tutors in charge of a large language model, Google's LearnLM. When a student needed help, the AI drafted a reply. Before that message went out, a human tutor had a chance to revise it to the point where they would feel comfortable sending it themselves.

The students did not know whether they were talking to a human or a chatbot. Think of it as an AI drafting for a human editor, with the student on the other end of the conversation.

What did the researchers find?

Students working with the supervised AI did slightly better than those who chatted online with human tutors alone. They solved new kinds of problems on later topics 66.2 percent of the time, compared with 60.7 percent with human tutors. They also had longer conversations on average with the supervised AI and human combination than with a human alone.

The safety and accuracy results were the headline. The AI "hallucinated," meaning it stated factual errors, in five of 3,617 messages. Human tutors approved about three out of four of the drafted replies with few to no edits. And students who got the combination of AI and human tutoring corrected misconceptions and reached correct answers over 90 percent of the time, compared with about 65 percent when they got a static, pre-written response to their questions. The researchers called the AI "a reliable source" of instruction.

Why did it work so well?

Eedi's chief impact officer, Bibi Groot, pointed to one factor above the rest: the AI had detailed information about each student. It knew what topics they had covered over the previous 20 weeks, which ones they had struggled with and which they had mastered. That guided its strategy about whether a student needed an extra push or just more support, something she said an "out-of-the-box, vanilla" model cannot do.

"They don't know anything about what the teacher is teaching in the classroom," she said of general chatbots. "They don't know what misconceptions or what topics the students are struggling with and what they've already mastered." She also made a point about people. Even excellent human tutors cannot read a student's full course history before a session, and they are under pressure to reply quickly, which can lead to cognitive overload. An AI can read all of that context almost instantly.

So the two halves complement each other: the AI brings memory and speed, and the human brings judgement and a safety check. The trial cannot say which piece mattered most, because it had no arm with the AI working alone. Groot's own account credits the detailed student data for how the AI performed, and treats the human tutors as the safeguard around it.

What does this mean for teachers?

It suggests a different picture of AI in tutoring from the one most families imagine. Instead of an AI replacing a tutor, it can act as a front line, with humans stepping in when a student is derailing the conversation or holds a misconception the AI cannot dislodge. That is how Groot described it. It also shows what "reliable" may take: not just a good model, but good information about the student, and a person checking the output.

For a classroom teacher, the practical lesson is about where checking happens. An everyday chatbot has no reviewer between it and your students. If you are considering an AI tool for students, ask who checks what it says, what it knows about your class, and what happens when it gets something wrong. Our procurement scorecard turns those questions into a checklist.

What should you be careful about?

Three cautions apply. The trial was run by Eedi, the company whose platform it tested, which is a reason to read the results as encouraging and not conclusive. It was small, with 165 students, and it covered maths for teenagers in Britain, so it may not carry over to other subjects, ages or settings. And what it tested was a supervised system with detailed student data, not a general chatbot, so it says little about how reliable an everyday chatbot is on its own.

There is also a question of scale. Eedi employs about 25 tutors across several time zones, available from 9 a.m. to 10 p.m. every day, and Groot noted that giving students broader access would require hiring "an army of tutors." A model that needs a human to approve every reply is safer, but it is also more expensive to run than one that does not.

Frequently asked questions

Does this mean AI tutors are accurate? It means a supervised AI tutor with detailed student data was very accurate in this trial. A general chatbot without a reviewer or that data is a different thing.

Did students know they were talking to an AI? No. In the trial, students did not know whether they were talking to a human or a chatbot.

Could a teacher supervise an AI like this? In principle, but the trial used a small team of expert tutors and a platform built for the task. It is not a description of what a teacher can do alone in class.

Should schools wait for more evidence? It is reasonable to want independent, larger studies. In the meantime, the trial is a useful guide to what to ask vendors: who reviews the output, and what data the tool uses.

Related guides