Ninety-six percent of the students messaged the AI tutor at least once. Then most of them rarely used it.
That is the short version of one of the first large-scale experiments on AI tutors in real schools: two years, 18 Tennessee middle schools, and an AI tutor built to do what good human tutors do, which is to ask questions instead of handing over answers. The tutor was carefully designed. The students it was designed for mostly ignored it.
The story of why is more useful than any verdict, and it matters if you are choosing, buying or recommending an AI tutor. It starts in 2023, when Khan Academy released Khanmigo, a "Socratic-method" chatbot that withholds answers, gives hints and guides students with questions. Soon after ChatGPT arrived, students had worked out how to use it as an answer machine, and the shortcut showed up in lower test scores. Khanmigo was the fix. In August 2026, researchers published some of the first large-scale experimental evidence on whether the fix worked.
What did the researchers actually do?
Philip Oreopoulos of the University of Toronto and Nina Low ran a two-year cluster randomised trial in 18 Tennessee middle schools, from 2024 to 2026 (Hechinger Report; working paper). The students were low-achieving, at least one grade level behind their peers, and their schools had built an extra remedial maths period into the school day. Some students were randomly assigned to use Khan Academy, with Khanmigo available as an AI assistant configured to coach rather than give answers. The others got their usual remedial maths work, which often included online tools such as Waggle, IXL, Zearn and DeltaMath.
That design matters. Because assignment was random and the sessions were part of the school day, the researchers avoided a problem that undermines many edtech studies, where only motivated students opt in. Nearly everyone assigned to the tutor had it available, and in the second year the researchers could see how students used it.
What did they find?
The students who received the Khan Academy intervention did better in maths than those in regular remediation. The gain was modest: about 1.3 national percentile ranks per term, or roughly 0.06 to 0.08 standard deviations over a school year, and about 0.14 standard deviations for a full year of active participation. In other words, Khan Academy appeared to help, and the gains looked like those from Khan Academy practice without an AI assistant. As the Hechinger Report put it, "the Khanmigo tutor didn't add anything."
The researchers' explanation is the most useful part of the study. Message-level records are available for the second year, the year of strongest implementation and largest effects, and they show that access was nearly universal but engagement was thin. Ninety-six percent of students messaged Khanmigo at least once. Yet the median student sent it no messages on two-thirds of the days they practised, and messaged it in only 17 percent of the exercise sessions in which they made a mistake.
The messages students did send were mostly one- or two-word answers to Khanmigo's own questions, or clicks on suggested prompts. Only 14.5 percent contained a mathematical question or a step of reasoning.
Why did students stop using it?
According to Hechinger, many students tried at first to get the AI to give them the answer. When Khanmigo refused and asked questions or offered hints instead, they largely stopped using it. The paper's authors wrote that "seeking help with one's own confusion remained a choice, and most students declined it most of the time."
That is an uncomfortable finding for anyone who believes the answer to answer-machine chatbots is a Socratic one. The design that protects learning, which asks questions and does not hand over answers, is also the design that a student who wants to finish quickly has the least reason to use. A tutor that is good but unused helps nobody. The paper's conclusion is blunt: the binding constraint appears to be engagement, and realising the promise of AI tutoring will require getting students to use it, not just giving them access.
What did Khan Academy say?
Sal Khan praised the study and wrote a detailed post about it. In an interview with Hechinger he said the findings confirmed what Khan Academy was already seeing internally, that students were not engaging much with Khanmigo. He said he did not regret releasing the tutor before knowing whether it would be effective, and was confident the early release did "no harm."
The tutor tested was the version released in 2023, and Khan Academy has since changed it. According to the Hechinger report, Khanmigo is now built into Khan Academy instead of sitting in a separate tab. If a student asks for help before trying a problem, it offers little more than a hint and encouragement, and after a student gets a problem wrong it now pops up automatically to work through the problem together. Schools can also turn it off. Khan said the company plans to give students credit for using Khanmigo after a mistake, so that a redo with the AI can count toward the five correct answers in a row they need before moving on. Separately, Khan Academy's chief learning officer has said that only about 15 percent of students with access to Khanmigo regularly engage with it (EdTech Innovation Hub), which is consistent with the trial.
What are the limits of this evidence?
It is one product, tested in one setting: low-achieving middle school students in a remedial maths period, in one state. The version tested has since been redesigned, and this trial did not test the redesign. The study is a working paper circulated by the National Bureau of Economic Research and has not been through peer review. And it measures what happened when students were assigned to use the tool. It says little about students who are self-motivated and choose to use one.
It would also be a mistake to read the result as "AI tutors do not work." The study found that Khan Academy practice helped, and that the AI assistant did not add to it in this setting, mainly because students did not use it much. Hechinger's reporter put the point well: use and engagement are only one piece of the puzzle, and we still need evidence that when students regularly use an AI tutor, they actually learn more.
What should schools and families take from it?
The first lesson is to measure use from day one. Any pilot should track what share of students use the tool each week and how they use it, not only whether they like it. Our six-week pilot plan builds this in. The second is to look at what students actually type. A tool whose students mostly send one-word answers or click suggested prompts is not being used as a tutor, whatever its design intends. Ask a vendor to show you real sessions, and see our procurement scorecard for the questions to ask.
The third lesson is for parents. If your child is offered an AI tutor, watch what happens in the first sessions. Does your child attempt problems, respond to the tutor's questions and come back to it, or do they look for a way around it? Our guide to choosing an AI tutor for your child covers what to look for. The trial suggests that the best tutor is the one your child will actually talk to, and that being well designed is not enough.
Frequently asked questions
Does this mean AI tutors do not work? No. It means that in this trial, the AI assistant added little on top of Khan Academy practice, mainly because students used it sparingly. Whether higher engagement would produce larger gains is still an open question.
Was the tutor giving students wrong answers? The paper's finding is about engagement, not accuracy. The authors did not report accuracy as the problem.
Should schools avoid Khanmigo? The trial does not say that. It shows what happened with the 2023 version in one setting, and the product has changed. If you are considering any AI tutor, run your own pilot and measure use.
What would make students use a tutor more? Khan Academy has added an automatic prompt after mistakes and plans to give credit for redoing a problem with the AI's help. This study does not test whether those changes raise use.
Related guides
- A six-week pilot plan
- The procurement scorecard
- 8 questions before choosing an AI tutor
- Education News: latest research, a live feed of new education studies