An AI quiz maker can produce twenty questions in the time it takes to find a pencil. That speed is the appeal, and it is also the risk, because writing good questions is hard and AI is not automatically good at it. Research on AI-generated multiple-choice questions has found that teachers do not always rate the output highly, particularly for the wrong-answer options, which are called distractors. The takeaway is not "don't use AI." It is "review before you use."
This guide describes the problems to look for, gives you a review checklist you can apply in about a minute per question, shows a before-and-after example, and explains how to get better questions in the first place. The short answer is to solve every question yourself, check that each wrong option reflects a real mistake, and keep the ones that pass.
What goes wrong in AI-generated questions?
The most serious problem is a wrong answer key, and it happens most often in maths and science, where a confident answer can be plainly incorrect. A related problem is two right answers, when a "wrong" option is arguably correct. Trivial distractors are more common still. These are wrong options that no student would choose, so the question tests very little.
Other problems are subtler. Cueing happens when the correct answer is longer, more specific or worded differently from the others, so students can guess without knowing. Questions often ask only for surface recall, such as definitions, when you wanted reasoning. They can be misaligned, testing content you have not taught or using different terms, and ambiguous, with vague stems that could mean several things. Some carry cultural or contextual bias, using examples that assume background some students lack.
What should you check in each question?
Take about one minute per question and work through five areas. On accuracy, solve the question yourself and confirm that you reach the marked answer and that only one option is defensible. On alignment, make sure it matches a learning objective you taught and that its vocabulary and method match your class. On clarity, check that the stem is a complete question, that there are no double negatives and that the reading level suits your students.
On distractors, check that each wrong option reflects a real misconception, that the options are similar in length and form, and that "all of the above" and "none of the above" are avoided, because they let students succeed by elimination. On depth, ask whether the question requires more than recalling a phrase, and whether the set as a whole has a mix of easy, medium and hard questions. If a question fails on accuracy, fix it or bin it. If it fails elsewhere, decide whether one edit will save it.
What does a weak question look like next to a good one?
Here is the kind of question a vague prompt tends to produce.
What is a noun?
A) A word that names a person, place or thing
B) A type of animal
C) A kind of rock
D) A season of the year
The answer is obvious even to a student who was not listening, because three of the four options are absurd. It tests almost nothing. Now compare a question built so that each wrong answer is a real mistake.
A jacket costs $80 and is 25% off. What is the sale price?
A) $55 B) $60 C) $20 D) $105
Option A comes from subtracting 25 dollars instead of 25 percent. C is the discount, not the price. D adds the discount instead of taking it off. Only B is correct. If you see a class split between A and C, you know what to reteach, and that diagnostic value is why good distractors matter.
How do you prompt for better questions?
Vague prompts get vague questions, so provide context.
Write 8 multiple-choice questions for grade 7 on adding and subtracting fractions with unlike denominators. Use the common-denominator method. Each wrong option should reflect a common student error, and include a one-line note naming that error. Provide the answer key with worked solution.
The "name the error" instruction is especially useful, because it makes distractors meaningful and gives you diagnostic information. It also makes your review faster, since you can see the reasoning behind each wrong option and judge whether it is realistic.
How much checking is enough?
You do not need to rebuild every question, but you do need to look at every one. A 2023 study of LLM-generated multiple-choice distractors found that, on average, only 53% were rated high quality by teachers, meaning suitable to use as they were. Models have improved since, but the finding still describes the right attitude: the first draft is a draft.
A quick way to review a batch is to solve every question blind, without looking at the key, and to fix or bin any where you disagree with it. Then sort the questions into three piles, keep, fix and bin, and do not agonise over the bin pile. Fix the distractors last, because rewriting one weak option usually takes less than a minute. Finally, spot-check the mix. If all your questions are recall, ask for a second batch that asks students to apply or explain.
Why should you use the wrong answers as data?
A good distractor tells you why a student is wrong. If many students choose option C, and you know C is "added the denominators," you know exactly what to reteach. This is the biggest advantage of well-built multiple-choice questions over a plain score, and it is the reason to spend time on distractors.
What other question types can AI draft?
Multiple choice is only one format. You can ask AI for short-answer questions with a marking guide, "explain why" prompts, error-analysis questions in which students find a mistake in a worked solution, and questions that ask students to create an example of their own. These tend to test understanding more deeply, and they need the same review before students see them.
How should you use the questions with students?
How a quiz is used matters as much as how it is written. Low-stakes quizzes at the start of a lesson, on material from last week, give students practice at retrieving what they learned. Go through the answers straight away and ask students why each wrong option is wrong. Avoid using AI-generated questions in a high-stakes test unless you have reviewed and, ideally, trialled them. A flawed question in a graded test costs you trust as well as marks.
Where interactive practice differs
A quiz gives you a score. Interactive practice gives students feedback on each step. When a student answers incorrectly, a tutor can ask a follow-up question to find the misunderstanding rather than simply marking the answer wrong. This is the approach of Tutor AI, where questions are part of a guided Socratic conversation and exercises are interactive, not just a list.
Should you keep a question bank?
Yes. Every question that passes your checklist goes into a bank labelled by topic and difficulty. Over a year you build a reliable resource that AI helped you start and you finished, and it is far quicker to draw from than to regenerate each time.
The short version
Always solve the questions yourself, check distractors, alignment and depth, ask the AI to name the misconception behind each wrong option, and save what passes.
Frequently asked questions
How many AI-generated questions can I trust without checking? None. Even a good generator produces occasional errors, and the ones you miss reach students. Check them all, and expect to fix or bin a fair share.
Can AI mark student answers to these questions? For multiple choice, the key does the marking, so accuracy of the key is what matters. For short answers, AI marking can be inconsistent, so sample and check.
Should students write quiz questions too? It is a strong activity. Writing a good question with plausible wrong answers requires understanding, and you can use the best ones.
How many questions make a good quiz? For a low-stakes retrieval quiz, five to ten. The point is practice and feedback, not coverage.