AI models gave students answers in 97% of unprompted tutoring tests, Comprendo finds
The company ran 779 simulated eighth-grade algebra conversations across six frontier models and found few factual errors, but prompting exposed a trade-off between diagnosing misconceptions and withholding answers
Comprendo tested six frontier AI models across 779 simulated eighth-grade algebra tutoring conversations
Comprendo has tested whether six frontier AI models can tutor simulated eighth-grade algebra students rather than simply solve problems for them, finding zero outcomes that met its full teaching criteria across 298 conversations conducted without specific tutoring instructions.
The models handed students the answer in 97% of those conversations. A dedicated tutoring prompt reduced answer-giving, but no model performed strongly across both parts of the task measured by Comprendo: identifying the student’s misconception and leaving the student to complete the reasoning.
The evaluation, covered 779 conversations and approximately 3.7 million tokens. Comprendo tested GPT-5.6 Terra, Claude Sonnet 5, Grok 4.5, DeepSeek V4 Pro, Kimi K3 and Qwen3.8 Max.
Each model received a basic prompt similar to one a student might submit and a separate tutoring prompt. The simulated students were given misconceptions drawn from documented mathematics education research, allowing evaluators to check whether a model found the intended error.
Trevor Selby, Founder of Comprendo, drew a distinction between accuracy and instruction in a LinkedIn post about the results: “We ran 779 simulated tutoring conversations to test whether AI models can teach, not just answer.”
“Out of the box, models don't tutor. Without a tutoring prompt, they handed the student the answer in 97% of conversations. Zero successful tutoring outcomes in 298 tries,” Selby added.
Accuracy was less of a problem than instructional behavior: “The models were rarely factually wrong. We found 3 false statements across all 779 conversations. The problem is that handing over answers isn't teaching, and there's no standard way for buyers to tell the difference ahead of time.”
That distinction drove the scoring. A successful conversation required the model to diagnose the planted misconception, avoid false claims and withhold the answer.
Tutoring prompts improve restraint but expose a trade-off
Under the naive prompt, none of the six models produced a clean tutoring win. This result was consistent across the 298 conversations in that condition.
Adding a tutoring prompt substantially reduced answer leakage. Performance still varied, with models giving away between one and nine answers across 65 conversations each.
Diagnostic accuracy also differed. Prompted models correctly identified the planted misconception in between 88% and 100% of relevant conversations.
Comprendo measured restraint in two ways. The narrower test checked whether the tutor explicitly stated the answer. Its stricter test also failed models that completed the student’s reasoning while leaving only arithmetic or another minor step unfinished.
Strict clean-win rates ranged from 28% to 54% across the six models. DeepSeek V4 Pro recorded the highest rate at 54%, followed by Grok 4.5 at 52%. Claude Sonnet 5 recorded the lowest at 28%.
The rankings changed under the narrower definition. Models that diagnosed errors most consistently often performed more of the reasoning for the student, while those showing greater restraint missed more misconceptions.
This left the upper corner of Comprendo’s evaluation empty. No model combined the highest diagnostic accuracy with the strongest strict restraint.
Accuracy was high, but price did not predict tutoring quality
Factual mistakes were uncommon. The evaluation identified three false statements across all 779 conversations, each confirmed by a human reviewer.
No model incorrectly told a control student that a correct answer was wrong. The models also found the “students-and-professors” reversal, identified in mathematics education literature as a difficult misconception, harder to diagnose than the other errors tested.
Cost produced another separation between conventional model comparisons and tutoring performance. The cheapest model cost approximately $0.002 in API fees per conversation and achieved the highest strict clean-win rate.
The most expensive model cost roughly 12 times more per conversation but produced what Comprendo described as an average result. API costs varied by a factor of 16 across the six models.
Those calculations used OpenRouter pricing from August 2026. Comprendo notes that model versions and prices change frequently, limiting how long any one cost comparison remains current.
Simulated conversations are not classroom evidence
The evaluation tested general-purpose frontier models, not commercial AI tutoring products. No students, teachers or classrooms participated.
Comprendo argues that testing foundation models remains useful because AI tutoring products are likely to build on them. However, commercial products may add prompts, fine-tuning, curriculum content, guardrails and other systems that could change their performance.
The company used 13 scenarios, comprising 10 documented algebra misconceptions and three controls in which the student’s work was correct. Each scenario was repeated five times.
An AI judge scored every transcript against an answer key before Comprendo audited its decisions using human-labeled samples. The company says the judge’s errors were generally conservative, producing false alarms and undercounting correct diagnoses rather than overlooking problems.
There are still methodological limits. The strict withholding rubric was created for writing feedback and adapted for mathematics dialogue. It agreed with a human expert in 77% of an adversarial sample. Comprendo also notes that the model playing the simulated student belonged to the same model family as one of the tutors being tested.
Comprendo is an applied research and technical strategy company that offers independent measurement and evaluation services for AI in education. The evaluation was conducted by the company rather than through a classroom trial.