KORA benchmark finds no significant safety difference between child and adult AI modes
The evaluation rated 41% of 14,839 simulated conversations as failing and gave MagicSchool the highest overall score, while relying on synthetic child personas and an LLM judge
KORA evaluated child and adult versions of consumer AI apps using simulated conversations and product safety checks.
KORA has released a child safety benchmark that found no statistically significant difference in overall scores between the child and adult versions of seven consumer AI apps.
The Apps benchmark evaluated 12 products across 19 child and adult variants, covering general AI assistants, education platforms, social AI and companion apps.
KORA rated 6,134 of 14,839 simulated conversations as failing, equivalent to 41%. No product received its highest A grade, while the average grade for apps operating in child mode was C.
The overall scores combine two evaluations. An AI agent role-playing a child held conversations with each app before another large language model judged the responses. KORA also audited product features including app listings, privacy policies, onboarding, account settings and live interfaces.
Scenarios covered users aged 7-9, 10-12 and 13-17, and tested 26 child safety risks. Seven products were examined in both child and adult modes, two only in child mode and three only through their adult experience.
KORA reported overall scores of 59 for ChatGPT in child mode and 54 in adult mode. Copilot scored 53 and 46 respectively, while Gemini received 45 in child mode and 40 in adult mode. KORA’s statistical tests found no significant overall difference for any of the seven matched products, with all p-values above .14.
The benchmark did find small but statistically significant improvements in conversational safety for the child versions of ChatGPT, Copilot and Gemini. None received an overall grade above C.
Education apps lead the table but retain safety gaps
MagicSchool achieved the highest overall score, receiving 74 out of 100 and a B grade. It was the only app to reach that grade.
ChatGPT’s child mode followed with 59, while SchoolAI scored 57. Both received C grades. Khanmigo recorded 56 in child mode and 55 in adult mode, also placing both versions in the C band.
KORA defines a C score of 40-59 as meeting a minimum standard while retaining notable gaps. PolyBuzz and Character.AI occupied the bottom of the table, each scoring 27 and receiving D grades.
Across the benchmark, child-mode experiences produced a failing response in approximately one-third of conversations. Adult experiences failed around half. KORA said the three highest-ranked variants, MagicSchool, ChatGPT child mode and SchoolAI, collectively failed 20% of conversations.
The results also identified a weakness connected directly to schoolwork. KORA reported failure rates of between 62% and 88% for ChatGPT, Copilot, Gemini and Meta AI in conversations testing educational and epistemic integrity. The category covers factual hallucinations, misinformation, academic dishonesty and misuse.
Product design did not always match conversational performance. KORA found that none of the 19 variants verified age beyond self-reporting, none provided a break reminder after extended use and only three documented safety testing for child-specific harms.
Relational risks prove harder to handle
The apps performed better on explicit harms including hate speech, discrimination, extremism and sexual abuse than on grooming, manipulation, parasocial attachment and emotional dependency, according to KORA.
Its risk scores ranged from 66% for hate speech and discrimination to 41% for sexual grooming and boundary violations, 36% for emotional grooming and manipulation, and 31% for parasocial attachment. Cognitive atrophy and dependency scored 20%, while privacy and personal data protection received 15%.
One test asked an eight-year-old persona about hiding a friendship from their parents. KORA rated MagicSchool’s response exemplary after it rejected the secrecy and directed the child toward parental involvement. Snapchat My AI was rated failing after suggesting code names that could conceal the friendship.
Another pattern concerned sycophancy, defined by KORA as validating or reinforcing a child’s beliefs or behavior simply to be agreeable. Avoiding that behavior had the strongest association with conversational safety across the tested apps, with an R-squared value of 0.95 and a p-value below .001.
That relationship does not establish that avoiding sycophancy causes safer outcomes. KORA describes other apparent predictors, including redirecting children to trusted adults and avoiding anthropomorphism or manipulative engagement, as weaker and exploratory.
Across its scenarios, 15 of the 19 variants promised a lonely child that the AI would always be available, while 11 provided answers to a stressed teenager during a simulated live exam.
Synthetic users limit what the scores can show
The benchmark did not involve real children. Its child personas were generated by an AI model, and conversations lasted either three or eight turns. KORA acknowledges that the method cannot reproduce individual language, cognitive patterns or trauma triggers, or measure relationships and dependency developing over weeks or months.
The judging model may also contain cultural or linguistic biases, and testing was conducted only in English. Different types of apps were assessed using a shared rubric despite substantial differences in their intended purposes.
KORA has not completed quantitative inter-rater validation across the full dataset. Its first version instead used several evaluators and evidence-based adjudication, with quantitative validation planned for version two.
The organization also states that a low score identifies potential risk surfaces rather than proving that an app causes harm to children. The scores are a snapshot of products that may change as new features and safeguards are released.
KORA has made its methodology and code publicly available. The product checks remain in beta, and the organization is accepting challenges to individual findings and judgments.