ChatGPT boosts student scores while critical-thinking training broadens ideas
A randomized trial with 1,053 first-year Bocconi University students found GPT-4o improved conventional performance scores, while causal-reasoning training increased the diversity of students’ ideas
A randomized trial involving 1,053 Bocconi University students found ChatGPT improved conventional performance scores while critical-thinking training increased idea diversity
A new study shows giving university students access to ChatGPT helped them produce higher-scoring, more expert-like work, but a separate critical-thinking intervention changed something the traditional grading rubric largely missed: the range and distinctiveness of their ideas.
The randomized controlled trial, conducted by researchers at Bocconi University in collaboration with OpenAI Economic Research, involved 1,053 first-year undergraduates studying economics, management and finance.
Students were split into four groups. One received access to ChatGPT Edu using GPT-4o, another completed training in causal reasoning, a third received both, and a control group received neither.
They were then given 45 minutes to tackle the same real-world business problem, producing a short consultant-style recommendation for increasing awareness and use of their university’s merchandise among alumni.
On the conventional performance measure, ChatGPT made the clearest difference. Students with GPT access scored an estimated 0.86 points higher on a five-point scale than the control group, whose estimated score was 2.09. Their recommendations were also more similar to those produced independently by three domain experts.
When researchers controlled for features including coherence, writing style, number of ideas and idea diversity, around half of the GPT effect remained. That led the researchers to conclude that ChatGPT was improving some of the substance of students’ answers as well as their presentation.
But the students given critical-thinking training produced a different kind of gain.
Critical-thinking training produced less conventional ideas
The causal-reasoning intervention taught students to think about cause and effect, explain why a proposed solution might work and consider conditions under which it might fail.
Those students showed substantially more mechanism-based reasoning and falsification logic in their final responses. Their solutions also contained a more diverse mix of ideas and were more distinct from those produced by other participants.
ChatGPT, meanwhile, substantially increased the number of ideas students included, but did not produce the same increase in between-student idea diversity.
Causal-reasoning training alone did not improve scores on the standard marketing rubric. In some specifications, its initial effect was slightly negative.
The problem was not that the training had failed. The attributes it encouraged were simply not what the assessment rewarded.
Responses containing more coherent logic and a larger number of ideas tended to score better. But answers that emphasized mechanisms, falsifiability or ideas further from the typical solution space were more likely to receive lower scores.
In other words, the students generating less conventional responses were not necessarily being rewarded for doing so.
The researchers argue that this exposes a weakness in assessment design when AI can already generate polished, conventionally strong answers. If originality and diversity of thought are valued outcomes, grading systems have to measure them.
Using AI and thinking critically were not competing skills
The group receiving both ChatGPT access and causal-reasoning training retained the characteristics of each intervention.
Their performance scores and number of ideas were similar to students with ChatGPT alone, while their idea diversity was comparable with students who had received causal training. The combination also produced stronger coherent logic and more evidence of mechanisms and falsification than GPT access by itself.
That undercuts a simple choice between teaching students cognitive skills and teaching them to work with AI.
OpenAI CFO Sarah Friar made the same point when sharing the research on LinkedIn, writing: “AI and critical thinking aren’t mutually exclusive. They're actually complementary.”
She also drew a distinction between presentation and judgment: “AI can help us deliver polished work much faster. But polished doesn’t always equate to thoughtful.”
The experiment provides evidence for that distinction, but within a tightly defined task. Students were working on a business case with established performance criteria in an area where the researchers expected the LLM to have substantial relevant knowledge.
The paper does not establish that the same performance advantage would appear across different subjects, assignments or types of intellectual work.
A grading problem as much as an AI problem
The findings also complicate the question of what educators should evaluate when students have access to generative AI.
Twenty trained master’s students assessed each response, with three independent ratings per submission, using a five-point rubric focused on two established marketing goals. The researchers separately compared students’ work with recommendations created by a marketing professor, a university merchandising manager and an expert alumnus.
ChatGPT-assisted answers moved closer to those expert responses and produced greater agreement among evaluators.
Causal reasoning pushed students in another direction, toward ideas that were more varied and less typical of the overall cohort. Yet being further from the conventional answer was itself associated with lower evaluation scores.
For universities revisiting assignments and assessment in response to generative AI, the experiment therefore points beyond the question of whether students should be allowed to use ChatGPT. It also tests what happens when the assessment system continues to reward a polished standard answer after AI becomes good at producing one.
There is one question the trial cannot answer. Researchers observed the quality of what students submitted, not what they subsequently learned or retained. They say they cannot determine whether the portion of ChatGPT’s performance advantage attributed to better substantive recommendations represented knowledge acquired by students or output they obtained without acquiring that knowledge themselves.