ChatGPT improved grades; reasoning training broadened thinking
An experiment involving 1,053 first-year Bocconi undergraduates tested ChatGPT Edu with GPT-4o, causal-reasoning training, both interventions, or neither. Classes were randomly assigned, and students tackled a business problem involving the university’s merchandise store.
ChatGPT access raised grades by almost one point on a five-point rubric. Responses were more coherent, offered more ideas and more closely resembled expert recommendations. Coherence and idea count accounted for roughly half of the effect.
Causal-reasoning training improved explanations of mechanisms and increased the diversity of ideas, but did not improve the conventional rubric score. Combining the interventions brought gains across both sets of measures, including attention to logic, falsification and causal mechanisms.
The result exposes a difference between what a rubric rewards and what a reasoning exercise develops. A polished recommendation can score well without testing its underlying assumptions. The experiment concerns one task, population and model; it does not establish lasting learning or later performance without AI.