Research

AI-generated text

ChatGPT improves polish while causal-reasoning training boosts originality in student work

A randomized experiment at Bocconi University with over 1,000 first‑year undergraduates, run in collaboration with OpenAI Economic Research, tested the separate and combined effects of access to ChatGPT (GPT‑4o) and a causal‑reasoning exercise.

ChatGPT improves polish while causal-reasoning training boosts originality in student work

Researchers at Bocconi University, in collaboration with OpenAI Economic Research, conducted a randomized experiment involving more than 1,000 first‑year undergraduate students. The task was a real‑world business case: students prepared marketing recommendations for the university’s merchandise store.

Students were randomly assigned by class period to one of four groups: (1) access to ChatGPT (GPT‑4o) only, (2) a causal‑reasoning critical‑thinking exercise only, (3) both interventions, or (4) neither. The causal‑reasoning training was independent of AI and taught cause‑and‑effect reasoning through a game, examples, questions, and feedback.

Evaluation methods

Submissions were graded by trained human raters using a five‑point rubric. Separately, the researchers applied automated text analysis to measure each submission’s number and variety of ideas, evidence of causal reasoning, and similarity to recommendations from three experts.

Main findings

  • ChatGPT access: students with GPT‑4o access scored nearly one full point higher on the five‑point rubric. Their responses contained more ideas, exhibited clearer logical structure, and were more similar to expert recommendations. The study notes that students still engaged actively with the tool — they had to decide what to prompt, evaluate AI responses, and select content for their final submission.

  • Causal‑reasoning exercise: students who completed the causal thinking exercise provided clearer explanations of why ideas might work or fail and showed more evidence of seeking explanations and questioning assumptions. However, they did not receive higher scores on the traditional rubric, which focused on two standard marketing goals: increasing awareness of and use of the university store.

  • Idea diversity: automated analysis showed that the causal‑training group generated a wider and more distinct set of ideas compared with peers. This dimension was not captured by the rubric, illustrating that conventional grading can reward clarity and structure while overlooking whether a student produced an idea that others did not.

  • Combined intervention: students who received both ChatGPT access and the causal exercise displayed benefits of each approach. Their idea variety matched that of the causal‑training‑only group; their rubric scores and number of ideas were similar to the ChatGPT‑only group. Their work also showed stronger logical coherence and more evidence of causal explanation. Overall, this combined group improved across the broadest set of measures.

Implications for assessment and instruction

The randomized design lets the researchers separate the distinct effects of AI access and critical‑thinking instruction, and their interaction — a useful contribution to growing evidence about AI’s role in education and how to structure learning supports.

A key implication is that as AI enables students to produce more polished, expert‑like answers, educators should reconsider what assignments and assessments measure. Solely evaluating final answers may hide important dimensions of student learning, such as originality, reasoning, and the ability to generate diverse approaches. Many educators are already discussing how tasks and rubrics should evolve to reward these qualities alongside clarity and correctness.

Conclusion

The experiment shows complementary roles for AI tools and critical‑thinking training. GPT‑4o improved the polish, coherence, and expert‑likeness of student work; causal‑reasoning training broadened idea generation and strengthened explanations of cause and effect. Combining both interventions produced the most comprehensive gains, suggesting that teaching students both to use AI effectively and to reason causally can better prepare them for future challenges.