News
Study of 27,000 Students: AI Boosts Homework Grades but Cuts Exam Scores by Up to 24 Percent

A three-year study tracked nearly 27,000 students in China: using AI chatbots for homework raised grades by 18 percent, but after two years, entrance exam scores fell by as much as 24 percent.
Contents
Researchers from Stockholm and Hong Kong tracked nearly 27,000 Chinese students for two and a half years to calculate the real cost of using artificial intelligence to help with homework. In the short term, students come out ahead: homework grades rise and the time spent on assignments drops. In the long term, they pay for it with entrance exam scores, which can fall by as much as a quarter.
The study was conducted by David Strömberg of Stockholm University along with Victor Lei and Yanhui Wu of the University of Hong Kong. The paper was published on August 20, 2026 as a working paper by the Centre for Economic Policy Research (CEPR) and covered a single county in central China, where access to AI chatbots spread unusually fast among teenagers.
The Grades-Exams Paradox
Previously, homework grades were a good predictor of exam results, better homework usually signaled a better test score. After chatbots became widespread, that relationship reversed. The authors found that students with the highest homework grades increasingly performed worse on exams, because AI had started doing the thinking for them instead of merely supporting it.
The gap widened over time. In the first six months after AI was introduced into homework routines, monthly test scores fell by 20 percent relative to the group of students who did not use chatbots. After two years, the difference became entrenched in the highest-stakes exams: the zhongkao, the entrance exam for high school, and the gaokao, the national college entrance exam.
Who Loses the Most
The losses were not evenly distributed. The steepest score drops appeared in social science subjects, with somewhat smaller declines in STEM subjects and languages. The hardest hit were younger students, previously top-performing students, and boys, the groups that had done best before the AI era or had the most room to fall behind.
The authors note that about 80 percent of students who used AI showed patterns typical of outsourcing their thinking, copying ready-made answers instead of working to understand the material. This group accounted for most of the decline in exam scores.
How You Use It Determines the Outcome
The study does not show that AI in education is inherently harmful. The key factor turned out to be the difference between treating a chatbot as a tutor and treating it as a machine for generating ready-made answers. Students who, despite having access to AI, spent the same amount of time on assignments as peers without it suffered minimal consequences in exam results.
Students whose exam scores remained high did not limit themselves to mindlessly copying and pasting answers to save time.
In other words, the same students who used AI to explain concepts and work through problems step by step, rather than to generate a final answer, achieved results close to those of the group that did not use chatbots at all. This distinction, a tool that supports learning versus a tool that replaces thinking, became the central conclusion of the paper.
Relevance for Poland
The findings matter beyond China's education system, because they point to a universal mechanism: shorter work time and higher current grades at the cost of deep understanding of the material. In Poland, the topic of AI in schools has gained fresh weight following the education ministry's decision to place artificial intelligence at the center of education policy, the Chinese study provides concrete numbers that could serve as an argument in the debate over how to teach chatbot use without losing cognitive skills.
For teachers and parents, the practical takeaway is simple: a high grade on homework completed with AI help is no longer a reliable signal of mastery of the material. Exams without access to AI tools remain the only reliable test of real knowledge, and the gap between them and current grades may grow precisely where students most readily reach for ready-made answers.
The authors caution that the results apply to a single county in central China and a specific set of chatbots popular there, so the scale of the effect should not be mechanically applied to other education systems. Still, the direction of the relationship, short-term gains followed by long-term losses from passive AI use, appears consistently across every measure in the study.
