OECD AI Schoolwork Study Finds Lower Scores, Not a 1.5-Year AI Loss
The OECD AI schoolwork findings carry a sharp warning: students using chatbots for specific assignments generally scored lower than students who avoided them.
The finding arrives beside an even larger number. Average OECD reading scores fell 28 points between 2015 and 2025, equal to about 1.5 years of learning. Mathematics scores fell 22 points, or slightly more than one year.
Those figures do not show that chatbots caused a decade of decline. Public generative AI tools only became widely available after late 2022. The decline began earlier and reflects many pressures, including weaker family support, absenteeism, distraction, and pandemic disruption.
Still, the overlap matters. AI reached classrooms while foundational performance was already falling. Schools now face a difficult choice between restricting shortcuts and teaching students to use the technology deliberately.
Asian education systems make that choice especially visible. Chatbot use is widespread in Singapore, Korea, Macao, Thailand, Vietnam, and the Philippines. Yet several Asian systems also remain among the strongest performers.
The emerging divide is not simply AI users against non-users. It is answer-generating AI against instructionally guided AI, with teachers deciding whether the tool replaces effort or strengthens it.
What the OECD AI Schoolwork Findings Actually Show
PISA 2025 found an association between chatbot use and performance, not proof that AI caused lower scores.
The OECD released its latest Programme for International Student Assessment results on September 8, 2026. PISA tests how 15-year-old students apply reading, mathematics, and science knowledge to unfamiliar problems.
The 2025 assessment covered 760,000 students across 91 countries and economies. It was the largest PISA exercise conducted so far. Science was its main academic domain.
This cycle also asked students how they used AI chatbots for schoolwork. The questions covered summarizing assigned texts, researching new topics, drafting written assignments, and getting general learning help.
Chatbots had already reached most students. Only 14% across OECD countries said they had never, or almost never, used AI for any surveyed schoolwork purpose.
About 46% used AI at least weekly to help them learn. The figure confirms that school systems are studying an established behavior, not an experimental classroom trend.
The performance pattern was unfavorable for several assignment-specific uses. Students who did not use AI to summarize, research, or draft generally outperformed users in science.
Frequency also mattered. Moderate users often scored above students who used chatbots rarely or almost every day. Drafting assignments showed a weaker advantage for moderate use than other tasks.
General learning produced a different result. Students using AI weekly to help them learn performed at roughly the same level as non-users after adjusting for socioeconomic background.
The complete schoolwork analysis explicitly warns against treating these relationships as causal. Different students adopt AI for different reasons and under different conditions.
A struggling student might turn to a chatbot more often because the work is already difficult. That pattern would produce lower scores among users without proving that AI created the original difficulty.
Access also complicates the comparison. Students who never used AI tended to come from more disadvantaged backgrounds. Frequent users were more likely to come from advantaged backgrounds.
Socioeconomic adjustments reduce that imbalance but cannot remove every unmeasured difference. Motivation, prior knowledge, teacher practices, assessment design, and the quality of individual prompts can all affect results.
The safest conclusion is narrower than the alarming headline. Outsourcing specific academic tasks correlates with weaker tested knowledge, while moderate learning-oriented use does not show the same penalty.
That distinction sets up the real policy question. Schools cannot evaluate AI by counting logins alone. They need to examine what cognitive work remains with the student.
The 1.5-Year Figure Describes a Decade of Reading Decline
The widely repeated 1.5-year loss refers to reading performance since 2015, not a measured penalty caused by AI.
Average reading performance across OECD countries dropped 28 points from 2015 through 2025. The organization translates that difference into approximately one and a half years of learning.
Mathematics declined by 22 points during the same decade. Science performance fell more modestly over ten years and remained broadly stable between 2022 and 2025.
The PISA results also show that one in five students performed poorly across science, mathematics, and reading. That share was 16% in 2022.
The chronology rules out a simple explanation centered on ChatGPT or similar services. The broader slide began before generative AI became a common consumer product.
PISA does not follow the same individual students across the decade. It samples comparable groups of 15-year-olds, creating a system-level view rather than a personal learning timeline.
It also measures performance at one point. The assessment cannot observe every step connecting a student's chatbot activity with the knowledge displayed during testing.
The OECD therefore presents two findings that should remain separate. Academic performance has declined over a decade, and current AI-use patterns correlate with different performance levels.
Their intersection still deserves attention. Generative AI arrived during a prolonged learning crisis, when many students already lacked secure reading and mathematics foundations.
A chatbot can draft a coherent paragraph even when its user cannot build the argument independently. It can summarize a chapter without requiring the reader to identify its central claims.
That creates a measurement gap between completed work and retained ability. Homework quality can rise while unaided performance stays flat or falls.
Teachers then receive a misleading signal. A polished submission might represent stronger learning, useful editing support, extensive machine assistance, or nearly complete delegation.
Students face a similar illusion. Fast completion can feel like mastery because the immediate output looks correct. The misunderstanding becomes visible only during an unaided test or later task.
This does not make every chatbot interaction harmful. It means the educational value depends on whether AI preserves retrieval, reasoning, revision, and explanation.
The OECD's Secretary-General, Mathias Cormann, framed the distinction clearly. Technology can strengthen learning, he said, but not when it substitutes for attention, effort, and understanding.
That statement shifts responsibility beyond students. Product design, teacher instructions, assignment structure, and school policy determine how easily AI becomes a substitute.
The 1.5-year decline raises the stakes because schools have little room for another source of shallow engagement. It does not provide a causal estimate of that source.
Why Summarizing and Drafting Create the Sharpest Conflict
AI becomes educationally risky when it completes the mental operation that an assignment was designed to exercise.
A summary assignment is rarely about producing a shorter document. Its learning purpose is to make students select evidence, distinguish main ideas, and reconstruct meaning.
When a chatbot performs those steps, the student receives the visible product without necessarily practicing the hidden skill. The same problem applies to preliminary research.
Research assignments teach question formation, source comparison, uncertainty management, and evidence selection. Asking for a ready-made overview can remove those decisions from the student's workflow.
Drafting presents an even clearer conflict. Writing is not merely the final expression of knowledge. It helps people discover weak reasoning, missing evidence, and unresolved contradictions.
A generated first draft can erase that productive friction. Students start from fluent prose and become editors before they have formed their own argument.
This mechanism helps explain the OECD AI schoolwork findings. Non-users generally outperformed users on specific tasks, while moderate users sometimes did better than rare or intensive users.
The pattern resembles a curve rather than a ban-or-adopt choice. Some assistance can support learning, but extensive delegation reduces the effort that creates durable knowledge.
The OECD's separate digital education review reaches a similar conclusion. Generative AI can help when teaching principles guide its use.
Educational tools can ask a learner to explain a step before offering feedback. They can provide hints, generate practice questions, or challenge an unsupported claim.
Generic chatbots usually optimize for satisfying the request. If a student asks for an answer, the system has little reason to preserve the assignment's intended difficulty.
That difference separates tutoring from completion. A tutor manages assistance so the learner retains responsibility for the reasoning. An answer engine can remove that responsibility instantly.
The problem is not limited to cheating. A student may openly use AI under permissive school rules and still practice less than the curriculum expects.
Academic integrity policies address authorship, disclosure, and prohibited assistance. Learning design asks a deeper question: what thinking must students still perform themselves?
Schools need both layers. Disclosure can tell a teacher that AI was involved, but it does not show whether the interaction strengthened understanding.
Process evidence becomes more useful. Students can retain notes, source trails, successive drafts, prompt histories, and short explanations of rejected suggestions.
A student knowledge base can support that kind of traceable process. The educational value comes from revisiting evidence and reasoning, not collecting generated answers.
Assessment also needs to compare assisted and unaided performance. A student might produce an AI-supported essay, then defend its argument verbally or write a related passage without assistance.
Such designs make AI use visible without treating every interaction as misconduct. They also reveal whether support transferred into independent ability.
This is the central reversal. AI can improve the artifact while weakening the learning process that the artifact was meant to represent.
Asia Shows That High Adoption and High Performance Can Coexist
Asian results challenge any claim that frequent chatbot access automatically produces weak academic performance.
Singapore offers the clearest counterexample. Its students remained among the world's strongest performers while reporting extensive AI use across school tasks.
About 66% of Singaporean students used chatbots at least weekly to help them learn. The OECD average was approximately 46%.
Weekly use was also high for individual tasks. In Singapore, 46% used AI for preliminary research, 41% for summarizing, and 45% for drafting written assignments.
Only 4% reported no or almost no AI use across the surveyed activities. That compares with 14% across OECD countries.
The country's PISA profile provides another important figure. About 81% learned to assess AI-generated information during school lessons.
The OECD average for that classroom activity was about 63%. Across participating systems, reported exposure ranged from 31% to more than 80%.
Singapore therefore combines high access with unusually broad instruction in evaluation. Students encounter AI, but many also practice questioning its output.
Korea presents another version of the pattern. Its overall schoolwork AI-use index stood above the OECD average, while it remained among the top-performing systems.
Korea and Singapore have both pursued broad digital learning policies. Singapore's national approach explicitly includes artificial intelligence tools.
Other Asian economies show why broad regional labels remain risky. Japan reports comparatively low schoolwork use, while Singapore reports very high use. Both perform strongly.
Chinese Taipei combines strong performance with lower average AI use and a large digital learning policy dating from 2021. Macao combines high performance with extensive adoption.
The mainland China results cover Beijing, Shanghai, Jiangsu, and Zhejiang rather than the entire country. Readers should not treat that sample as a national estimate.
The better comparison concerns instructional control. Several high-performing systems give teachers a central role in adapting digital tools to classroom goals.
OECD education director Andreas Schleicher described a Chinese calligraphy classroom where students still wrote with brushes and ink. Teachers then used AI to analyze the handwriting.
The classroom example is useful because the technology did not produce the student's work. It delivered feedback on work the student had already performed.
That structure keeps the essential practice intact. Students form each character, teachers define the learning objective, and software adds a diagnostic layer.
Compare that with asking a chatbot to draft an essay. The second workflow transfers the core act of composition to the model.
This is why “AI in education” is too broad to function as an outcome variable. Calligraphy feedback, automated drafting, guided tutoring, and instant summarization are different interventions.
National averages can also hide classroom variation. One teacher might require evidence checks and oral defenses. Another might allow unrestricted generation with limited supervision.
Asian adoption therefore does not disprove the OECD's student-level associations. It shows that system design can change what adoption means.
High-performing Asian systems pressure countries that frame policy as a choice between prohibition and unrestricted access. A third route is visible: guided exposure with teacher authority.
What the Numbers Still Cannot Tell Us
PISA identifies a warning signal, but its observational design cannot isolate the effect of AI from the students and schools using it.
The survey relies on student reports about frequency and purpose. It does not capture every prompt, response, correction, or classroom rule surrounding those interactions.
Two students can select the same frequency category while using AI very differently. One might request explanations, while another submits generated text with minimal review.
The broad phrase “help me learn” creates additional uncertainty. It can describe tutoring, translation, brainstorming, practice generation, answer checking, or direct solution requests.
Science scores provide the primary outcome in this PISA cycle. They cannot represent every effect on writing quality, creativity, confidence, research skill, or long-term retention.
The assessment also cannot establish direction. Lower-performing students might adopt AI because they need help, making AI use partly a response to difficulty.
Frequent users might face demanding workloads that encourage automation. Schools with weak support systems might rely on consumer chatbots as inexpensive substitutes.
Another possibility runs in the opposite direction. Repeated task delegation might reduce practice and weaken later unaided performance. PISA cannot determine each mechanism's contribution.
Even the relationship between AI literacy instruction and scores requires caution. Advantaged students report those learning opportunities more often.
Schools that teach verification effectively may also have stronger teachers, better resources, and more supportive families. Those factors can influence performance independently.
The OECD nevertheless found a promising pattern. Frequent general-learning users who also assessed AI information in class achieved slightly higher science performance than infrequent or non-users.
That result supports guided use, but it does not prove a particular lesson caused the difference. The report describes an association inside a complicated learning environment.
A further uncertainty concerns product change. The chatbots used during 2025 differ from later systems in accuracy, interface design, personalization, and educational safeguards.
Future tools might become better tutors. They might also become better at completing assignments invisibly, widening the gap between apparent output and retained competence.
Schools cannot wait for perfect evidence because adoption is already pervasive. Yet they should avoid turning preliminary correlations into universal bans.
A blanket ban can drive use out of sight. It can also deny students supervised practice with tools they will encounter in higher education and work.
Unrestricted adoption carries the opposite risk. It lets product defaults determine pedagogy, even though consumer chatbots were not designed around individual curricula.
The cautious policy response is controlled experimentation. Schools can define permitted tasks, preserve unaided assessments, teach verification, and compare outcomes over time.
They should measure more than submission quality. Useful indicators include delayed recall, independent transfer, source accuracy, revision quality, and the ability to explain reasoning.
Those measures would test the mechanism behind the OECD AI schoolwork findings. They would show whether a chatbot supports cognitive work or merely conceals its absence.
Three Signals Will Show Whether Schools Are Learning From PISA
The next phase will be decided by assessment design, teacher-led AI literacy, and evidence that guided use transfers into independent performance.
The first signal is a shift toward process-aware assessment. Schools should disclose how AI was used and require students to demonstrate the same knowledge without assistance.
Watch for oral defenses, supervised writing, prompt records, source logs, and staged drafts. These approaches separate polished output from genuine understanding.
If school systems adopt them widely, the OECD warning will strengthen the case for redesign instead of prohibition. If assignments remain unchanged, invisible delegation will remain difficult to measure.
The second signal is broader instruction in evaluating generated information. About six in ten OECD students already report doing this during school lessons.
The quality of that instruction matters more than the checkbox. Students need to test claims, locate original evidence, identify uncertainty, and recognize persuasive but unsupported prose.
Singapore's 81% exposure rate provides a useful benchmark, not a complete explanation for its results. Other systems must show that evaluation exercises change student behavior.
If AI literacy expands while independent performance improves, guided use gains support. If exposure rises without better outcomes, schools will need stronger instructional designs.
The third signal is longitudinal evidence comparing different forms of assistance. Researchers need to separate tutoring, feedback, brainstorming, summarization, research, and drafting.
They also need delayed tests. Immediate task performance can improve even when students retain less knowledge a week or month later.
The next PISA cycle will help, but schools should not rely on a three-year assessment alone. Districts and ministries can run controlled evaluations inside existing courses.
The OECD is already preparing a media and AI literacy domain for PISA 2029. That assessment should reveal whether students can engage critically with AI-mediated information.
Product developers also face a clear test. Education modes should resist giving complete answers before students attempt the underlying problem.
Teachers need controls for hint depth, source requirements, age suitability, and activity records. Without those features, “AI tutor” remains more marketing category than educational guarantee.
Parents and students can apply the same standard now. Ask whether a chatbot is extending thought, checking thought, or replacing thought.
For summaries, read first and draft the main points before requesting feedback. For research, collect original sources before asking AI to compare them.
For writing, form the argument and produce an initial structure before using a model for critique. Keep an unaided version of essential practice.
The reported OECD warning is serious, but its most dramatic interpretation is unsupported. AI did not receive a causal 1.5-year penalty in PISA.
The real finding is more useful. Students often score lower when chatbots perform specific schoolwork tasks, while guided and moderate learning use follows a different pattern.
That leaves schools with an immediate question: are their AI policies protecting the effort that creates knowledge, or only improving the work students submit?



