OECD AI Test Scores Reveal a Learning Reversal for School Students
- Sophie Larsen

- 2 hours ago
- 13 min read
OECD AI test scores reveal a troubling conflict: students who avoid chatbots for schoolwork often outperform classmates who use them. The finding comes from PISA 2025, the OECD’s international assessment of 15-year-old students.
The result challenges a popular assumption about AI in education. Chatbots can produce polished homework, summarize readings, and answer questions within seconds. Yet better assisted work does not automatically build knowledge that students can use independently.
That distinction puts general-purpose chatbots on one side of an emerging divide. On the other are structured systems that make students reason, practice, and explain their answers. The evidence does not support banning AI from education. It supports asking whether the technology helps students think or simply completes the thinking for them.
PISA Finds an AI Test Score Gap
The OECD’s central finding is an association between chatbot use and weaker independent performance, not proof that AI caused the lower scores.
The PISA 2025 results provide one of the broadest views yet of student AI use. PISA, the Programme for International Student Assessment, measures how 15-year-olds apply knowledge to unfamiliar problems.
By 2025, a majority of students in most participating education systems had used AI chatbots for schoolwork. Only 14 percent of students across OECD countries reported almost never using them for the purposes examined.
Those purposes included summarizing assigned texts, conducting preliminary research, drafting written assignments, and getting general learning help. General learning assistance was the most commonly reported use.
Almost-daily use remained less common. Across OECD countries, about one in five students reported regular use. That matters because frequency and purpose were both connected with different performance patterns.
Students who reported no chatbot use for specific school tasks generally earned higher science scores than users. The pattern appeared across most participating countries and economies.
The relationship was not a simple straight line. Students using AI weekly for general learning performed similarly to non-users after researchers adjusted for socioeconomic background.
Students who used AI almost every day, or only occasionally, tended to score below non-users. Moderate users sometimes outperformed limited and frequent users when summarizing texts or researching new topics.
Drafting was different. The performance advantage for moderate users was less pronounced when AI helped write assignments. That result fits the concern that text generation can replace more cognitive work than research assistance does.
The comparison uses predicted science scores adjusted for students’ socioeconomic profiles. It still comes from observational, self-reported usage data. Students were not randomly assigned to chatbot and non-chatbot groups.
Lower-performing students might seek AI help more often. Students with heavier workloads, weaker support networks, or lower confidence might also adopt chatbots differently.
The OECD explicitly cautions that its findings do not establish a negative causal effect. They instead describe a relationship shaped by who adopts AI, why they use it, and how often.
That warning should prevent an easy headline from becoming a false verdict. PISA did not prove that every student loses knowledge after opening ChatGPT. It did reveal a consistent gap that schools cannot safely dismiss.
The gap also matters because PISA tests unaided application. A chatbot can improve the work submitted at home while leaving the student less prepared when the tool disappears.
This difference between completed work and retained ability creates the real conflict. Education systems have spent years evaluating visible outputs. Generative AI can now improve those outputs without reliably improving the learner.
OECD AI Test Scores Expose the Homework Reversal
AI can make an assignment look better while weakening the connection between that assignment and what a student knows.
The OECD’s earlier education outlook identified this problem before the new PISA results arrived. General-purpose AI can raise task performance without producing lasting learning gains.
A student might ask a chatbot to summarize a chapter, organize an argument, or draft an essay. The final submission can become clearer even when the student performs less reading, planning, and revision.
Those skipped steps are not administrative friction. They are often the learning activity itself.
Writing requires a student to retrieve information, weigh evidence, arrange ideas, and notice gaps in understanding. Summarizing requires deciding which facts matter and how they connect.
When a chatbot performs those operations, the student receives an answer without necessarily building the mental structure behind it. The immediate result looks productive, but the independent test exposes what was retained.
Researchers sometimes describe this pattern as cognitive offloading. The learner transfers part of a mental task to an external system, reducing the effort needed to complete it.
Offloading is not automatically harmful. Calculators, spell-checkers, search engines, and reference books all move work outside the mind. Students can use those tools while developing deeper skills.
The difference lies in what remains for the learner to do. A calculator still requires someone to select an operation and interpret the result. A chatbot can select the argument, compose the explanation, and supply the conclusion.
That broader substitution creates what researchers call metacognitive laziness. Metacognition means monitoring one’s understanding, strategies, and errors while learning.
Students exercise it when they recognize confusion, reconsider a weak answer, or ask for targeted help. A fluent chatbot response can suppress those signals by making incomplete understanding feel resolved.
The PISA results include another clue. Students who avoided AI chatbots were also more likely to report stronger help-seeking behavior.
Human help usually involves negotiation. A teacher, parent, or classmate can ask what the student tried and where the reasoning failed. A general chatbot often answers immediately unless prompted to behave differently.
That convenience changes the cost of avoiding effort. Before generative AI, outsourcing a complete essay required money, coordination, or obvious misconduct. Now a plausible draft can appear during a short browser session.
The reversal becomes clearest when the tool is removed. Students may submit stronger assisted work, then struggle on a closed assessment that requires the same knowledge.
A 2026 study of Chinese secondary education offers a sharper version of that pattern. The learning penalty study analyzed 30 months of records from 26,811 students in grades 7 through 12.
The researchers tracked homework, completion time, monthly closed-book exams, and high-stakes entrance examinations across nine subjects. AI adoption was associated with better homework productivity but worse unaided performance.
That study also requires caution. Adoption was not randomly assigned, and researchers inferred some outsourcing behavior from administrative data. It cannot establish a universal penalty for every student or classroom.
Still, its structure addresses a weakness in ordinary homework analysis. It compares assisted work with assessments where students must perform without the tool.
The OECD AI test scores point in the same direction across a much broader international sample. Together, these findings challenge grades based mainly on take-home outputs.
If schools continue treating polished homework as direct evidence of learning, AI will make that signal less reliable. The assignment may increasingly measure access, prompting, and editing rather than independent mastery.
The Data Does Not Support a Simple AI Ban
The strongest evidence separates unguided answer generation from AI use designed around teaching and active participation.
The OECD’s warning is easy to flatten into a claim that AI always damages learning. Its own findings reject that interpretation.
Weekly use for general learning was associated with science performance similar to non-use. Moderate use for certain tasks also produced better results than rare or intensive use.
This pattern suggests that frequency matters, but frequency alone is not enough. The purpose and instructional design around the tool can change the outcome.
A student who requests a finished answer is not having the same experience as one who asks for a hint. Neither resembles a supervised lesson where an instructor challenges incorrect reasoning.
This distinction appears in controlled research. A World Bank study tested a six-week, teacher-supported AI tutoring program with secondary students in Edo State, Nigeria.
The program used Microsoft Copilot to support English learning after school. Teachers supervised structured activities, supplied prompts, and helped students evaluate the system’s responses.
Students in the intervention achieved a 0.31 standard-deviation improvement on an assessment covering English, AI knowledge, and digital skills. The Nigeria tutoring trial therefore offers evidence that guided use can improve learning.
That result does not contradict the OECD warning. It clarifies the mechanism behind it.
The Nigerian program did not give students an unrestricted answer engine and assume learning would follow. It placed the model inside a curriculum, a schedule, and a teacher-led process.
The chatbot functioned as a tutor within an instructional system. It did not replace the system.
More recent experimental evidence also complicates a purely negative story. Researchers Zara Contractor and Germán Reyes randomly assigned undergraduates to learn an unfamiliar topic with or without generative AI.
Participants completed unaided assessments immediately and one week later. AI access raised immediate knowledge scores by 0.27 standard deviations, and those gains persisted one week later.
The randomized experiment found stronger delayed benefits among augmentation users. Those students asked AI to explain concepts instead of generating text for them.
Automation users followed the opposite pattern. Their short-term improvements in output quality disappeared after AI access was removed.
This contrast provides a practical framework for interpreting the OECD AI test scores. The relevant question is not whether a chatbot appeared during study.
The question is which cognitive steps remained with the student. Explanation, retrieval, comparison, and self-correction can support learning. Automatic production can bypass them.
Age and context also limit direct comparisons. The randomized experiment involved undergraduates, while PISA examines 15-year-olds. The World Bank intervention operated in a specific Nigerian setting.
A structured after-school program cannot predict what happens when millions of students use different chatbots privately. A small experiment cannot resolve long-term effects across every subject.
Even so, the positive studies expose the weakness of blanket prohibition. Removing AI would also remove tutoring options that can expand feedback and practice.
The policy choice is therefore not unlimited use versus total exclusion. Schools need rules that distinguish productive assistance from substituted effort.
A chatbot might be allowed to ask Socratic questions, generate practice exercises, or explain why an answer is wrong. It might be restricted from producing final text for assessed work.
Students can also document their interaction history and explain which suggestions they accepted. That process makes reasoning visible instead of treating the final document as the only evidence.
Tools for students should support retrieval, reflection, and source comparison. A personal student workspace can help organize materials, but the learner still needs to evaluate and connect them.
The OECD report makes instructional design the dividing line. General-purpose availability is not a learning strategy. Schools must decide which uses preserve effort and which uses erase it.
Schools Now Face an Assessment Problem
Chatbots place the greatest pressure on schools that still equate unsupervised output with individual understanding.
Generative AI did not create every weakness in homework and coursework. It made existing weaknesses easier to exploit and harder to ignore.
Teachers have long known that take-home work reflects uneven levels of assistance. Some students receive extensive family help, private tutoring, or editorial feedback. Others work alone.
AI adds an always-available collaborator that can generate the substance of an answer. The final text may remain difficult to separate from legitimate student work.
Detection software does not solve this problem reliably. False positives can punish original writing, while revised AI text can evade detection. Schools need assessment designs that reveal the student’s process.
The pressure extends beyond deliberate cheating. A student can use a chatbot within a permissive policy and still avoid the practice needed for later performance.
This makes intent an incomplete standard. A well-meaning learner might repeatedly request polished explanations, feel productive, and discover the knowledge gap only during an exam.
The United Kingdom’s qualifications regulator has addressed the integrity side directly. Its 2026 coursework guidance states that submitting AI-generated work as one’s own is cheating.
Yet punishment addresses only one part of the learning problem. Students can follow disclosure rules and still depend too heavily on generated answers.
Schools need to observe competence through multiple formats. Short oral defenses, supervised writing, iterative drafts, practical demonstrations, and live problem-solving can provide stronger evidence.
These formats are not immune to coaching. They do make it harder for a final AI-produced artifact to stand in for the student’s understanding.
Assessment redesign also carries costs. Oral examinations require staff time. Supervised work reduces flexibility. Project-based evaluation can introduce subjective grading and coordination burdens.
Those costs will fall unevenly. Well-funded schools can train teachers and create smaller assessment groups. Systems with staff shortages may rely more heavily on scalable written assignments.
AI vendors face pressure as well. General-purpose products usually optimize for useful, complete responses. That goal can conflict with teaching, which sometimes requires withholding an answer.
An effective educational mode might ask the learner to attempt a solution first. It could provide progressively stronger hints, identify misconceptions, and require explanation before moving forward.
The system would also need age-appropriate privacy protections and transparent data practices. Student conversations can contain academic records, personal concerns, and sensitive behavioral signals.
Teachers need control over curriculum alignment. A chatbot that offers correct but off-syllabus methods can confuse students or undermine a carefully sequenced lesson.
Hallucinations remain another problem. A confident false explanation can teach a misconception, especially when the learner lacks enough knowledge to challenge it.
The OECD reports that about six in ten students had been asked during lessons to assess AI-generated information. Among them, 46 percent did so sometimes, 33 percent often, and 22 percent very often.
Those figures show that evaluation skills have entered many classrooms. They also leave a substantial share of students without regular practice checking AI output.
Academic integrity concerns are already widespread among educators. In the OECD Digital Education Outlook, 72 percent of lower-secondary teachers believed AI could enable students to submit generated work as their own.
At the same time, 37 percent of lower-secondary teachers reported using AI for their jobs in 2024. Another 57 percent agreed that AI helps write or improve lesson plans.
That combination captures the institutional tension. Schools see useful applications for teachers while worrying that students can bypass learning.
Policy must account for both sides. Teachers might use AI to create differentiated exercises, draft feedback, or simulate tutoring conversations. Students still need protected opportunities to struggle productively.
Productive struggle does not mean withholding all assistance. It means giving enough support to move forward without removing the reasoning that builds mastery.
The OECD AI education report therefore shifts responsibility toward institutions and product designers. Telling students to “use AI responsibly” is too vague to guide behavior.
Rules should identify allowed actions, required disclosures, and unaided competencies. They should also explain why some friction remains educationally necessary.
Correlation Is the Central Uncertainty
The PISA comparison is important because of its scale, but its design cannot determine whether chatbot use caused weaker science performance.
Students arrive at AI with different abilities, motivations, and resources. Those differences can influence both adoption and test scores.
A struggling student might use a chatbot more because existing instruction has failed. Lower performance would then precede the AI use rather than result from it.
High-achieving students might avoid chatbots because they already have strong study habits. Their higher scores could reflect those habits, not the absence of AI.
Self-reported frequency adds another limitation. Students may interpret “help me learn” differently, forget occasional use, or underreport behavior that seems prohibited.
The categories also combine many interaction styles. One weekly user might request practice questions. Another might generate an entire assignment during the same reported frequency.
PISA adjusted the science comparisons for socioeconomic background. That improves the analysis, but it cannot control every difference between users and non-users.
The OECD acknowledges this directly. Its report says the relationships might reflect a complex mix of who adopts AI and how they use it.
Headline language should therefore remain careful. “Students using AI scored lower” describes an observed association. “AI made students score lower” makes a causal claim the PISA data cannot support.
The Chinese panel study offers stronger timing and outcome comparisons, but it also remains observational. Researchers did not randomly assign long-term chatbot access across thousands of children.
Controlled trials solve some selection problems. However, they usually cover smaller samples, shorter periods, or highly structured interventions.
The Nigeria trial shows that guided AI tutoring can help under specific conditions. It does not demonstrate that unsupervised chatbot use will produce the same result.
The undergraduate experiment found learning gains when AI augmented explanation. Its participants, subject matter, and controlled environment differ from everyday secondary school use.
These limitations do not cancel the warning. They define what schools should investigate next.
If different data sources repeatedly show stronger assignments but weaker unaided performance, educators have a credible reason to redesign assessment. They do not need to wait for a universal causal estimate.
At the same time, premature certainty can produce poor policy. A complete ban might drive use underground, block beneficial tutoring, and prevent schools from teaching AI literacy.
The more defensible response is controlled adoption. Schools can compare classes, subjects, and tool designs while measuring both assisted output and later unaided performance.
They should also monitor differences across student groups. AI might support learners with limited tutoring access while harming those who use it primarily to avoid effort.
Accessibility needs require special care. Some students use AI for language support, reading assistance, or organizing thoughts. Restrictions should not remove accommodations without equivalent alternatives.
Subject differences matter too. Generating a history essay replaces different mental work from receiving feedback on a mathematics proof.
Long-term measurement is essential. A tool might improve a quiz after one week but weaken retention after a semester. Another might initially slow work while building durable reasoning skills.
The central uncertainty is therefore not whether AI changes learning. It already changes the sequence, effort, and evidence around learning.
The uncertainty concerns which designs produce lasting knowledge, for which students, under which conditions. The OECD findings provide a baseline, not a final answer.
Three Signals Will Show What Happens Next
The next phase should be judged by independent performance, instructional product design, and assessment reform rather than chatbot adoption alone.
The first signal is how researchers analyze the PISA 2025 database. Country-level studies can test whether the association persists across subjects, school policies, and student backgrounds.
A consistent gap after stronger controls would reinforce the OECD warning. Large variation by purpose or school practice would strengthen the case for targeted rules instead of broad restrictions.
Researchers should separate drafting, summarizing, research, feedback, and tutoring. Combining them under general AI use hides the mechanisms schools most need to understand.
The second signal is whether major chatbot providers change educational modes. A meaningful product shift would keep students active instead of optimizing every interaction for immediate completion.
Useful indicators include delayed hints, attempt requirements, teacher controls, curriculum alignment, and explanations tailored to diagnosed misconceptions. Independent trials should test whether these features improve unaided retention.
Marketing claims will not settle the question. Vendors need comparisons that remove the tool after instruction and measure whether students can transfer knowledge to new problems.
The third signal is assessment redesign during the next academic cycle. Schools and examination bodies will reveal their priorities through grading rules, supervised tasks, and disclosure requirements.
More oral defenses, staged drafts, and in-class demonstrations would show that institutions no longer trust final take-home products by themselves. Continued dependence on unsupervised text would preserve the current mismatch.
Schools should publish evidence from those changes where privacy rules allow it. Educators need to know whether redesigned assessments improve validity without creating unreasonable workloads.
Parents and students also need clearer expectations. A rule that merely permits or prohibits “AI” cannot distinguish tutoring from automatic completion.
The most useful standard asks a direct question: can the student still explain, reproduce, and apply the work without the chatbot?
OECD AI test scores have turned that question into an international policy issue. They do not justify panic, and they do not support complacency.
Students should examine how they currently use chatbots. Ask for explanations, counterarguments, practice questions, and feedback on an original attempt. Avoid letting the system replace the first draft of thought.
Teachers should test learning after the tool disappears. Product teams should measure retained competence, not only faster completion or higher satisfaction.
The next round of evidence will show whether schools can preserve AI’s tutoring value without outsourcing the learner’s effort. That outcome depends on design choices being made now.


