how-is-ai-used-in-education.md — opusjake_os ARTICLE
// OPUSJAKE BLOG · HOW IS AI USED IN EDUCATION

How Is AI Used in Education? What the Evidence Actually Shows

2026-09-179 MIN READBY · OPUSJAKE
HINTS NOT ANSWERS
how is ai used in educationai in educationai tutoringedtechacademic integrity

How is AI used in education? Four jobs, in practice: teacher prep (lesson plans, worksheets, differentiation), student tutoring and homework help, administrative work (scheduling, translation, IEP drafts), and AI literacy taught as a subject. Teacher prep is the biggest and works. Tutoring is the one with real evidence on both sides, and the evidence splits on a single design choice: whether the tool hands over answers or hands over hints.

That last point is the whole article. Two randomized controlled trials published within weeks of each other in 2025 came to opposite conclusions about AI and learning, and the difference was not the model. It was the prompt behind it.

TL;DR

  • Six in ten US public K-12 teachers use AI for work and three in ten use it weekly, per Gallup's May 2026 poll of 2,069 teachers. Only 18% have received formal guidance on how to use it.
  • A Harvard crossover trial of 194 physics students found more than double the median learning gain from a purpose-built AI tutor versus in-class active learning, in less time (Scientific Reports, June 2025).
  • A trial of nearly 1,000 Turkish high school math students found unrestricted GPT-4 produced 48% better practice performance and then 17% worse exam scores once access was removed (PNAS, June 2025). The guardrailed tutor arm showed no such harm.
  • Detection is not a solution. Seven GPT detectors misclassified human-written TOEFL essays as AI at a 61.22% average false positive rate (Liang et al., Patterns).
  • The policy gap is worse than the tool gap. RAND found 45% of principals reported any AI policy, and over 80% of students said no teacher had taught them how to use AI for schoolwork.

The four jobs, ranked by how well they work

Most coverage of AI in education treats it as one thing. It is four, and they have wildly different success rates.

The four jobs AI does in schoolsFour stacked rows ranking how AI is used in education by evidence quality: teacher preparation, tutoring and homework help, administrative work, and AI literacy instruction. Teacher prep is highlighted as the job with the clearest measured benefit and the lowest risk.// FIG 01 · JOBSFour uses, four very different evidence bases01 TEACHER PREP5.9 hrs/wk saved, weekly usersPlans, worksheets, rubrics. Low risk.02 TUTORING2x gain or 17% lossOutcome depends on the guardrail.03 ADMINdrafts, translation, schedulingWorks. Human signs every output.04 AI LITERACY<20% of students taughtSmallest today. Biggest gap.Adoption is concentrated in row 1. The argument is all about row 2.

Teacher prep is the quiet majority of real usage. Gallup's earlier June 2025 survey of 2,232 teachers put the top tasks as preparing to teach (37%), making worksheets and activities (33%), and modifying materials for individual student needs (28%), and found weekly users saved an estimated 5.9 hours a week. That is roughly six weeks of reclaimed time across a school year. No learning-outcome controversy attaches to a teacher generating four reading levels of the same passage in ninety seconds.

Tutoring is the contested one, covered below.

Administrative work is the least discussed and probably the second most valuable: translating a newsletter into seven home languages, drafting the first version of an IEP, turning a messy observation note into structured feedback. The rule is the same as in any business deployment, which I laid out in how to use AI in business: a human signs the output, every time, and the AI never has final authority over a record about a child.

AI literacy as a subject is the smallest and the most obviously underserved. Stanford HAI's 2026 AI Index education chapter reports that more than 90% of countries now offer computer science at primary or secondary level, and that China and the UAE mandated AI education starting in the 2025-26 school year. US coverage is thinner and uneven by state.

The adoption numbers disagree, and that is useful

If you search this question you will find three different national figures and no explanation of why.

  • Pew Research Center surveyed 1,458 US teens aged 13 to 17 between September 25 and October 9, 2025, and published in February 2026 that 54% had used chatbots for schoolwork help. Ten percent do all or most of their schoolwork that way. More than 40% use them for researching a topic or solving math problems.
  • RAND's survey panels found 54% of middle and high school students and 53% of ELA, math, and science teachers used AI for school in 2025, both up more than 15 points in a year.
  • Stanford HAI's 2026 AI Index reports four out of five US high school and college students now use AI for schoolwork.

These are not contradictions. Pew and RAND measure secondary students; the AI Index pools a broader population that includes undergraduates, where usage runs far higher. The practical read: in a high school classroom, assume roughly half your students have used AI on schoolwork and a tenth are leaning on it heavily. In a university lecture hall, assume nearly everyone.

Two trials, opposite results, one variable

Here is the part worth reading twice.

In June 2025, Kestin, Miller, Klales and colleagues published a randomized crossover trial in Scientific Reports. One hundred ninety-four Harvard physics students each experienced both conditions across two weeks of material on surface tension and fluid flow: an in-class active learning session led by experienced instructors, and an at-home session with a custom AI tutor called PS2 Pal. Median learning gains in the AI condition were more than double, and students spent less time. Note the control: active learning is already the research-backed best practice in physics instruction, not a lecture.

Also in June 2025, Bastani and colleagues published a field experiment in PNAS with nearly 1,000 Turkish high school math students across three arms. During AI-assisted practice, the unrestricted GPT-4 arm scored 48% better than control and the guardrailed tutor arm scored 127% better. Then the researchers took the AI away and gave everyone an exam. The unrestricted arm scored 17% worse than students who never had AI at all. The guardrailed arm scored about the same as control.

Raw chatbot versus guardrailed tutorTwo side-by-side panels comparing an unrestricted GPT-4 chatbot against a tutor constrained to give teacher-designed hints, showing practice performance versus exam performance once AI access is removed, based on the Bastani et al. PNAS 2025 field experiment with high school math students.// FIG 02 · GUARDRAILThe same model, two prompts, opposite outcomesRAW CHATBOTgives the answerDuring practice+48%Exam, AI removed-17%VERDICT · worse than no AIHINT TUTORteacher-written hints onlyDuring practice+127%Exam, AI removedflatVERDICT · no learning harmWithholding the answer is the entire intervention. It costs one paragraph of prompt.Source: Bastani et al., PNAS 122(26), June 2025. ~1,000 Turkish high school math students.

The authors describe the failure mode plainly: without guardrails, students use the model as a crutch. They look fluent while the scaffolding is there and collapse when it goes. The fix that worked was not a better model. It was instructing the tutor to give teacher-designed hints instead of answers. The Harvard tutor was tuned the same direction: brief replies, a few sentences at most, to avoid dumping a solution.

Two caveats before anyone quotes this at a board meeting. The Harvard sample was 194 motivated students at one elite university, in proctored Zoom sessions, across two lessons. The Turkish study was a different country, subject, and age group. Neither generalizes cleanly. What does generalize is the mechanism, and it matches everything known about productive struggle: learning happens in the gap between not knowing and knowing, and a chatbot that closes that gap instantly deletes the learning along with the difficulty.

Why detection is not the answer

The obvious reaction to student AI use is to detect it. That does not work, and the way it fails is unfair.

Liang and colleagues at Stanford tested seven widely used GPT detectors against 91 human-written TOEFL essays and published the results in Patterns. The average false positive rate was 61.22%. At least one detector flagged 89 of the 91 essays as AI-generated. Native-speaker essays were classified near-perfectly. Prompting the model to enrich the vocabulary of those same TOEFL essays dropped the false positive rate to 11.77%, which tells you exactly what the detectors are measuring: not authorship, but how native the prose sounds.

The downstream cost is already visible. RAND found half of surveyed students worry about being falsely accused of using AI to cheat, with high schoolers more worried than middle schoolers. Running a classroom on a 61% false positive instrument against your multilingual students is not an integrity policy. It is a liability.

Treat a detector score the way you would treat an anonymous tip: a reason to open a conversation, never the conversation's conclusion. Ask the student to walk you through their draft history or explain a choice they made in paragraph three. That takes four minutes and it is actually diagnostic.

The gap is governance, not technology

The tools are in the building already. The instructions are not.

Gallup's February-March 2026 survey, fielded through the RAND American Teacher Panel with 2,069 US public K-12 teachers, found only 18% had received any formal guidance from administrators on how AI tools should be used. Thirty-four percent received no guidance at all, and 48% got informal verbal guidance only. Among teachers who did receive guidance, most said it neither encouraged nor discouraged AI use, which is guidance in the same sense that a shrug is an answer.

RAND's panels found the same shape from the other direction: 45% of principals reported school or district policies on AI use, and only 34% of teachers reported a policy covering academic integrity specifically. The 2026 AI Index adds the sharpest version of it: about half of middle and high schools have an AI policy, and just 6% of teachers say those policies are clear.

Meanwhile 35% of district leaders said they provided students any AI training, and over 80% of students reported that no teacher had explicitly taught them how to use AI for schoolwork. So the median student is using a tool nobody showed them how to use, under a policy nobody wrote down, graded by a teacher who got a verbal shrug about it.

That is a governance problem, and governance is cheap compared to procurement. The same control-first thinking applies here as anywhere else you deploy this stuff, which I worked through in responsible AI.

What to do this term

Concrete, in order, for whoever is reading this.

If you teach: label every assignment with one of three tags before you hand it out. AI-permitted (use it, cite it). AI-assisted (use it for brainstorming and editing, disclose what you used it for). AI-free (done in class, on paper or on a locked device). You do not need a district policy to do this. You need one line on the assignment sheet, and it removes most ambiguity your students are currently guessing at.

If you build or configure a tutor: prompt it to withhold answers. Hints, one step at a time, questions back. Keep replies to a few sentences. This is the single highest-leverage configuration decision in the whole field, and it is a paragraph of system prompt. The structure of that paragraph matters more than most people think, and I broke down how to write one in how to write AI prompts.

If you are a student: the finding that should worry you is the 17%. Using a chatbot to finish the problem set feels like progress and measures as regression on the exam. Make it ask you questions instead of answering yours. I collected the specific phrasings that work, plus the study and research patterns worth keeping, in STUDENT CODES.

If you run a school: write the assignment-level taxonomy above into a one-page policy, train teachers on it (optional training is how you get 18%), and verify every claim a vendor makes against a named study. Also expect fabricated citations in anything a student or staff member generates, which is a known failure mode with known mitigations: see how to reduce AI hallucinations.

The bottom line

AI in education is not one intervention with one result. It is four jobs, and the evidence sorts them clearly. Teacher prep saves real hours and nobody has shown it harms anyone. Admin work is fine with a human signature. AI literacy is underbuilt almost everywhere. And tutoring, the job everyone argues about, produces double the learning gains or a 17% deficit depending on one design decision: whether the thing hands over the answer.

Everything else in this debate is downstream of that decision. Schools that get it right will not be the ones that bought the most licenses. They will be the ones that wrote down what the tool is allowed to say.

If you want the working version of this, not the survey version: I send one practical AI build every week to people who ship things. Join the newsletter. And if you are the one doing the assignments, start with STUDENT CODES, the prompt set that makes the model quiz you instead of doing it for you.

// FREQUENTLY ASKED
How is AI used in education right now?

Four jobs, in descending order of how well they work. Teacher prep is the biggest and least controversial: lesson plans, worksheets, differentiated versions of the same material, rubrics, and parent emails. Gallup's May 2026 poll of 2,069 US public K-12 teachers, fielded through the RAND American Teacher Panel, found six in ten teachers use AI for work and three in ten use it weekly. Second is student tutoring and homework help, which is where the outcome evidence gets complicated. Third is administrative work: scheduling, IEP drafting, translation for multilingual families, and grading support. Fourth, and smallest, is AI literacy taught as its own subject. Almost everything a vendor pitches as transformation is one of those four with a new interface on top.

Does AI actually improve learning outcomes?

It depends entirely on whether the tool gives answers or gives hints, and two randomized trials make the split obvious. Kestin and colleagues at Harvard ran a crossover trial with 194 physics students, published in Scientific Reports in June 2025, and found students using a purpose-built AI tutor showed median learning gains more than double the in-class active learning condition, in less time. Bastani and colleagues ran a trial with nearly 1,000 Turkish high school math students, published in PNAS in June 2025, and found that students with unrestricted GPT-4 scored 48% better during practice but 17% worse than the control group on an exam once AI access was removed. Same technology, opposite results. The Harvard tutor was tuned to be brief and Socratic. The harmful arm was a raw chatbot.

What percentage of students use AI for schoolwork?

The honest answer is that it depends on how you ask, and the spread is wide enough to matter. Pew Research Center surveyed 1,458 US teens aged 13 to 17 in late 2025 and published in February 2026 that 54% had used chatbots for schoolwork help, with 10% doing all or most of their schoolwork that way. RAND's nationally representative panels found 54% of middle and high school students reported using AI for school in 2025. Stanford HAI's 2026 AI Index puts the figure at four in five US high school and college students, because it pools a broader population including undergraduates. Use the number that matches your population. A single national percentage is close to meaningless for any individual classroom.

Are AI detectors reliable for catching cheating?

No, and the failure is not random. Liang and colleagues at Stanford tested seven widely used GPT detectors against 91 human-written TOEFL essays and published in Patterns in 2023 that the detectors misclassified them as AI-generated at an average false positive rate of 61.22%. At least one detector flagged 89 of the 91 essays. The features that read as machine-written, low sentence-length variance and predictable word choice, are also the features of competent second-language academic writing. The cost lands on real students: RAND found half of surveyed students already worry about being falsely accused of AI cheating. Detector output is a prompt to have a conversation, never evidence on its own.

What should a school do first about AI?

Write an assignment-level policy before buying anything. RAND found 45% of principals reported any school or district AI policy and only 34% of teachers reported one covering academic integrity. Gallup's 2026 data is starker: 18% of teachers had received formal guidance, and most guidance that existed neither encouraged nor discouraged use. That vacuum is the actual problem, not tool selection. The fix costs nothing. For each assignment, label it AI-permitted, AI-assisted with disclosure, or AI-free and done in class. Then teach students the difference, which most schools skip. Over 80% of students told RAND no teacher had explicitly shown them how to use AI for schoolwork.

// BUILD WITH OPUSJAKE

OpusJake is Jake Schincariol's operating system for building with AI: agents, workflows, prompts, and the free resources behind them. Get the next move every week.

STATUS · ONLINE · OPUSJAKE © OPUSJAKE // CRT V1