The Proctors Are Back. They Are Guarding the Wrong Struggle.
By Firoz Azees
Faced with AI, the most advanced institutions on earth reached for the proctor and the oral exam. A Brazilian state school system chose differently. Your child's school is choosing right now, and most parents cannot see the choice being made.
12 min readOn May 11, 2026, the faculty of Princeton University voted, near-unanimously, to end something that had survived 133 years: the honor-code exam, taken without supervision, on the strength of a student's signed word. Proctors now stand in every exam room. One of the most advanced institutions on earth looked at AI in school assessment and reached for the instrument of the nineteenth century.
I want you to sit with that vote for a moment, because your child's school is making the same choice right now, at smaller scale, with less ceremony. And most parents cannot see the choice being made.
TL;DR: Schools have produced two responses to AI. The first is the fortress: proctors, in-person handwriting, the revival of the oral exam. The second is the bolt-on: an AI-literacy period added to the timetable while every other subject is assessed exactly as before. Both responses protect the old cognitive struggle — the hours of manual drafting AI has already taken. Neither grades the new struggle your child faces every evening: interrogating confident machine text, specifying intent before the machine averages it away, choosing between infinite cheap variants. A third response exists. It is running in a Brazilian state school system and a handful of classroom platforms, and it grades the process instead of the product. Below: what each response protects, the evidence on what unguided AI use does to a learner, and the five questions to put to your child's school.
The weld, briefly
I wrote the first half of this argument in Your Child Is Producing More Than Ever. That Is Not the Same as Learning. The short version: for two hundred years, the submitted essay was the receipt for the thinking, because a child could not produce the essay without doing the wrestling first. AI cut that weld. Production and thinking separated, quietly and completely, and the report card kept certifying as if nothing happened.
That piece asked whether the thinking happened. This piece asks the harder question: what thinking should a school demand now, and is anyone on earth grading it?
Response one: the fortress
Princeton is not alone. UCL's law faculty now requires that at least half of every module be assessed in formats where AI cannot do the work for the student — in practice, in-person written exams and oral examinations, per Times Higher Education's survey of the retreat. The same reporting puts roughly 45% of institutions in some form of assessment redesign. The oral exam, an instrument from the medieval university, is being revived across campuses as an anti-cheating device.
Understand what the fortress protects. The exam hall with a proctor at the door certifies that a student can retrieve vocabulary, organise paragraphs, and draft by hand, alone, against a clock. That was the old cognitive struggle, and for two centuries it was a fair proxy for thinking, because no machine could do it.
No working adult performs that struggle any more. The lawyer the UCL student becomes will never again draft unassisted against a clock, nor will the engineer, nor the doctor writing case notes, nor me writing this. Every one of them works with a machine in the loop, and the quality of their work turns on how they handle the machine. The proctor is a museum guard. The room being watched certifies a skill that working life outside the room dropped years ago.
A three-year study across 290 students and 25 subjects found oral assessment outperforming written work on validity and reliability, and students preferred it. I am not against the viva. A student defending an argument face to face is real assessment. My objection is to what the revival is for: it is being installed to catch cheating, to protect the old proxy, and not to examine what the student did with the machine.
Response two: the bolt-on
The second response looks progressive and dodges the same work. The UAE made AI a mandatory subject across public schools from kindergarten to Grade 12, with around a thousand trained teachers, and no written exams for it — practical projects instead. China mandated AI literacy across every stage of its system, roughly eight hours a year, woven into teacher certification.
I live between the UAE and India. I have watched this rollout with genuine respect, and the respect does not change the diagnosis. AI became a subject that sits beside the curriculum. A new wing on the building. Walk into any other classroom in the same school, English or history or economics, and the essay is still graded exactly as it was in 1995. The load-bearing wall of the whole structure, assessment, is untouched. A child can complete the AI-literacy module on Tuesday and hand in a machine-drafted history essay on Wednesday, and the school has no instrument that notices the difference.
What the evidence says happens in that gap
The MIT Media Lab put EEG caps on 54 students and had them write essays with and without ChatGPT. The group using the model showed up to 55% lower neural connectivity, and 83.3% of them could not repeat a single correct sentence from the essay they had submitted minutes earlier (Kosmyna et al., "Your Brain on ChatGPT"). A student quoted in NPR's coverage of the Brookings report on AI in schools put the mechanism in one line: "It's easy. You don't need to use your brain."
A newer randomised controlled trial with over 1,200 participants went further than the EEG correlations. AI assistance improved performance while in use, then measurably reduced persistence and independent effort afterwards — with the effect setting in within about ten minutes of use.
So the parents' fear is not imaginary, and I will not wave it away. Unguided, an answer-machine drains the effort out of a learner. But read what these studies measured: students using AI the old way, on old-struggle tasks, with nobody teaching or grading the interaction. Measuring cognitive decline there is measuring muscle loss in a man who switched from walking to driving. Of course the muscle goes if he just sits there. The finding is real. The conclusion (go back to walking everywhere, post a guard on the car) is where the elite institutions went wrong.
The struggle did not disappear. It moved up.
Watch a capable person work with AI for one evening and you can name the new struggle yourself. It has three loads, and each one is heavier than it looks:
- Verification. The machine writes 2,000 fluent words with total confidence, including its fabrications. The student's effort shifts from producing the sentence to interrogating it. Where did this claim come from? What did the machine leave out to make the paragraph clean? Sustaining that suspicion against beautiful prose is hard executive work, and nobody is grading it.
- Intent. In the old struggle you discovered what you thought by writing. With a machine, vague intent gets you the average answer — the same smooth essay it hands the next thousand students. Diagnosing why an output is generic and re-specifying what you meant demands a command of the topic that the copy-paste student never builds. Nobody is grading that either.
- Judgment. When four versions of an argument cost nothing to generate, the work becomes choosing — which structure carries the logic, which line is yours and which is the machine's filler. Taste, exercised under abundance. No rubric in your child's school has a row for it.
At Ivanooo we call the capacity that spans these three loads Direction: whether the human is steering the machine or being steered by it. We measure it in professionals. The schools' version of the question is identical. The essay's polish can no longer carry the signal. The interaction can.
Response three: the frontier exists
Here is what the Princeton coverage will not tell you: the redesign the fortress refuses is already running, and not in the places prestige would predict.
| The fortress | The bolt-on | The frontier | |
|---|---|---|---|
| What changes | The exam room (proctors, handwriting, vivas) | The timetable (an AI subject added) | The assessment itself |
| What is graded | The final product, machine-free | The final product, as before | The process: the transcript, the drafts, the interaction |
| The struggle it serves | The old one (manual drafting) | Neither | The new one (verification, intent, judgment) |
| Who runs it | Princeton, UCL, most elite institutions | UAE, China, national curricula | A Brazilian state system, sandbox schools |
In Brazil, the state of Espírito Santo runs essay assessment through a platform called Letrus. Students write inside it. The system gives process feedback while the writing happens and hands teachers a structural view of the whole class. In US and international sandbox schools, teachers using Flint or SchoolAI set up a bounded AI space against their own rubric; the student spars with the machine (brainstorming, counter-arguing, outlining) and the teacher grades the recorded conversation, not the final paragraph. Inside Google Docs, Brisk's Inspect Writing plays back the typing history, so a teacher can watch thought arrive word by word or watch 400 words land in one paste.
The research world has caught up to the same idea this year. One 2026 study treats a student's prompts as "observable externalisations" of thinking in progress; another spent a semester grading students on their AI dialogues and mapped four distinct engagement patterns. The transcript is becoming legible as evidence. The frontier grades it. The fortress shreds it and hands the child a pen.
Hold the inversion in your head, because it is the sentence I most want you to keep: a Brazilian public school system redesigned the essay for the machine age before Princeton did. Prestige is not a leading indicator here. The institutions with the most reputation staked on the old proxy have the least freedom to abandon it. Their nostalgia is structural.
The five questions to ask your child's school
Do not ask "what is your AI policy." Every school has a policy paragraph now. Ask these, in order:
- When my child submits written work, does any part of the grade attach to how the work was produced, or only to the final product?
- If my child used AI on an assignment, can the school show me the interaction (the prompts, the drafts, the changes) or is the only evidence the polished output?
- Which assessments certify that my child can catch a confident machine error? Name one.
- Your school teaches AI literacy. In which other subjects has the assessment changed because of it?
- If two students hand in essays of equal polish, one wrestled the machine through five rounds and one pasted a prompt, does your grading tell them apart?
A fortress school answers with proctoring arrangements. A bolt-on school answers with the AI module's project list. A school on the frontier answers question five without flinching, because its instruments were built for exactly that pair of students.
The proctor at Princeton's door is guarding a room where the old struggle is performed one last time, for the certificate. Outside that room, your child's real cognitive life is happening in a chat window, unwatched, ungraded, and untaught. The school that learns to read that window is the school preparing your child for working life as it now runs. The rest are preparing museum guides. See also What Is Education For When Knowledge Is Free? for the deeper question underneath this one.
FAQ
Should schools ban AI for essays? A ban is unenforceable at home and abandons the child at exactly the skill the working world will demand. The workable alternatives are bounded platforms and process-visible assessment, both running today in real classrooms — banning is the fortress with less honesty.
Is proctoring exams a bad idea? Proctoring certifies unassisted recall and drafting, and certifies it well. The problem is coverage, not validity: a school assessing only inside the fortress certifies nothing about the machine-in-the-loop work that fills adult life. UCL's own formula concedes half the ground.
Doesn't the MIT study prove AI harms learning? It shows that students using ChatGPT the copy-paste way, on traditional essay tasks, engaged less and remembered less. That is a finding about unguided use of the old assignment, not about every possible design. The frontier platforms exist precisely because the old assignment is the problem.
How does a school grade the process instead of the product? Three working forms: writing inside a platform that records drafting (Letrus), grading the student's recorded sparring with a bounded AI (Flint, SchoolAI), and reviewing typing-history playback in the document itself (Brisk). In each, the evidence of thinking is the interaction, not the polish.
My child's school has an AI curriculum. Is that enough? An AI subject beside an unchanged assessment system is the bolt-on. The test is question four above: has any other subject's grading changed? If the history essay is graded as it was in 1995, the curriculum is a new wing on a building whose load-bearing wall nobody touched.
What is the "new cognitive struggle" in one sentence? Verifying confident machine text, specifying your intent precisely enough to escape the machine's average answer, and judging between cheap variants — the three efforts that stay hard when producing prose becomes free.