No independent test has found an AI writing detector accurate enough to accuse a student on. The largest published evaluation found every tool it tested scored below 80% accuracy. A Stanford study found detectors wrongly flagged 61.3% of essays by students writing in a second language. Curtin, UQ and the ANU have all now switched detection off.
What the testing found
The most cited evaluation is Weber-Wulff et al., published in the International Journal for Educational Integrity in 2023. It put fourteen detection tools, Turnitin among them, through a controlled test. Not one reached 80% accuracy. Only five got past 70%. Six of the fourteen produced false positives — human writing labelled as machine-written.
The authors’ own summary is that the tools are “neither accurate nor reliable”. That is a direct quote from a peer-reviewed paper about the entire product category, not a vendor comparison or a blog post.
Worth knowing about their bias direction, because it is the opposite of what most people assume: the tools leaned toward calling text human. They miss AI writing more often than they invent it. The false positives are the rarer error, which is exactly why they are so damaging when they land — nobody is expecting one.
Who gets wrongly flagged
In 2023 a Stanford team ran seven widely used GPT detectors over 91 TOEFL essays, all written by humans, all by students whose first language is not English. The detectors wrongly flagged 61.3% of them on average. Nearly one in five was flagged by every detector in the test. Over 97% were flagged by at least one.
The same detectors were close to perfect on essays written by native-speaking American eighth-graders. The difference is not that one group cheated. It is that detectors work by measuring how predictable the writing is, and someone writing carefully in a second language produces exactly the flat, regular prose the tools associate with a machine.
That mechanism catches other students too. Anyone who writes in short, plain sentences produces low-variability text. Researchers have warned that neurodivergent students are more likely to be hit by false positives, and Bloomberg documented the case of an autistic student flagged and given a zero. Neither of those is a large empirical study, so treat it as a warning rather than a measurement — but it is a warning from people who have looked closely.
At the false-positive rate Turnitin publishes for itself, a university submitting 75,000 papers a year is wrongly accusing about 750 students.
What Turnitin says about Turnitin
This is the part worth reading twice, because it comes from the vendor rather than a critic. Turnitin publishes a false-positive rate of “under 1%” for documents containing more than 20% AI writing. At the level of individual sentences, the figure it gives is around 4% — roughly one highlighted sentence in twenty-five may be human-written.
And in its own FAQ: the percentage on the AI writing indicator “should not be used as the sole basis for action”. The company selling the detector is telling schools not to act on the number alone. That instruction is frequently the first casualty once the score is in front of somebody.
One more piece of context. OpenAI built a classifier to detect its own model’s output and withdrew it in July 2023, citing low accuracy. It caught about a quarter of AI-written text and falsely flagged nearly one in ten human samples. The company with the most training data and the strongest commercial reason to solve this could not.
Institutions that have switched detection off
| Institution | Decision | Effective |
|---|---|---|
| Vanderbilt University (US) | Disabled Turnitin’s AI detector | August 2023 |
| Australian National University | Not used for academic integrity matters | January 2024 |
| Australian Catholic University | Stopped using the AI indicator (per ABC News) | March 2025 |
| University of Queensland | AI Writing Indicator withdrawn | Semester 2, 2025 |
| University of Waterloo (Canada) | Detection tool withdrawn | September 2025 |
| Curtin University | AI writing detection disabled | 1 January 2026 |
Curtin is the one that matters most to a WA school, because it is the university a good share of your Year 12s will walk into. From the start of this year they are marking without it. Waterloo went further and said why: their own IT testing found the product flagging human-written text as “100% generated by AI”, more than once.
None of these institutions decided that AI use stopped being a problem. They decided the detector was not evidence.
What works instead
- Treat a score as a reason to look, never as a findingA flag is a prompt to open a conversation. It is not a result, it cannot be shown to a parent as proof, and on the vendor’s own instruction it cannot be the sole basis for action.
- Design the task so the process is visibleDrafts, planning notes, an in-class writing stage, a short verbal follow-up on the argument. Work that shows its own development is far harder to outsource and far easier to defend if it is ever questioned.
- Say what is allowed, in writing, per taskMost students who use AI on work where it was not wanted were never told clearly. A one-line statement on the assessment sheet — what is permitted, what must be declared — removes most of the ambiguity before it becomes a case.
- Ask students to declare use rather than hide itA disclosure line turns AI use into something discussable. Detection turns it into something to conceal, which makes the concealment better rather than the use rarer.
- Never open a misconduct process on a percentage aloneThe Australian Catholic University processed thousands of flags in a year and later dismissed many of them, according to ABC reporting. The cost of that lands on students, on staff time, and on how much your families trust the school.
Why we publish this
We run AI training for schools, and we could sell more of it by leaving this vague. We have told this to every staffroom we have stood in, so it may as well be written down where a Head of Learning Area can check it before booking anything.
Every figure above is linked to its source below. If any of it changes, the page changes.
Sources
- Testing of detection tools for AI-generated text (Weber-Wulff et al.) — International Journal for Educational Integrity, vol. 19, art. 26 (2023)
- GPT detectors are biased against non-native English writers (Liang, Yuksekgonul, Mao, Wu, Zou) — Patterns (Cell Press), vol. 4, art. 100779 (2023)
- Guidance on AI detection and why we’re disabling Turnitin’s AI detector — Vanderbilt University (2023)
- Update on the Turnitin AI detection tool — Curtin University (2025)
- Turnitin Similarity Report and AI Writing Indicator update — The University of Queensland (2025)
- Discontinuing use of AI detection functionality in Turnitin — University of Waterloo (2025)
- Turnitin’s AI writing detection capabilities — FAQs — Turnitin (2026)
- Understanding the false positive rate for sentences of our AI writing detection capability — Turnitin (2023)
- AI detection tools falsely accuse international students of cheating — The Markup (2023)
Common questions
- Can Turnitin detect ChatGPT?
- Sometimes, and not reliably enough to accuse anyone on. Turnitin publishes a document-level false-positive rate under 1% and a sentence-level rate around 4%, and states the score should not be the sole basis for action. Independent testing put every tool it evaluated below 80% accuracy.
- A student has been flagged. What should we do?
- Treat it as a prompt to talk, not a finding. Ask to see drafts and planning, and ask the student to talk through their argument. If there is no evidence beyond the percentage, there is no case — that is the vendor’s own position as much as ours.
- Are detectors unfair to some students more than others?
- Yes, and this is the best-evidenced problem with them. A Stanford study found seven detectors wrongly flagged 61.3% of human-written essays by students writing in a second language, while being near-perfect on native-speaker work. Anyone who writes in short, plain sentences is at raised risk for the same reason.
- If we cannot detect it, how do we assess fairly?
- By making the process visible rather than policing the output — drafts, in-class stages, a short verbal follow-up. It is more work to set up once and considerably less work than running misconduct cases you cannot substantiate.
- Does this mean students should be allowed to use AI freely?
- No. It means the rule has to be stated per task and the assessment designed around it. The schools handling this best are specific about what is permitted and ask for it to be declared, rather than relying on software to catch what was never clearly forbidden.