The number in a Turnitin AI Writing Report is widely misread. Students confuse it with the similarity score, and institutions treat it as proof. This guide explains what the percentage measures, what Turnitin itself publishes about its error rates, and where the tool is documented not to work.
Four things about the AI Writing Report that account for most of the confusion around it.
The AI writing percentage is separate from and independent of the similarity score, and AI highlights do not appear in the Similarity Report at all. Most questions along the lines of 'is 25% bad?' are really about the similarity score, which measures matched sources and is a completely different thing.
Since July 2024, Turnitin does not surface any score above 0% and below 20%. It shows an asterisk instead. Their stated reason is that testing found a higher incidence of false positives in that band, so the number was withheld rather than shown.
The percentage covers prose sentences in long-form writing. Turnitin states the model does not reliably detect AI text in poetry, scripts, code, bullet points, tables or annotated bibliographies, so a mixed-format document can show a percentage that does not match its highlights.
A submission must contain at least 300 words of prose and no more than 30,000 to generate a report at all. Short assignments, discussion posts and problem sets fall outside what the tool is built to assess.
These are Turnitin’s figures, not ours. They are worth reading closely, because the document-level number and the sentence-level number are very different.
Turnitin's stated rate for incorrectly flagging a fully human-written document, but only for documents scored at 20% AI writing or above.
Turnitin's stated likelihood that any individual sentence highlighted as AI-written was in fact written by a person.
Turnitin reports that more than half of falsely highlighted sentences appear directly next to genuine AI writing, most often at the transitions.
In August 2023 Vanderbilt University disabled Turnitin’s AI detector and explained the arithmetic behind the decision: the university submitted 75,000 papers in 2022, so a 1% false positive rate would mean roughly 750 papers wrongly flagged in a single year. They also cited a lack of transparency about how the model works, and research showing detectors are more likely to flag writing by non-native English speakers.
Turnitin’s own instructor guidance is consistent with that caution. It states the model “may not always be accurate (it may misidentify human-written, AI-generated, and AI-paraphrased text), so it should not be used as the sole basis for adverse actions against a student,” and advises using a highlighted result “to initiate a conversation, not to draw a conclusion.”
Sources: Turnitin, Using the AI Writing Report and Turnitin, Understanding AI writing detection: false positive rates.
Turnitin's guidance and ours agree on the shape of this: the score opens a conversation, it does not settle one.
Confirm whether the number is the AI writing percentage or the similarity score. They measure different things, appear in different reports, and a high similarity score usually means quoted or cited sources rather than AI use.
Read the specific sentences the report marked. If the flagged passages are formulaic transitions, method descriptions or plain factual summary, that is the writing style most likely to be misread as AI.
Version history, drafts, notes and outlines are stronger evidence of authorship than any detector score in either direction. Ask the student to talk through how the piece came together.
A second detector built on different models is a useful cross-check. Where two independent tools disagree, that disagreement is itself information, and it is a reason to slow down rather than act.
Not a substitute for your institution's process, and not a way to predict its output.
Our checker uses its own detection models rather than reproducing Turnitin's. That is the point of a second reading: two tools that fail in the same way tell you nothing, two that fail differently tell you where to look.
Because the result is not locked behind an institutional licence, a student and a teacher can look at the same output together and talk about the specific sentences it marked.
Every sentence carries its own score, so partial AI use in an otherwise original piece is visible rather than averaged away into a single document percentage.
Heavily rewritten AI is the case we handle worst, and the checker is English-only. Both are documented on our technology page rather than left for you to discover.
How to use an AI checker in a classroom without turning a score into an accusation.
AI detection for higher education admissions and research integrity.
The models behind the score, our measured accuracy, and the cases we get wrong.
Paste any text and get a verdict, a confidence score, and the specific sentences behind it. Useful as a second reading alongside an institutional report, not as a prediction of one.
Free to try. No account, no card.