What went wrong when schools treated the number as a verdict
Over the last couple of years several universities have switched off the AI detection in their plagiarism software, after students who had done nothing wrong were accused on the strength of a percentage. A number came out of a tool that had never been calibrated, a teacher or a committee treated it as a verdict, and no process around it caught the mistake. A score is a fair reason to look closer at a piece of work. The harm starts when it becomes the reason to punish, and the rest of this page is the process that keeps those two apart.
What a calibrated result tells you
When a tool prints “AI DETECTED” in capital letters, it has told you nothing about how often that alarm goes off for no reason. What Cobalynx gives you is a score counted on a population, which means we can tell you how often the texts that scored like this one turned out to be AI on the frozen set we test against. It is meant to be a second opinion inside your own process, and you should not act on it by itself.
We also publish how often each confidence band has been wrong on that same frozen set, with the raw counts and the confidence intervals. A high confidence flag still comes with a false positive rate that we have counted and that is not zero. When we say inconclusive, we mean we could not tell, and it is not a polite way of saying probably AI. When we say likely human, that is no proof either, since no statistical method can certify that a person wrote something. How the scores get produced is on the methodology page.
Four kinds of evidence we would look at before the score
Every serious integrity framework we have read ends up in the same place, which is that a detector score may open a conversation, never close one. So before you act on any flag, ask the student for the things below, in that order.
- Version history. Ask the student to share the revision timeline of the document, which is under File in Google Docs and comes with AutoSave in Word. Hours of small edits building up over days is strong evidence that they wrote it. One big paste is something to ask them about.
- Earlier work. Ask to see something the same student wrote earlier in the year, and read the two side by side. You are looking for the voice and the level you remember. People do get better over a term, but they do not usually turn into a different writer between one essay and the next.
- Drafts and notes. Ask if they kept an outline, notes, or a list of the sources they looked at, since real work leaves that kind of thing lying around and it is hard to fake after the fact.
- A conversation. Ask the student how they went about the essay, what they were trying to argue, why they picked the sources they picked and what they left out. Somebody who wrote it can tell you, and this one conversation usually settles the case.
If all four point the same way, you have a case that rests on how the work was done, and that is a case you can stand behind. If they do not, the score was wrong about this student, and it is much better to find that out at this stage than after you have accused them.
The bias you need to know about
The mainstream tools flag writing by people whose first language is not English far more often than they should. In 2023 a team at Stanford put TOEFL essays through seven detectors and found that over 60% of them were wrongly flagged [Liang et al., 2023]. Careful, formal prose that has been run through a grammar tool, which is what your most conscientious students hand in, gets flagged for the same reason. We count our own false positive rate on non-native (ESL) writing, we publish it, and we hold it to the same gate as every other kind of writing. Ask whichever company you buy from for that number, and if they do not publish one, give their flags less weight.
A fair process, written down for your syllabus
Put a line up front in the syllabus saying what kind of AI use is allowed, and say that flagged work leads to a conversation first and not straight to a penalty, so that nobody is surprised later on. Do not act on a score from any tool on its own, and that includes ours. Ask the student for evidence of how they wrote it before you make up your mind. Take extra care with the students who write English as a second language, since the research says the tools get them wrong more often. And write down what the tool told you, meaning the probability it gave and the error rate the company publishes for that probability. If all you write down is the word flagged, in six months nobody is going to remember what it meant. Your students have their own version of this page, the false-accusation playbook. It tells them to offer the same evidence listed here and to ask you what raised the concern.
What Cobalynx can and can't do here
Nothing that comes out of a detector, ours included, is proof that a student did anything wrong. We cannot tell you who typed the words, and we are not going to pretend that we can. What we can do is give you a score with an error rate we have counted and published, so it can be one part of a conversation, and refuse the job when the text is too short to judge. Why two detectors will disagree about the same text, and what a believable accuracy claim looks like, is in Can AI detectors be trusted?, and the tool itself is free, with no signup.
Sources
- Liang, Yuksekgonul, Mao, Wu, Zou, “GPT detectors are biased against non-native English writers,” Patterns (2023), arxiv.org/abs/2304.02819.
- Weber-Wulff et al., “Testing of detection tools for AI-generated text,” International Journal for Educational Integrity (2023), arxiv.org/abs/2306.15666.
- Public university decisions to disable AI-detection features after false-accusation concerns (UCLA, UC San Diego, Vanderbilt, among others; accessed Aug 2026).
Common questions
Is an AI-detection score enough to fail a student?
No, it is not, and that goes for our score as much as anyone else's. A score is a probability from a piece of software that has a measured error rate, and the institutions that treated it as a verdict ended up with false accusations that made the news. Before you draw any conclusion we would put the score next to the student's version history, their earlier work, and a conversation with them about what they wrote.
What does calibrated confidence tell me?
It tells you two things, and we would keep them apart in your head. The first is how often the texts that scored the way yours did turned out to be AI when we ran our frozen test set, and the second is how often we were wrong when we gave that same level of confidence, so that when we tell you we are not sure, we mean that we could not tell and we are not dressing it up as an alarm. How much AI writing there was in your classroom to begin with matters as well, since that changes what a flag means for you, and we have put the error rates for each band and the base-rate explainer on the evidence page so you can work it out for your own class.
Why did several universities stop using AI detectors?
Because the scores were being used as verdicts, and the false positives, which landed hardest on students writing in a second language and on careful, formal writers, were doing real harm to real people. The lesson we would take from it is that a score from a tool that was never calibrated, used inside a process that was never fair, is worse than having no score at all, and we do not think it means that detection is useless.
How do I handle a flagged non-native speaker fairly?
With more care than usual, because the research, including a widely cited Stanford study from 2023, shows that the mainstream tools flag writing by non-native speakers far more often than they should. We would put the flagged work next to the student's earlier writing, ask to see their drafts, and give the evidence of how they wrote it more weight than the score. We measure our own error rate on non-native writing and publish it on the evidence page for this reason.