This happens to people who did nothing wrong

If a piece of software has told somebody you did not write your own work, the research is on your side. In 2023 a team at Stanford ran essays from the TOEFL exam through seven of the popular detectors, and more than 60% of them came back as machine written [Liang et al., 2023]. The same year another group tested fourteen tools and found that none of them reached 80% accuracy [Weber-Wulff et al., 2023]. People have pasted the U.S. Constitution into these things and been told a machine wrote it [reported widely, 2023–2026].

The writing they get wrong most often is the writing teachers ask for. Correct grammar, a formal structure and a careful second draft all push the score up, so the student who worked hardest on the essay is the one most likely to be flagged by mistake. If the score is all they have, the accusation rests on that mistake.

What we would do, in order

  1. Do not reply yet. Whatever you write while you are angry will be read later by somebody who was not in the room. Give it a few hours, or sleep on it.
  2. Save the accusation the way it arrived. Keep the message itself, the name of the software they used, the number it gave them and the date they ran it. Vague claims get smaller once the specifics are written down.
  3. Open up your version history. In Google Docs it is under the File menu, and in Word you have it if AutoSave was on while you worked. What you want is the whole timeline of the essay being written: the rough first paragraph, the section you deleted and wrote again, the two hours in the middle where nothing moved because you went out to get food. That timeline is the strongest proof of authorship you have, and no scanner produces anything like it, for you or against you.
  4. Gather everything else you have. Any outline you made, the notes you took, the sources you saved as you went along, earlier drafts, your browser history from the nights you were writing, and some older graded work, so that whoever reads all of this can see what your writing normally looks like.
  5. Ask for the evidence, in writing. Ask which software they used, what score it gave, and what the written policy says that score is supposed to mean. Keep it polite. A lot of accusations end right here. The truthful answer is “a probability from a piece of software with a known error rate”, and not many people want to put that in an email.
  6. Cite the research. The two studies at the top of this page are linked at the bottom, and several universities switched their detectors off after false accusations made the news, which you can name as well. You are pointing at numbers that were published. Anybody reading your appeal can go and read them too.
  7. Appeal if you have to. Academic integrity offices and HR departments both have a second step. Use it if the first conversation goes nowhere. Bring the folder, and keep the tone you used in the first email.

A reply you can adapt

“I wrote this myself and I would like to settle it with evidence. I can share my full version history and my drafts, which show the document being written over several days. Could you tell me which detection tool flagged it and what score it reported? The published research shows that these tools have high false-positive rates on formal writing and on non-native English, so I would like us to look at the process evidence together before any decision is made on the score alone.”

If this happened at work

The steps are the same, with one difference. Most employers have no written policy on this, so ask what standard was applied to your work and if anybody told you about it beforehand. The edit history in a shared document, a commit log or the revisions tab in your CMS does the job a student's version history does. In 2025 a Cornell study of the freelance platforms found that suspicion did not fall evenly, and that the polished, formal writers were suspected most [Cornell, 2025]. So bring the edit history into the conversation early, before anybody has settled on how the writing feels to them.

Run your own check first

Before you answer anybody, paste your text into the Cobalynx checker. What comes back is a probability with a published error rate behind it, and when the evidence is thin the result is “inconclusive” and does not pick a side. A dated result from a tool that publishes its false-positive rate is something you can put in the folder. Leave out any screenshot from a tool that claims to be certain, since certainty is the part the research does not support.

Two scanners often disagree about the same essay, and we have written up why detectors disagree with each other. Our own error rates are on the page that reads them straight out of the evaluation files, so cite the figure with the date and engine version printed beside it, because those numbers change when the engine does. And if nobody has accused you of anything and you are just nervous the night before a deadline, the student guide is the calmer version of this page.

Three words that settle most of these meetings

The people across the table will use words that sound technical and never get defined, so learn three of them, with a definition you can quote back. A false positive is human writing that was wrongly flagged, which is most likely what happened to you. A calibrated probability is what a score has to be before 80% means 80%. The base rate is why a 1% error rate still produces dozens of wrong flags in a school where most students wrote their own work.

What Cobalynx can and cannot do for you

We cannot prove that you are human. No tool can, and any tool that offers to certify a text as “human-verified” is claiming more than the science allows. What we can give you is a second opinion with published error rates you can cite, a page that explains how the score is produced, and this guide to the process evidence that shows who wrote something. No detector, ours included, replaces the version history you already own.

If you are the one about to accuse someone

Teachers and managers read this page as well, usually the evening before a hard conversation. What a score can and cannot support, and how you keep false accusations close to zero, is in the guide for teachers. Anything else is probably in the FAQ.

Sources

  1. Liang, Yuksekgonul, Mao, Wu, Zou, “GPT detectors are biased against non-native English writers,” Patterns (2023), arxiv.org/abs/2304.02819 — 61.3% average false-positive rate on TOEFL essays across seven detectors.
  2. Widely reproduced detector false positives on the U.S. Constitution and other historical texts; see the overview in Wikipedia, “Artificial intelligence content detection” (accessed Aug 2026).
  3. Cornell University research on bias in who gets suspected of AI use on freelance platforms (2025; accessed Aug 2026).
  4. Weber-Wulff et al., “Testing of detection tools for AI-generated text,” International Journal for Educational Integrity (2023) — all 14 tested detectors below 80% accuracy, arxiv.org/abs/2306.15666.

Common questions

Can a detector score alone get me failed or fired?

It should not, and more and more often it does not, because a detector score is a probability from a piece of software with a measured error rate, and several universities have already switched their detectors off after false accusations that made the news. If the score is the only evidence against you, then ask in writing for the name of the software, for its published false-positive rate, and for someone to look at your drafts and your version history.

Does Grammarly or heavy editing make writing get flagged as AI?

Yes, it can, and this catches a lot of people out. When you run your writing through Grammarly or you edit it very heavily, what you end up with is smoother and more formal than what you started with, and that smoother, more formal writing is exactly the kind of thing the detectors were trained to call AI. It happens most of all to people whose first language is not English. It is a false positive, and it is one more reason why your version history is worth more than any score.

If I check my essay myself before submitting, am I safe?

Not completely, no. Scanners are not calibrated to each other, so a low score here does not make your school's software agree with it. What a Cobalynx scan gives you is a probability with a published error rate behind it, on the evidence page, and a dated result you can put in your folder. It is worth having, and it still does not guarantee anything about what another tool will say.

What should I say to my professor or manager first?

The first thing to say is something like this: I wrote this myself, I can show you the drafts and the version history, and I would like to know what made you think it was AI. You do not need to argue about the score at that point, since what you are really doing is asking them to explain their process, and most of the time that is where the whole thing starts to fall apart for them.