Start with what the critics have got right

People have pasted the U.S. Constitution into these tools and been told it was written by a machine, and the same has happened with Macbeth and with the Book of Genesis [reported widely, 2023–2026]. In 2023 a group of researchers tested 14 of the popular detectors side by side and found that none of them reached 80% accuracy [Weber-Wulff et al., 2023]. That same year a team at Stanford found that the mainstream tools wrongly flagged more than 60% of the essays written by people whose first language was not English [Liang et al., 2023]. Our own testing says paraphrasing makes our own detector worse, and the published research reports the same across the field. All of that is true, and we would be wary of any company selling a detector that will not say so on its own site.

So why does every product page say 99%?

Most of them are quoting a number with no test conditions attached, and an accuracy number without its test conditions is unfalsifiable marketing, which means neither you nor the school about to pay for it can test it. Think about what would have to be true for that 99% to mean anything. You would need the length of the texts, since a tool that does well on a 2,000-word essay can fall apart on a 100-word email. You would need the language, and which model wrote the machine ones, since a tool that learned on 2023 output has no promise of doing as well on what the models write in 2026. You would need to know where the company put its threshold, and how much a person edited the machine text before it was tested. Change any one of those and the same tool can look brilliant on one test and barely better than a coin toss on the next. A company that picks a test set that suits it, and a threshold that suits it, and then prints one percentage on its home page has not necessarily lied to you. It has given you a number nobody outside the company can test. The rankings disagree with each other for the same reason, and nearly every “best AI detector 2026” list you will find was written by a company that sells one.

Why two detectors disagree about the same text

Say you paste an essay into one of these tools and it comes back “83% AI,” and the same essay in another one comes back “human.” Nothing strange has happened, and neither tool has caught the other out. You have seen two guesses from two pieces of software trained on different piles of text, drawing their lines in different places, neither of them tested against a shared standard that would make them agree. A raw score only means something next to a count of how often that score has been wrong before, and most of the companies selling these tools have never published that count.

What calibration changes

Calibration means going back over the guesses a tool has made and counting how often each level of confidence turned out to be right. When we built Cobalynx we took a set of texts that we knew the origin of, some written by people and some by models, froze it so that no one could swap the hard cases out later, ran the tool over all of it and wrote down what happened at every level of confidence. So when it tells you “likely AI at 90%” today, you can look up how often it was wrong the last time it said that.

It changed how the tool behaves as well. It says inconclusive when there is not enough to go on, and we would be suspicious of any tool that never once says it does not know, since a tool like that is still making the same mistakes and not reporting them. How often we flagged a person, how often we missed a model and how often we said we could not tell, with the raw counts and the confidence intervals, are on the evidence page in full, including the ones that do not make us look good. How the scores are produced is on the methodology page.

What you can expect, situation by situation

  • Long text: the longer the piece of writing you give it, the more reliable the answer gets. The band where our accuracy came out best is 1,000+-word text, though that is a detection rate rather than a safety rating, and the false-positive rate in that band is the worst we measured, so it is the last place to trust a single verdict on its own. A 500-word essay sits in the band that runs from 300 to 1,000 words, which still does a good deal better than a coin toss but carries more uncertainty, and you should treat it that way. The numbers for each band are on the methodology page.
  • Short text (under about 150 words): no detector on the market is any good on this, and we would not trust ours either. If what you have is a chat reply or one review or a short email, the honest thing for a tool to say is that it is too short to judge. We would rather say that than hand you a verdict we could not stand behind, which is why the tool has a minimum length.
  • Paraphrased or “humanized” text: this is what happens when someone takes writing that came out of a model and runs it through a rewriting tool so that it reads differently. Every tool on the market then gets worse at spotting it, ours included. We did this to our own test set to see how much worse we got, and what we found is on the evidence page.
  • Polished, formal, or non-native writing: this is where the documented false positives live, and where treating a score as proof does the most harm. If it has happened to you, we have written up what we would do.

How we would judge any detector, including ours

If we were choosing one of these tools for our own school or our own company, there are five things we would ask the people selling it before we paid them anything. Do they publish how often they were wrong, with the raw counts and the confidence intervals? Where did they draw their line? Will they show us the numbers that make them look bad, meaning how often they flagged people who wrote in a second language and how often they missed text that had been rewritten, right next to the numbers that make them look good? Do they hold a test set they do not train on, and publish what it says even when it reads badly for them? And does the tool ever say that it does not know? A company that can answer all five can still be wrong, but at least you will know how often. Our own answers are on the evidence page.

What Cobalynx can and can't do here

We can tell you how often texts that scored like yours turned out to be AI when we ran them over the set we test against, and for every level of confidence we give you, how often we were wrong at that level. When the text is too short, or there is not enough to go on, we would rather say we do not know than guess. What that score means for you depends on how much AI text was in your situation to begin with, and the base-rate explainer does that arithmetic for you, while the why-cobalynx page puts the whole checklist in one table. We cannot tell you that a person wrote something, since no statistical method can, and finding no AI signal in a text is no evidence that a human wrote it. A score, ours or anybody else’s, is one input to a judgment a person has to make, and never the verdict on its own. For a particular decision about an essay, an email or a chat transcript, the FAQ goes through the common cases, and the tool itself is free and does not ask you to sign up.

Sources

  1. Weber-Wulff et al., “Testing of detection tools for AI-generated text,” International Journal for Educational Integrity (2023), arxiv.org/abs/2306.15666 — none of 14 tested detectors reached 80% accuracy.
  2. Liang et al., “GPT detectors are biased against non-native English writers,” Patterns (2023), arxiv.org/abs/2304.02819.
  3. Detector false positives on the U.S. Constitution and other historical texts; overview in Wikipedia, “Artificial intelligence content detection” (accessed Aug 2026).
  4. Communications of the ACM, “Can AI Detectors Be Trusted?” (accessed Aug 2026) — survey of the reliability debate.

Common questions

Why does every detector claim 98% or 99% accuracy?

Because the way they usually say it, you have no way of checking it. If they do not tell you which texts they used for the test, or where they drew their line, or what the raw counts were, then almost any tool can come up with a headline number like that on a sample that happens to suit it. When a group of researchers tested fourteen of the popular tools side by side in 2023, not one of them got to 80% accuracy once the conditions changed a little, and those were the same tools whose home pages said 98% and 99%. So we would ask the company for its counts and its confidence intervals, and we have put ours on the evidence page.

Did AI detectors really flag the U.S. Constitution?

Yes, they did, and people have reproduced that result many times over, along with the same kind of flag on other famous old texts. Very formal, conventional prose reads a lot like what these tools learned to call machine written, and that is the same mistake that hurts real people who write formally or who write English as a second language.

Two detectors gave my text opposite verdicts. Which is right?

It could be that neither of them is. Tools that were never calibrated, and that were trained on different data with their lines drawn in different places, disagree with each other all the time, and what tends to happen if you paste the same paragraph into three of them in a row is that you get three different answers. The question we would ask is how often each of those tools has been wrong when it gave you that level of confidence, and if one of them has never published that number at all, then it is asking you for trust it has not earned.

Do “humanizer” tools prove detection is useless?

No, they do not, but they do real damage, since a paraphrase attack makes every detector worse, and we have measured how much worse it makes ours. What we do about it is publish those numbers, and when the evidence is thin we say that we do not know, so that you can at least tell which case you are in.