The frozen evaluation
Every number on this page came out of one run of our test. On August 28, 2026 we took Cobalynx detection engine v2 and ran it against a set of documents that we had frozen and held back for that purpose. That is the same engine that answers you when you press the button on the landing page, and every time it changes we run the test again before the change goes out, so there is no separate benchmark build and no gap between what we tested and what you get. This is a Cobalynx internal frozen-set evaluation, run by us, and the whole page should be read knowing that.
There are 1522 documents in the evaluation set. 520 of them are clean human documents, 528 are clean AI documents, 160 are human documents written by non-native (ESL) writers, 310 are AI documents that were put through a paraphrase attack, and 4 are further AI documents that belong to no named slice, which are scored in the run but left out of the per-slice rates. We mixed clean writing, non-native writing and attacked writing on purpose, since if you only ever test a detector on easy text you will get numbers that flatter it.
The human writing in the set was collected before large language models existed, so none of it could have come from one. The AI writing spans both older model generations and current ones, and a part of it was deliberately rewritten to get past a detector. None of the writing we train on is ever allowed into the frozen evaluation set.
The verdicts are read at one fixed operating point. We call a text “likely AI” when its calibrated probability came out at 0.96 or higher, we call it “likely human” when the probability came out below 0.50, and anything that fell in between we report as inconclusive. Every rate below comes with a Wilson 95% confidence interval, which is the range the true rate could sit in given how many documents we had to work with, and the range is the part to read rather than the single number in the middle of it.
Headline rates at the operating point
| Measure | Rate | Count | Wilson 95% CI |
|---|---|---|---|
| False-positive rate — clean human writing (non-abstained) | 2.6% | 13 of 493 | 1.5% – 4.5% |
| False-positive rate — all human writing incl. non-native (non-abstained) | 2.1% | 14 of 652 | 1.3% – 3.6% |
| False-positive rate — non-native (ESL) human writing (non-abstained) | 0.6% | 1 of 159 | 0.1% – 3.5% |
| Detection rate (TPR) — clean AI text, all model eras (non-abstained) | 98.1% | 510 of 520 | 96.5% – 99.0% |
| Detection rate (TPR) — clean AI text, all model eras, abstentions counted as misses | 96.6% | 510 of 528 | 94.7% – 97.8% |
| Detection rate (TPR) — 2025–26-generation AI text (non-abstained) | 100.0% | 252 of 252 | 98.5% – 100.0% |
| Detection rate (TPR) — 2025–26-generation AI text, abstentions counted as misses | 100.0% | 252 of 252 | 98.5% – 100.0% |
| Adversarial recall — paraphrase-attacked AI text (non-abstained) | 55.0% | 142 of 258 | 48.9% – 61.0% |
| Adversarial recall — paraphrase-attacked AI text, abstentions counted as misses | 45.8% | 142 of 310 | 40.3% – 51.4% |
When a row says “non-abstained” it was worked out over the documents where we gave a verdict, leaving out the ones where we said we did not know, and when a row says “abstentions counted as misses” every one of the times we said we did not know was charged against us as if we had got it wrong. We give you both since either one of them on its own could be made to look better than it really is, and if you ever read a rate from another company you should ask them which of the two they are showing you.
The rows that say “all model eras” cover the whole of the clean-AI slice, which mixes writing from older model generations with output from current ones. The rows that say “2025–26-generation” are the same measurement taken only on the current-generation samples. Put the two side by side and the AI people are using today comes out easier for us to catch than the mixed headline suggests, since it is the older material pulling the all-era raw rate down.
Before you read anybody’s accuracy claims, ours included, the EU has already made a finding about this whole field. In July 2026 the EU Code of Practice on AI-content transparency concluded that forensic detection of unwatermarked AI text is “not yet considered reliable enough,” and that finding came after people with no stake in the answer had tested the commercial detectors and got numbers a long way below what the marketing had told them. How we work inside that reality is written up on our methodology page.
The next table is the ranking quality before we apply the band where we say we do not know, which is what people in the field call the abstention-free AUROC, and it tells you how well the raw score would have sorted the AI documents from the human ones if you had lined them all up by score and had not drawn a line anywhere.
| Slice | AUROC |
|---|---|
| Clean evaluation split | 0.993 |
| Clean split including non-native human writing | 0.994 |
| Paraphrase-attacked AI vs clean human writing | 0.868 |
If a text is flagged here, how likely is it really AI?
What a “likely AI” verdict would mean in your world. We worked out our error rates on a frozen set of documents that we built to be roughly 55.1% AI text by construction (838 of 1,522 documents), and the place where you work is almost certainly not 55.1% AI. If you are in a newsroom where AI text hardly ever turns up, then even a low false-positive rate means that a good share of the “likely AI” answers you get are going to be wrong, and if you are looking after a spam queue where nearly everything is AI, then nearly all of them are going to be right. The slider below does that arithmetic for you with the numbers we published, and when you play with it you will see why a score from any detector, ours included, must never be the only evidence against a person.
Uses our published measured rates: 98.1% detection on clean AI text (510/520) and a 2.1% false-positive rate on human text (14/652), and both of those only count the cases where we gave a verdict at all. You can see the full tables with their confidence intervals on the evidence page.
| Share of AI text in your context | Flagged “likely AI” per 10,000 verdicts | Of which human (false accusations) | Flagged verdicts that are really AI | “Likely human” verdicts that are really AI |
|---|---|---|---|---|
| 1% (a mostly-human context) | 310.6 | 212.6 | 31.6% | 0.02% |
| 20% (a mixed context) | 2133.3 | 171.8 | 91.9% | 0.49% |
We would read the first row again. When only 1% of the texts are AI, a “likely AI” verdict is right just 31.6% of the time. The remaining flags would be false accusations. A teacher holding one of those flags would still need drafts, notes or other evidence to decide what happened. These figures come from our published rates.
Three honest limits. The first is that these are our numbers on a frozen set of documents that we measured on August 28, 2026, and they are not a guarantee about your texts, or your writers, or the models that came out after that date. The second is that the rates above only apply when we gave a verdict at all, and on that set we abstained (“inconclusive”) on 5.8% of documents and did not guess on those. The third is that our non-native English sample had a false-positive rate of 0.6% (1 of 159). That smaller sample does not establish the error rate for an individual school or language group. A person still needs to look at the work and the evidence of how it was written.
What the verdict labels mean
These are the same three numbers that you see on the landing page, counted across the whole set of documents at once, with the clean writing, the non-native writing and the attacked writing all taken together, since a rate counted over the easy slices only would not tell you anything.
| Outcome | Rate | Count | Wilson 95% CI |
|---|---|---|---|
| Said “likely AI”, text was actually human | 2.1% | 14 of 670 | 1.2% – 3.5% |
| Said “likely human”, text was actually AI | 16.5% | 126 of 764 | 14.0% – 19.3% |
| Answered “inconclusive” instead of guessing | 5.8% | 88 of 1522 | 4.7% – 7.1% |
Abstention rates per split
This table shows how often we said “inconclusive” for each slice of the test set, when the other option would have been a confident answer that was wrong. We said it more on the hard slices, and that is what we want it to do.
| Slice | Abstention rate | Count | Wilson 95% CI |
|---|---|---|---|
| Clean evaluation split (human + AI) | 3.3% | 35 of 1048 | 2.4% – 4.6% |
| Clean AI documents | 1.5% | 8 of 528 | 0.8% – 3.0% |
| Clean human documents | 5.2% | 27 of 520 | 3.6% – 7.4% |
| Non-native (ESL) human documents | 0.6% | 1 of 160 | 0.1% – 3.5% |
| Paraphrase-attacked AI documents | 16.8% | 52 of 310 | 13.0% – 21.3% |
What these numbers do not cover
| Your use case | How to interpret the published results |
|---|---|
| English essays and longer prose | The tables describe a mixed writing population. They are not an essay-only accuracy guarantee. Keep drafts and version history alongside any score. |
| English written by non-native speakers | The evaluation includes 160 non-native human documents. Read that slice's rate and confidence interval, not just the overall rate. |
| Rewritten AI text | The evaluation includes 310 paraphrase-attacked AI documents. Read both detection and inconclusive rates for this harder population. |
| Emails, cover letters and business writing | We have not published separate false-positive measurements for these writing types. The overall rate does not establish their accuracy. |
| Short text or other languages | Text under 150 words is refused. Languages other than English are outside the measured population. |
Every rate above was measured on the population described at the top of this page, and a rate is only as good as the population it came from, so if the writing you are checking looks nothing like what is in that set, these figures are the wrong yardstick for it. Text under our minimum length is not scored at all. Writing in languages other than English is outside what we measured. Text that has been through a rewriting tool is represented here, and the table above shows you how much harder it is for us. And a population we have not measured has no number on this page at all, since a number we did not produce would not be evidence of anything.
There is also a limit on what any of this can be used for. A score is one input to a judgment a person makes, and it should never be the only evidence used to punish somebody. What a verdict can and cannot support, and how to contest one, is set out on our methodology page.
Inspect and try the public examples
Read the exact text before running it through the checker. These are the same demonstration samples offered on the homepage. They show how the tool behaves on individual examples; they are not a representative benchmark and do not establish an accuracy rate. The measured results above come from the larger frozen evaluation.
AI-generated
Generated by a large language model in August 2026, unedited.
Load this example in the checker. Review it, then press Scan text to run a fresh check.
Human-written (1880)
Mark Twain, “The Awful German Language” (A Tramp Abroad, 1880) — public domain, written more than a century before AI text generation.
Load this example in the checker. Review it, then press Scan text to run a fresh check.
AI, then paraphrased
The same AI draft rewritten by a second model pass — the paraphrase pattern used to evade detectors. Committed August 2026.
Load this example in the checker. Review it, then press Scan text to run a fresh check.