The direct answer

We froze an evaluation set on August 28, 2026, and there were 1522 documents in it, and then we ran Cobalynx detection engine v2 over the whole thing, which is the same engine that is serving every scan you run today, so there is no gap between what we tested and what you get. On that set, at the operating point we publish, the false-positive rate on human writing came out at 2.1% (14 of 652; 95% CI 1.3%–3.6%). How well it caught AI text depended a great deal on how long the text was, so the catch rate is split by length band, and on the frozen set it came out like this: at 300–1,000 words we detect 98.0% of clean AI documents (492 of 502 non-abstained) at a 2.0% human false-positive rate (11 of 551); at 1,000+ words, 100.0% (18 of 18) with 2 false positives in 19 human documents; below 300 words the clean-AI evaluation slice is empty, so these results do not establish accuracy for that length.

That is a longer answer than the “99% accurate” you would get from a competitor. How accurate a detector is depends on how long the text was, what language it was written in and how much it was edited, so the answer to “how accurate” has to be a table, and the short-text and paraphrase cases further down are the ones that matter most if you are the person being scored.

How often each verdict is wrong

If you are reading a result, the accuracy question that matters to you is narrower than the overall rate, and it is how often the verdict you have just been shown turns out to be wrong. This table answers that one, from the same frozen run as the numbers above.

What happenedCountsRate95% interval
Said “likely AI”, text was actually human 14 of 670 2.1% 1.2% – 3.5%
Said “likely human”, text was actually AI 126 of 764 16.5% 14.0% – 19.3%
Answered “inconclusive” instead of guessing 88 of 1522 5.8% 4.7% – 7.1%

The “inconclusive” row is there on purpose, and we would not read it as damage. When a piece of text really is borderline we abstain and do not guess, since a confident answer that turns out to be wrong is the one outcome this whole product was built to avoid.

The numbers a marketing page would leave out

Start with what happens when somebody takes AI text and rewords it deliberately to get past a detector. Every detector gets worse at catching that, ours included, and our recall under that kind of attack came out at 55.0% of non-abstained adversarial documents (142 of 258; 45.8% counting abstentions as misses). If you want to see how the attack works there is an entry for it in the glossary.

Then there is short text. Below our word floor we give you no answer at all, so there is no accuracy figure here for a case we declined to score, and if another tool quotes you an accuracy figure for 100 words of text then it is putting a number on a guess.

And there is writing by people whose first language is not English. Detectors as a group flag that kind of writing as AI far more often than they should, and that has been known since 2023, so we score that slice on its own and hold it to the same gate as everything else, and our false-positive rate on it came out at 0.6% (1 of 159; 95% CI 0.1%–3.5%).

How these numbers were produced

We ran one evaluation on a set of documents we had frozen and set aside beforehand, and every figure on this page is read straight out of that run rather than typed in by anyone. The full tables are on the evidence page and the pipeline is written up on the methodology page. The measured ranking quality is abstention-free AUROC 0.993 on the clean evaluation split, 0.994 when non-native human writing is included. AUROC describes how well scores separate the evaluated groups; it does not establish that an individual percentage is a reliable probability for a new population. See what calibration means and the evaluation limits before applying a result.

How to compare us with anyone else

If you want to compare us with another vendor, we would ask every one of them the same three questions: what decision threshold was the accuracy tested at, what test set was it tested on, and what were the raw counts and the intervals? In 2023, Weber-Wulff and her colleagues tested 14 of the popular detectors under a range of conditions and not one of them got to 80% accuracy [Weber-Wulff et al., 2023], and at the same time the vendors’ pages were saying 98 to 99%. Both of those cannot be true at once, and when you look into a gap like that the difference usually comes down to the conditions the test was run under. Ours are published, and every vendor should be held to the same, us included. The longer version of that argument is on Can AI detectors be trusted?

Sources

  1. Weber-Wulff et al., “Testing of detection tools for AI-generated text,” International Journal for Educational Integrity (2023) — all 14 tested detectors below 80% accuracy, arxiv.org/abs/2306.15666.