The frozen evaluation

Every number on this page came out of one run of our test. On August 28, 2026 we took Cobalynx detection engine v2 and ran it against a set of documents that we had frozen and held back for that purpose. That is the same engine that answers you when you press the button on the landing page, and every time it changes we run the test again before the change goes out, so there is no separate benchmark build and no gap between what we tested and what you get. This is a Cobalynx internal frozen-set evaluation, run by us, and the whole page should be read knowing that.

There are 1522 documents in the evaluation set. 520 of them are clean human documents, 528 are clean AI documents, 160 are human documents written by non-native (ESL) writers, 310 are AI documents that were put through a paraphrase attack, and 4 are further AI documents that belong to no named slice, which are scored in the run but left out of the per-slice rates. We mixed clean writing, non-native writing and attacked writing on purpose, since if you only ever test a detector on easy text you will get numbers that flatter it.

The human writing in the set was collected before large language models existed, so none of it could have come from one. The AI writing spans both older model generations and current ones, and a part of it was deliberately rewritten to get past a detector. None of the writing we train on is ever allowed into the frozen evaluation set.

The verdicts are read at one fixed operating point. We call a text “likely AI” when its calibrated probability came out at 0.96 or higher, we call it “likely human” when the probability came out below 0.50, and anything that fell in between we report as inconclusive. Every rate below comes with a Wilson 95% confidence interval, which is the range the true rate could sit in given how many documents we had to work with, and the range is the part to read rather than the single number in the middle of it.

Headline rates at the operating point

MeasureRateCountWilson 95% CI
False-positive rate — clean human writing (non-abstained) 2.6% 13 of 493 1.5% – 4.5%
False-positive rate — all human writing incl. non-native (non-abstained) 2.1% 14 of 652 1.3% – 3.6%
False-positive rate — non-native (ESL) human writing (non-abstained) 0.6% 1 of 159 0.1% – 3.5%
Detection rate (TPR) — clean AI text, all model eras (non-abstained) 98.1% 510 of 520 96.5% – 99.0%
Detection rate (TPR) — clean AI text, all model eras, abstentions counted as misses 96.6% 510 of 528 94.7% – 97.8%
Detection rate (TPR) — 2025–26-generation AI text (non-abstained) 100.0% 252 of 252 98.5% – 100.0%
Detection rate (TPR) — 2025–26-generation AI text, abstentions counted as misses 100.0% 252 of 252 98.5% – 100.0%
Adversarial recall — paraphrase-attacked AI text (non-abstained) 55.0% 142 of 258 48.9% – 61.0%
Adversarial recall — paraphrase-attacked AI text, abstentions counted as misses 45.8% 142 of 310 40.3% – 51.4%

When a row says “non-abstained” it was worked out over the documents where we gave a verdict, leaving out the ones where we said we did not know, and when a row says “abstentions counted as misses” every one of the times we said we did not know was charged against us as if we had got it wrong. We give you both since either one of them on its own could be made to look better than it really is, and if you ever read a rate from another company you should ask them which of the two they are showing you.

The rows that say “all model eras” cover the whole of the clean-AI slice, which mixes writing from older model generations with output from current ones. The rows that say “2025–26-generation” are the same measurement taken only on the current-generation samples. Put the two side by side and the AI people are using today comes out easier for us to catch than the mixed headline suggests, since it is the older material pulling the all-era raw rate down.

Before you read anybody’s accuracy claims, ours included, the EU has already made a finding about this whole field. In July 2026 the EU Code of Practice on AI-content transparency concluded that forensic detection of unwatermarked AI text is “not yet considered reliable enough,” and that finding came after people with no stake in the answer had tested the commercial detectors and got numbers a long way below what the marketing had told them. How we work inside that reality is written up on our methodology page.

The next table is the ranking quality before we apply the band where we say we do not know, which is what people in the field call the abstention-free AUROC, and it tells you how well the raw score would have sorted the AI documents from the human ones if you had lined them all up by score and had not drawn a line anywhere.

SliceAUROC
Clean evaluation split0.993
Clean split including non-native human writing0.994
Paraphrase-attacked AI vs clean human writing0.868

If a text is flagged here, how likely is it really AI?

What a “likely AI” verdict would mean in your world. We worked out our error rates on a frozen set of documents that we built to be roughly 55.1% AI text by construction (838 of 1,522 documents), and the place where you work is almost certainly not 55.1% AI. If you are in a newsroom where AI text hardly ever turns up, then even a low false-positive rate means that a good share of the “likely AI” answers you get are going to be wrong, and if you are looking after a spam queue where nearly everything is AI, then nearly all of them are going to be right. The slider below does that arithmetic for you with the numbers we published, and when you play with it you will see why a score from any detector, ours included, must never be the only evidence against a person.

%

1% rare (e.g. a trusted newsroom) · 5% · 10% occasional · 20% mixed · 50% half and half · 90% an AI-heavy queue. The labels are only there to give you a feel for the scale.

Uses our published measured rates: 98.1% detection on clean AI text (510/520) and a 2.1% false-positive rate on human text (14/652), and both of those only count the cases where we gave a verdict at all. You can see the full tables with their confidence intervals on the evidence page.

Two worked examples of what our published rates would have meant for every 10,000 texts that got a verdict, once when 1% of them were AI and once when 20% of them were
Share of AI text in your context Flagged “likely AI” per 10,000 verdicts Of which human (false accusations) Flagged verdicts that are really AI “Likely human” verdicts that are really AI
1% (a mostly-human context) 310.6 212.6 31.6% 0.02%
20% (a mixed context) 2133.3 171.8 91.9% 0.49%

We would read the first row again. When only 1% of the texts are AI, a “likely AI” verdict is right just 31.6% of the time. The remaining flags would be false accusations. A teacher holding one of those flags would still need drafts, notes or other evidence to decide what happened. These figures come from our published rates.

Three honest limits. The first is that these are our numbers on a frozen set of documents that we measured on August 28, 2026, and they are not a guarantee about your texts, or your writers, or the models that came out after that date. The second is that the rates above only apply when we gave a verdict at all, and on that set we abstained (“inconclusive”) on 5.8% of documents and did not guess on those. The third is that our non-native English sample had a false-positive rate of 0.6% (1 of 159). That smaller sample does not establish the error rate for an individual school or language group. A person still needs to look at the work and the evidence of how it was written.

What the verdict labels mean

These are the same three numbers that you see on the landing page, counted across the whole set of documents at once, with the clean writing, the non-native writing and the attacked writing all taken together, since a rate counted over the easy slices only would not tell you anything.

OutcomeRateCountWilson 95% CI
Said “likely AI”, text was actually human 2.1% 14 of 670 1.2% – 3.5%
Said “likely human”, text was actually AI 16.5% 126 of 764 14.0% – 19.3%
Answered “inconclusive” instead of guessing 5.8% 88 of 1522 4.7% – 7.1%

Abstention rates per split

This table shows how often we said “inconclusive” for each slice of the test set, when the other option would have been a confident answer that was wrong. We said it more on the hard slices, and that is what we want it to do.

SliceAbstention rateCountWilson 95% CI
Clean evaluation split (human + AI) 3.3% 35 of 1048 2.4% – 4.6%
Clean AI documents 1.5% 8 of 528 0.8% – 3.0%
Clean human documents 5.2% 27 of 520 3.6% – 7.4%
Non-native (ESL) human documents 0.6% 1 of 160 0.1% – 3.5%
Paraphrase-attacked AI documents 16.8% 52 of 310 13.0% – 21.3%

What these numbers do not cover

Choose the evidence that fits the writing you are checking
Your use caseHow to interpret the published results
English essays and longer proseThe tables describe a mixed writing population. They are not an essay-only accuracy guarantee. Keep drafts and version history alongside any score.
English written by non-native speakersThe evaluation includes 160 non-native human documents. Read that slice's rate and confidence interval, not just the overall rate.
Rewritten AI textThe evaluation includes 310 paraphrase-attacked AI documents. Read both detection and inconclusive rates for this harder population.
Emails, cover letters and business writingWe have not published separate false-positive measurements for these writing types. The overall rate does not establish their accuracy.
Short text or other languagesText under 150 words is refused. Languages other than English are outside the measured population.

Every rate above was measured on the population described at the top of this page, and a rate is only as good as the population it came from, so if the writing you are checking looks nothing like what is in that set, these figures are the wrong yardstick for it. Text under our minimum length is not scored at all. Writing in languages other than English is outside what we measured. Text that has been through a rewriting tool is represented here, and the table above shows you how much harder it is for us. And a population we have not measured has no number on this page at all, since a number we did not produce would not be evidence of anything.

There is also a limit on what any of this can be used for. A score is one input to a judgment a person makes, and it should never be the only evidence used to punish somebody. What a verdict can and cannot support, and how to contest one, is set out on our methodology page.

Inspect and try the public examples

Read the exact text before running it through the checker. These are the same demonstration samples offered on the homepage. They show how the tool behaves on individual examples; they are not a representative benchmark and do not establish an accuracy rate. The measured results above come from the larger frozen evaluation.

AI-generated

Generated by a large language model in August 2026, unedited.

Urban trees provide a remarkable range of benefits that extend far beyond their aesthetic appeal. They improve air quality by filtering particulate matter and absorbing pollutants, while simultaneously reducing urban heat island effects through shade and evapotranspiration. Studies have shown that neighborhoods with substantial tree cover experience measurably lower summer temperatures, which translates into reduced energy consumption for cooling and improved comfort for residents. Beyond their environmental contributions, urban trees offer significant social and economic advantages. Research consistently demonstrates that tree-lined streets are associated with higher property values, increased foot traffic for local businesses, and stronger community engagement. Access to green spaces has been linked to improved mental health outcomes, lower stress levels, and enhanced cognitive function in both children and adults. However, maintaining a healthy urban forest requires deliberate planning and sustained investment. Municipalities must consider species diversity to guard against pests and disease, ensure adequate soil volume and water access, and plan for long-term maintenance costs. When these factors are addressed thoughtfully, the return on investment is substantial: urban trees represent one of the most cost-effective interventions available for improving quality of life in cities. As climate pressures intensify, their role in urban resilience will only become more essential.

Load this example in the checker. Review it, then press Scan text to run a fresh check.

Human-written (1880)

Mark Twain, “The Awful German Language” (A Tramp Abroad, 1880) — public domain, written more than a century before AI text generation.

I went often to look at the collection of curiosities in Heidelberg Castle, and one day I surprised the keeper of it with my German. I spoke entirely in that language. He was greatly interested; and after I had talked a while he said my German was very rare, possibly a “unique”; and wanted to add it to his museum. If he had known what it had cost me to acquire my art, he would also have known that it would break any collector to buy it. Harris and I had been hard at work on our German during several weeks at that time, and although we had made good progress, it had been accomplished under great difficulty and annoyance, for three of our teachers had died in the mean time. A person who has not studied German can form no idea of what a perplexing language it is. Surely there is not another language that is so slipshod and systemless, and so slippery and elusive to the grasp. One is washed about in it, hither and thither, in the most helpless way; and when at last he thinks he has captured a rule which offers firm ground to take a rest on amid the general rage and turmoil of the ten parts of speech, he turns over the page and reads, “Let the pupil make careful note of the following exceptions.” He runs his eye down and finds that there are more exceptions to the rule than instances of it.

Load this example in the checker. Review it, then press Scan text to run a fresh check.

AI, then paraphrased

The same AI draft rewritten by a second model pass — the paraphrase pattern used to evade detectors. Committed August 2026.

City trees do a lot more for us than just look nice. They clean the air we breathe, catching dust and soaking up pollution, and they cool overheated streets with their shade and the moisture they release. Neighborhoods with plenty of canopy stay noticeably cooler through the summer months, so people run their air conditioning less and simply feel better outdoors. The advantages don't stop with the environment, either. Streets lined with trees tend to see homes sell for more, shops attract more passersby, and neighbors get to know one another. Spending time near greenery seems to ease stress and lift mood, and there's evidence it helps kids and adults alike think more clearly. None of this happens on its own, though. A city that wants a thriving canopy has to plant a mix of species so one pest can't wipe everything out, give roots enough soil and water to survive, and budget for pruning and care decades into the future. Get those pieces right and the payoff is hard to beat — few public investments deliver as much everyday benefit as trees do. And as the climate grows harsher, cities will lean on them more, not less.

Load this example in the checker. Review it, then press Scan text to run a fresh check.