The problem we were trying to fix
Before we built this we read what the companies in this market said about themselves, and then what happened when somebody outside the company tested their software, and the two did not match. A lot of them had put “99% accuracy” on the front page, and when the outside tests came in the number was a long way below that, and some of them had never published an error rate at all, so there was nothing to compare against in the first place. A few would sell you, on the same website, the paraphrasing tool that got past their own scanner. And there were real people on the other end of all this, students and job applicants and writers, who were being punished since one of these things had handed out a percentage and nobody had asked what the percentage was supposed to mean.
So we made the opposite bet, which was that the most useful scanner you could have is the one that tells you how often it gets things wrong. We ran the current engine version once, on a set of documents we had frozen beforehand so that it could not be gone back over and improved afterwards, and we published every number that came out of that run, including the ones that made us look bad.
Designed to minimize false accusations, not maximize accusations.
Twelve questions to ask any detector
The middle column says “a typical detector”, and that is not any one company. In August 2026 we worked our way down the tools that people had heard of, one by one, and wrote down what each of them said about itself and what other people had found when they put it to the test, and by the end it was the same story nearly every time. Where the table says what we did there is a link, so you can go and see for yourself.
| The question to ask | A typical detector | Cobalynx | Check it |
|---|---|---|---|
| What accuracy do you claim? | They give you one big number, and usually it was “99%” | We give no single accuracy number, since no one number could be true across every case. You get an error rate for each situation we tested, and a 95% confidence interval around each one so you can see how sure we were | /evidence |
| How often do you flag writing that a person wrote? | You would have had to dig for it, if they had put it anywhere at all | We counted it on our frozen set: 2.1% of human texts that received a verdict (14 of 652; CI 1.3%–3.6%); 2.6% on native-speaker writing (13 of 493) | /evidence |
| What about people who learned English as a second language? | The bias has been written about all over the industry, but we could not find a company that had run the test on its own tool and published the number | We ran the test and put the number on the page: 0.6% (1 of 159). On this run it came out lower than the rate for writers working in their first language, which is not a guarantee for any individual piece of writing, and it is the number to look at before you rely on a score for somebody who learned English later on | /evidence |
| What happens when the text is borderline? | You got a verdict anyway | You get “inconclusive”, which is what we said on 5.8% of our evaluation set (88 of 1522). When we were not sure we said so, since a guess would not have helped you | /methodology |
| What does the score mean? | They give you a percentage, and if you asked them what it was a percentage of you would not get much of an answer | Ours was worked out on a set of documents we had already frozen, so it says something like this: of the texts in that set that scored the way yours did, about that many turned out to have come from AI. We have done the base-rate arithmetic for you as well, since that is the part people get wrong most often | /methodology |
| Does “likely human” prove anything? | They let you assume that it did | No, it does not, and we can put a number on how far it falls short. On our evaluation set, 16.5% of texts we called likely human were actually AI (126 of 764) | /evidence |
| What about AI text that has been paraphrased (“humanized”)? | Nobody said a word about it | We ran that case on purpose and published what came back: 55.0% recall on paraphrase-attacked text (142 of 258 with a verdict). Every detector is weaker here, and we would rather you had the figure in front of you than had to go looking for it | /evidence |
| What do you do with a text that is too short to judge? | They scored it anyway | We refuse anything under 150 words. Short texts were where every tool we tried did its worst work, and a number we would not have trusted ourselves is not worth handing you | /methodology |
| Do you look at anything other than the writing style? | They looked at the style of the writing and nothing about the file it came in | We also read what the file itself carries, which is the Content Credentials (C2PA) and the document metadata, and that is evidence that came with the file, on top of what we can tell from the prose | /check-image-content-credentials |
| Do you sell a bypass or paraphrasing tool next door? | Some of them did, yes | No, and we never have. There is no humanizer or paraphraser on this site, and nothing that would help you get past a scanner, and we do not intend to add one | this page |
| Do you keep my text? | Often you could not tell, and you had to sign up before you found out | There is no signup and we keep nothing. The text you paste in gets processed in memory and then it is gone, and you can read how we prove that on the privacy page | /privacy |
| Where did you draw the line, and what did it cost? | They did not say, and often you could not find the threshold at all | We publish both thresholds and the false-positive rate we accepted when we set them, so you can see what we traded away to keep confident wrong accusations rare | /evidence |
What we ask of you
You do not have to take this page on trust. Every number in that table came out of the run written up on our evidence page, which tells you what the set we froze was made of (1,522 documents, as of August 28, 2026) and what thresholds we ran at. And whatever tool you end up using, ours included, never let a score be the only reason you punish a person, since when that rule gets broken it is usually somebody who did nothing wrong who ends up paying for it. If you have been flagged yourself, that page is the place to start. If you want to know how the score gets put together from beginning to end, that is on the methodology page.
Who builds this
Cobalynx is built and looked after by Roviant S.r.l., which is a limited liability company in Italy, and the same company writes the software, runs the servers, publishes the numbers and answers the mail, so there is no one standing in between. The people who work on it have been building with AI since 2020, and they have more than 15 years of software engineering behind them.
So one named company answers for every number on this site. When something we published turns out to be wrong, the correction goes up here under the same name that made the claim in the first place, and it says what the old claim was.