How the detector works

The software reads the writing itself, so the way the words and the sentences were picked and put together, since that turns out to be different, in ways you can count, when a person wrote it and when a model wrote it. When a person writes there are odd choices in it, personal ones, slightly “wrong” ones, the kind of thing an editor would have taken out, and when a machine writes it the result tends to come out more even. What it finds becomes a number, and that number was checked beforehand against a set of documents where we already knew who had written what.

It reads the whole document rather than the opening of it, so a few paragraphs of your own writing on top of generated text will not hide what is underneath. None of what it counts needs to know which AI wrote the text, so a big model and a small one go through the same reading, and it does not need to recognise the model that wrote your text, so a new one does not lock it out. What a new model family does to the accuracy is a different question, and the failure modes below are honest about it.

There are two things we do not do. We cannot read the watermarks that Google and Anthropic put into their text, since those watermarks are keyed and only the company that holds the key can read them (there is more on that in our watermark explainer). And we do not ask a chatbot “does this look like AI to you?”, since when people have tested that properly it came out close to a coin flip.

File provenance checks

Besides the text scan there is the file check, and that one looks at a file you upload for signs of where it came from. Every reading you get back is in one of three states, which are found, definitely not found, and could not be checked, and we never fold the three of them into one verdict, since folding them is what turns “we could not check” into “we checked and it was fine”, and those are not the same answer:

  • Images (PNG, JPEG). We check the C2PA Content Credentials with cryptography, and there are two different good answers rather than one. You get Content Credentials verified when the signature and the content hashes are intact and the certificate behind that signature traces back to an authority we recognise, and that is the strongest reading anything on this site produces. You get signed, but the signer is unverified when the signature and the hashes are intact and the certificate traces back to nobody we know, which happens because a person can make their own certificate and sign a file with it in a few minutes, so the file has not been changed since signing and we still cannot tell you who did the signing. In both cases you will see who the manifest says signed it and any note about AI use that was recorded, and in the second case that note is a claim somebody attached to the file rather than something we checked out. You get failed verification when a manifest is there but it is broken, which usually means the file was changed after it was signed. And you get no credentials when there is nothing there to look at, which is the normal case and is not evidence that a person made the image, since most files have no credentials in them at all, however they were made.
  • Documents (DOCX, PDF). Here all we have is metadata. We read the creator field, the last-modified-by field, the application and the producer fields, and we flag any names in there that we know belong to AI tools. That metadata was typed in by a piece of software and you could edit it yourself in a minute if you wanted to, so it is a weak hint about which program touched the file and it says nothing about how the words in it got written.
  • Pasted text. When you paste text in you have thrown away everything that lived at the file level, so all the text scan can do is look for invisible Unicode characters, which are the zero-width and control characters that some pipelines leave behind. If we find any we tell you, and we would read that as a weak hint that a machine was in the chain somewhere, since you could just as well have picked them up by copying from a web page, and they never go into the score.

The two kinds of evidence get different labels on the results page. If you upload a photo and its Content Credentials check out, what you see next to it is “cryptographic evidence — strong”. If you paste in an essay, what you see next to the score is “statistical analysis — pattern inference, indicative”. A signature either passes or it does not, and there is no argument about it, but a score from a text scan is always a guess, and it comes with an error rate that we had to work out and publish.

What the scores mean

The big number we show you is what we call a calibrated AI-likelihood score. It came out of one set of documents and it is tied to that set, so it is not a verdict on the text you pasted in. Say you paste in an essay and you get 90%. When we went back and looked at our frozen test set, out of all of the texts that had scored the way your essay did, about 9 in 10 of them had been written by AI, and 1 in 10 had been written by a person.

We built that test set so that roughly 55.1% of it was AI, and so 90% is not the probability in your newsroom or your classroom or your inbox, since the share of AI text where you are could be very different from what it was in our set, and that changes the answer more than most people would expect. The base-rate explainer below works out how to translate it into your own situation. The underlying ranking quality on the frozen test set is abstention-free AUROC 0.993 on the clean evaluation split, 0.994 when non-native human writing is included.

The same thing is true going the other way. When we say likely human we have not established who wrote it, and a share of the likely-human answers on our own test set were AI, since we went and counted them. We publish that number on the evidence page and on the landing page in the card called “what our verdicts mean”. There is no statistical method anywhere that can certify that a person wrote something, and if a product tells you that it can, then it is claiming more than it has.

Every result you get lands in one of three states.

  • Likely AI means the number went over our strict threshold. We put that threshold where we did since, when we ran it against our test set, it gave us a false-positive rate of 2.1% (14 of 652; 95% CI 1.3%–3.6%). We could have put it lower and caught more AI text, and we did not, since the price of that trade is paid by people like you getting accused of something on the strength of a guess.
  • Inconclusive means the number landed between our two thresholds, in the band where a verdict is not worth giving. That happens on short text, on text that was edited heavily, and on writing that sits in the middle, and in all of those cases we would rather tell you we do not know than pick a side and be wrong about a person. We give this answer on purpose and we give it often, and a detector that never says it does not know is lying somewhere.
  • Likely human means the evidence pointed away from AI, and the word to weigh there is likely, since we never report “human-verified,” and there is no technology today that can prove a person wrote something. Text that one of the big providers has watermarked can still read as “likely human” to every tool on the market, ours included, since none of us can see the watermark.

If a text is flagged here, how likely is it really AI?

What a “likely AI” verdict would mean in your world. We worked out our error rates on a frozen set of documents that we built to be roughly 55.1% AI text by construction (838 of 1,522 documents), and the place where you work is almost certainly not 55.1% AI. If you are in a newsroom where AI text hardly ever turns up, then even a low false-positive rate means that a good share of the “likely AI” answers you get are going to be wrong, and if you are looking after a spam queue where nearly everything is AI, then nearly all of them are going to be right. The slider below does that arithmetic for you with the numbers we published, and when you play with it you will see why a score from any detector, ours included, must never be the only evidence against a person.

%

1% rare (e.g. a trusted newsroom) · 5% · 10% occasional · 20% mixed · 50% half and half · 90% an AI-heavy queue. The labels are only there to give you a feel for the scale.

Uses our published measured rates: 98.1% detection on clean AI text (510/520) and a 2.1% false-positive rate on human text (14/652), and both of those only count the cases where we gave a verdict at all. You can see the full tables with their confidence intervals on the evidence page.

Two worked examples of what our published rates would have meant for every 10,000 texts that got a verdict, once when 1% of them were AI and once when 20% of them were
Share of AI text in your context Flagged “likely AI” per 10,000 verdicts Of which human (false accusations) Flagged verdicts that are really AI “Likely human” verdicts that are really AI
1% (a mostly-human context) 310.6 212.6 31.6% 0.02%
20% (a mixed context) 2133.3 171.8 91.9% 0.49%

We would read the first row again. When only 1% of the texts are AI, a “likely AI” verdict is right just 31.6% of the time. The remaining flags would be false accusations. A teacher holding one of those flags would still need drafts, notes or other evidence to decide what happened. These figures come from our published rates.

Three honest limits. The first is that these are our numbers on a frozen set of documents that we measured on August 28, 2026, and they are not a guarantee about your texts, or your writers, or the models that came out after that date. The second is that the rates above only apply when we gave a verdict at all, and on that set we abstained (“inconclusive”) on 5.8% of documents and did not guess on those. The third is that our non-native English sample had a false-positive rate of 0.6% (1 of 159). That smaller sample does not establish the error rate for an individual school or language group. A person still needs to look at the work and the evidence of how it was written.

Known failure modes

You cannot read a score properly without knowing where it goes wrong, and every tool of this kind goes wrong in the same few places.

  • Short text. This is the first place we go wrong. The evidence builds up as the text gets longer, and when you go below roughly 150 to 200 words every detector we know of gets a lot worse, and ours is in that group. So we refuse something that short, or we warn you loudly on the results page, and we do not guess at it. Our accuracy by length band is at 300–1,000 words we detect 98.0% of clean AI documents (492 of 502 non-abstained) at a 2.0% human false-positive rate (11 of 551); at 1,000+ words, 100.0% (18 of 18) with 2 false positives in 19 human documents; below 300 words the clean-AI evaluation slice is empty, so these results do not establish accuracy for that length.
  • Heavy paraphrasing and “humanizers.” If a writer rewrites AI text on purpose to get it past a detector, that works, and it has been shown to work. In 2024 the RAID benchmark (ACL 2024) found that fairly simple attacks got past nearly every published detector, and paraphrasing that is guided by a detector is stronger again. Under our own adversarial split our recall was 55.0% of non-abstained adversarial documents (142 of 258; 45.8% counting abstentions as misses). So if the text in front of you is something a writer had a reason to rewrite, a “likely human” result on it is weak evidence, and you should weigh it that way.
  • Unseen registers and new models. A detector only knows the models it was trained on and the kinds of writing it was shown. If you gave us a kind of writing that was thin in what we trained on, and legal boilerplate, poetry and very technical prose would all be examples of that, or if a brand-new model family came out next month that no one had seen, then we could get worse even though nothing on our side had changed. So every engine change gets a version number (the one you are using now is Cobalynx detection engine v2) and has to be run against the frozen test set again before it goes out, and if a new model era made us worse you would see it in the published numbers.
  • Non-native English speakers. This is the most serious bias documented in the field so far. In 2023 a study in Patterns found that commercial detectors called more than half of real TOEFL essays AI, and the reason was that the simpler vocabulary of a writer working in their second language reads as “predictable” to the software. So we hold our thresholds to a false-positive gate that has to pass on writing by people whose first language is not English before those thresholds can ship, and we publish the rate we got on that writing, which was 0.6% (1 of 159; 95% CI 0.1%–3.5%). The bias cannot be removed completely, and that is one more reason a score should never be treated as proof.

Why we publish our numbers

Most products in this space give you one accuracy figure, usually something like “99%!”, and they do not tell you what attacks it was tested against, or what the false-positive rate was, or where they had set the threshold when they got it. When people with no stake in the answer have gone and tested those same products, what they found was routinely far lower, and in July 2026 the EU's own Code of Practice on AI-content transparency concluded that forensic AI-text detection is “not yet considered reliable enough.” The right thing to do with a finding like that is to build around it, so we publish the evaluation and we publish the error rates, and you can see for yourself what “likely” means when we say it.

Our current numbers are an AUROC of 0.993, a false-positive rate of 2.1% (14 of 652; 95% CI 1.3%–3.6%), and a false-negative rate of 1.9% (10 of 520; 95% CI 1.0%–3.5%) at the default operating point, all of them read out of one run on August 28, 2026 against a set of documents whose contents were written down before the run started. If we change the model, or we move the thresholds, or we change the documents, the numbers on this page change with them, and the updated date at the top will say when. The full tables are on the evidence page.

Never sole evidence

A Cobalynx score should never be the only evidence used to punish a person, and that includes failing a student or firing a writer or rejecting a submission. The field as a whole has come to the same conclusion. In July 2023 OpenAI shut down its own AI-text classifier “due to its low rate of accuracy.” In August 2023 Vanderbilt turned off Turnitin's AI detection after false positives, and dozens of other universities did the same after it, and a lot of those false positives had landed on students who were writing in their second language. At any realistic false-positive rate, if you screen thousands of honest documents you will flag innocent people, and no choice of threshold gets you out of that.

So use a score the way it was built to be used, which is as one thing you weigh next to the drafts, the version history and a conversation with the person who wrote the piece. If all you have is the score then you do not have enough yet. It is the stated policy of this site that Cobalynx output must never be used as the sole grounds for an academic or employment sanction.

Intended purpose and limitations

What this system is for. Cobalynx is an informational tool. Give it an essay and you get a score that says how likely the essay is to be AI, worked out on one particular set of documents and not on yours, so it tells you something about the essay without settling the question. Give it a photo or a Word file and you get a reading of where the file came from, which either checks out or it does not. Both of those were built to be one input to a human judgment that you or a colleague is going to make afterwards. Verdicts are produced by an automated statistical system, and every verdict comes with the error rates we got for it and with a way for the system to tell you that it does not know.

What it is not for. Cobalynx is not designed or offered as a system for evaluating people. Say you were a teacher thinking of using it to give a student a grade, or a manager who wanted to sort the people who had applied for a job, or you had a matter on your desk that could get a person disciplined or fired or taken to court and you wanted a number to point to. In all of those cases the answer is to stop. A score can go into a fair process where a person looks at the other evidence as well and makes up their own mind, but using a verdict on its own as the whole basis for a decision or a sanction against a person is prohibited by our terms of service as a condition of use.

Who builds this

Cobalynx is built and maintained by Roviant S.r.l., which is a limited liability company in Italy, and the company that publishes these numbers is also the one that runs the service, so there is no one in between you and the people who produced them. The people who work on it have been building with AI since 2020, and they have more than 15 years of software engineering behind them. If you write to contact@cobalynx.com that is who reads it.

When a number turns out to be wrong we correct it here and say what the old claim was, and if you were the one who found it we say that too.

Sources

  1. Dugan et al., “RAID,” ACL 2024, arxiv.org/abs/2405.07940 (adversarial-attack findings)
  2. Liang et al., “GPT detectors are biased against non-native English writers,” Patterns, 2023, arxiv.org/abs/2304.02819
  3. Krishna et al., “Paraphrasing evades detectors of AI-generated text,” NeurIPS 2023, arxiv.org/abs/2303.13408; adversarial paraphrasing, arxiv.org/abs/2506.07001
  4. OpenAI, classifier retirement notice (Jul 20, 2023), openai.com/index/new-ai-classifier-for-indicating-ai-written-text/
  5. Vanderbilt University, “Why we're disabling Turnitin's AI detector” (Aug 16, 2023)
  6. TechPolicy.Press, “The EU's AI Transparency Code of Practice, Explained”: the Code's reliability finding (accessed Aug 16, 2026)
  7. Jabarian & Imas, NBER Working Paper 34223 (2025): independent detector audit and cost-based policy framework