Start with what the critics get right
AI detectors have flagged the U.S. Constitution, Macbeth, and the Book of Genesis as machine-generated [reported widely, 2023–2026]. A cross-tool academic study found none of 14 popular detectors reached 80% accuracy [Weber-Wulff et al., 2023]. Non-native English writing gets falsely flagged at rates above 60% by mainstream tools [Liang et al., 2023], and paraphrase tools measurably degrade every detector on the market. All of this is true, and any detector vendor who won't say it out loud is telling you something about their other claims too.
So why does every product page say 99%?
Because an accuracy number without its test conditions is unfalsifiable marketing. Accuracy depends on text length, language, the generating model, the decision threshold, and how much human editing happened — move any of those and the same tool can score brilliantly or barely above chance. A vendor who picks a favorable test set, a favorable threshold, and reports a single percentage has not lied, exactly; they have published a number that cannot be checked. That is why rankings conflict, and why nearly every “best AI detector 2026” list is written by a company that sells one.
Why two detectors disagree about the same text
Different training data, different thresholds, and no shared calibration standard. When one tool says “83% AI” and another says “human,” nothing paradoxical happened — you witnessed two uncalibrated probability estimates. A raw score is only meaningful next to a measurement of how often that score is wrong, and most tools never publish that measurement.
What calibration actually changes
Calibration is the difference between a confident guess and a quantified one. A calibrated detector reports a probability and the measured real-world error rate of that confidence band on a frozen public benchmark — so “likely AI at 90%” comes with a known answer to the question “and how often is that wrong?” It also changes behavior: a calibrated system says inconclusive when the evidence is thin, because a detector that never says “I don't know” is hiding its errors inside fake confidence. This is how Cobalynx is built, and the resulting numbers — false positives, false negatives, abstention rates, with raw counts and confidence intervals — are published in full, including the unflattering ones.
The honest picture, by situation
- Long-form text (500+ words): detection works meaningfully better than chance — this is the strongest measured territory, and where a calibrated verdict is genuinely informative.
- Short text (under ~150 words): nobody's detector is reliable. A chat reply, a single review, a brief email — the honest output is “too short,” not a verdict. We enforce a minimum length for exactly this reason.
- Paraphrased or “humanized” text: detection degrades, measurably, for everyone. We publish our paraphrase-attack numbers rather than claiming immunity.
- Polished, formal, or non-native writing: the documented false-positive zone. This is where treating any score as proof does real harm — see what to do if it happens to you.
How to evaluate any detector — including ours
Ask five questions. Does it publish error rates with raw counts and confidence intervals, not lone percentages? Does it state its decision thresholds? Does it report unflattering numbers — false positives on non-native writing, misses on paraphrased text — next to the flattering ones? Can an outsider reproduce the evaluation? And does it ever abstain? A tool that fails these questions is asking for trust it has not earned; a tool that passes them can still be wrong, but you will know exactly how often. Our answers, verifiable and reproducible, are on the evidence page — including the row where we fail our own checklist today (no third-party audit yet).
What Cobalynx can and can't do here
We can tell you the measured probability that a text is AI-generated, with a published error rate for every confidence band, and we can refuse to guess when the text is too short or the evidence too thin. We cannot certify any text as human — no statistical method can, and a missing AI signal is not evidence of human authorship. A score, ours or anyone's, is an input to judgment; it is never a verdict. If you're weighing a specific decision — an essay, an email, a chat transcript — the FAQ covers the common cases, and the checker itself is free, with no signup.
Sources
- Weber-Wulff et al., “Testing of detection tools for AI-generated text,” International Journal for Educational Integrity (2023), arxiv.org/abs/2306.15666 — none of 14 tested detectors reached 80% accuracy.
- Liang et al., “GPT detectors are biased against non-native English writers,” Patterns (2023), arxiv.org/abs/2304.02819.
- Detector false positives on the U.S. Constitution and other historical texts; overview in Wikipedia, “Artificial intelligence content detection” (accessed Aug 2026).
- Communications of the ACM, “Can AI Detectors Be Trusted?” (accessed Aug 2026) — survey of the reliability debate.
Common questions
Why does every detector claim 98% or 99% accuracy?
Because the claim is unfalsifiable as usually stated: without the test set, the decision threshold, and raw counts, a headline percentage can be produced by almost any tool on a favorable sample. Peer-reviewed cross-tool testing paints a different picture — none of 14 popular detectors reached 80% accuracy across conditions. Ask any vendor for counts and confidence intervals; ours are on the evidence page.
Did AI detectors really flag the U.S. Constitution?
Yes — that result has been reproduced widely, along with similar flags on other famous historical texts. Highly formal, conventional prose is statistically close to what detectors learn as AI-typical, which is exactly the false-positive mode that hurts real writers of formal or non-native English.
Two detectors gave my text opposite verdicts. Which is right?
Possibly neither — uncalibrated tools with different thresholds and training data routinely disagree. The meaningful question is not which label to believe but what error rate each tool has measured for the confidence it reported. A tool that publishes no error rate is asking for trust it has not earned.
Do “humanizer” tools prove detection is useless?
No, but they do real damage: paraphrase attacks measurably degrade every detector, ours included. We publish our adversarial numbers instead of claiming immunity, and we abstain rather than guess when the evidence is thin — an honest detector tells you which case you are in.