The direct answer

On our frozen evaluation of August 16, 2026 (1522 documents; engine v1 — the engine serving every scan today), the false-positive rate on human writing at the published operating point was 0.5% (3 of 609; 95% CI 0.2%–1.4%). Detection by length: at 300–1,000 words we detect 81.6% of clean AI documents (346 of 424 non-abstained) at a 0.6% human false-positive rate (3 of 525); at 1,000+ words, 100.0% (18 of 18) with 0 false positives in 20 human documents; below 300 words the clean-AI evaluation slice is empty by design — we abstain or refuse rather than guess.

If that answer looks longer than a competitor's “99% accurate”, that is the point: accuracy depends on text length, language, and editing, so the honest answer to “how accurate” is a table, not a number. The single-number version of this page would be marketing.

How often each verdict is wrong

The most useful accuracy question is conditional: given the verdict you just read, how often is it wrong? Measured on the same frozen run:

What happenedCountsRate95% interval
Said “likely AI”, text was actually human 3 of 442 0.7% 0.2% – 2.0%
Said “likely human”, text was actually AI 223 of 829 26.9% 24.0% – 30.0%
Answered “inconclusive” instead of guessing 251 of 1522 16.5% 14.7% – 18.4%

The “inconclusive” row is deliberate, not damage: on genuinely borderline text we abstain rather than guess, because a confident wrong answer is the one outcome this product exists to avoid.

The numbers a marketing page would hide

  • Paraphrase attack: AI text deliberately reworded to evade detection degrades every detector, ours included. Our measured adversarial recall: 33.5% of non-abstained adversarial documents (73 of 218; 23.5% counting abstentions as misses). We publish this instead of claiming immunity; the mechanics are in the glossary.
  • Short text: below the word floor we refuse to answer at all — there is no measured accuracy for a case we decline, and any tool quoting one is guessing.
  • Non-native (ESL) writing: detectors as a class false-flag non-native English at elevated rates, so we measure that slice separately and hold it to the same gate: our ESL false-positive rate is 1.7% (2 of 118; 95% CI 0.5%–6.0%).

How these numbers were produced

One sanctioned, seeded evaluation run on a frozen held-out corpus, with every dependent artifact checksummed and every execution of the measurement command disclosed in an append-only ledger. The full tables, the corpus composition, and the step-by-step reproduction pack are on the evidence page; the pipeline itself is documented on the methodology page. Calibration — what makes the probability we show mean what it says — is abstention-free AUROC 0.952 on the clean evaluation split, 0.938 when non-native human writing is included (what calibration means).

How to compare us with anyone else

Ask any detector vendor the same three questions: at what decision threshold was the accuracy measured, on what test set, and with what raw counts and intervals? Peer-reviewed cross-tool testing found none of 14 popular detectors reached 80% accuracy across conditions [Weber-Wulff et al., 2023] — while vendor pages advertise 98–99%. Both cannot be true, and the difference is always the test conditions. Our conditions are published; hold everyone, including us, to that standard. The longer version of that argument is Can AI detectors be trusted?

Has anyone independently verified this?

Not yet — and we say so on the evidence page rather than implying otherwise. The reproduction pack exists so a skeptic can re-run every number without trusting us, and the invitation to researchers and journalists is standing. When an independent evaluation exists, it will be linked here with its date.

Sources

  1. Weber-Wulff et al., “Testing of detection tools for AI-generated text,” International Journal for Educational Integrity (2023) — all 14 tested detectors below 80% accuracy, arxiv.org/abs/2306.15666.