The direct answer
On our frozen evaluation of August 16, 2026 (1522 documents; engine v1 — the engine serving every scan today), the false-positive rate on human writing at the published operating point was 0.5% (3 of 609; 95% CI 0.2%–1.4%). Detection by length: at 300–1,000 words we detect 81.6% of clean AI documents (346 of 424 non-abstained) at a 0.6% human false-positive rate (3 of 525); at 1,000+ words, 100.0% (18 of 18) with 0 false positives in 20 human documents; below 300 words the clean-AI evaluation slice is empty by design — we abstain or refuse rather than guess.
If that answer looks longer than a competitor's “99% accurate”, that is the point: accuracy depends on text length, language, and editing, so the honest answer to “how accurate” is a table, not a number. The single-number version of this page would be marketing.
How often each verdict is wrong
The most useful accuracy question is conditional: given the verdict you just read, how often is it wrong? Measured on the same frozen run:
| What happened | Counts | Rate | 95% interval |
|---|---|---|---|
| Said “likely AI”, text was actually human | 3 of 442 | 0.7% | 0.2% – 2.0% |
| Said “likely human”, text was actually AI | 223 of 829 | 26.9% | 24.0% – 30.0% |
| Answered “inconclusive” instead of guessing | 251 of 1522 | 16.5% | 14.7% – 18.4% |
The “inconclusive” row is deliberate, not damage: on genuinely borderline text we abstain rather than guess, because a confident wrong answer is the one outcome this product exists to avoid.
The numbers a marketing page would hide
- Paraphrase attack: AI text deliberately reworded to evade detection degrades every detector, ours included. Our measured adversarial recall: 33.5% of non-abstained adversarial documents (73 of 218; 23.5% counting abstentions as misses). We publish this instead of claiming immunity; the mechanics are in the glossary.
- Short text: below the word floor we refuse to answer at all — there is no measured accuracy for a case we decline, and any tool quoting one is guessing.
- Non-native (ESL) writing: detectors as a class false-flag non-native English at elevated rates, so we measure that slice separately and hold it to the same gate: our ESL false-positive rate is 1.7% (2 of 118; 95% CI 0.5%–6.0%).
How these numbers were produced
One sanctioned, seeded evaluation run on a frozen held-out corpus, with every dependent artifact checksummed and every execution of the measurement command disclosed in an append-only ledger. The full tables, the corpus composition, and the step-by-step reproduction pack are on the evidence page; the pipeline itself is documented on the methodology page. Calibration — what makes the probability we show mean what it says — is abstention-free AUROC 0.952 on the clean evaluation split, 0.938 when non-native human writing is included (what calibration means).
How to compare us with anyone else
Ask any detector vendor the same three questions: at what decision threshold was the accuracy measured, on what test set, and with what raw counts and intervals? Peer-reviewed cross-tool testing found none of 14 popular detectors reached 80% accuracy across conditions [Weber-Wulff et al., 2023] — while vendor pages advertise 98–99%. Both cannot be true, and the difference is always the test conditions. Our conditions are published; hold everyone, including us, to that standard. The longer version of that argument is Can AI detectors be trusted?
Has anyone independently verified this?
Not yet — and we say so on the evidence page rather than implying otherwise. The reproduction pack exists so a skeptic can re-run every number without trusting us, and the invitation to researchers and journalists is standing. When an independent evaluation exists, it will be linked here with its date.
Sources
- Weber-Wulff et al., “Testing of detection tools for AI-generated text,” International Journal for Educational Integrity (2023) — all 14 tested detectors below 80% accuracy, arxiv.org/abs/2306.15666.