The problem isn't detection — it's what scores get used for
Several universities publicly disabled AI detection after false accusations reached students who had done nothing wrong. What failed wasn't the idea of checking; it was the workflow: a percentage from an uncalibrated tool, treated as a verdict, applied without process. The same tool output that is genuinely useful as a reason to look closer is genuinely harmful as a reason to punish. This page is about staying on the right side of that line.
Reading a calibrated result
“AI DETECTED” tells you nothing about how often that alarm is false. A calibrated result does: Cobalynx reports the probability that a text is AI-generated and publishes how often each confidence band has been wrong on a frozen benchmark — the error rates, with raw counts and confidence intervals, are public. Three practical consequences for grading: a high-confidence flag still carries a known, nonzero false-positive rate; an inconclusive verdict means the evidence genuinely does not support a conclusion — it is not a soft “probably AI”; and a likely-human verdict is not proof of anything either, because no statistical method can certify human authorship. How the scores are produced is documented on the methodology page.
The multi-evidence protocol
Every serious integrity framework now converges on the same practice: a detector score may open a conversation, never close one. Before acting on any flag:
- Version history. Ask the student to share the document's revision timeline (Google Docs, Word AutoSave). Hours of incremental edits are strong evidence of authorship; a single paste is a real question worth asking about.
- Prior work. Compare voice and level against earlier writing from the same student, collected before the assignment.
- Drafts and process. Outlines, notes, bibliography trails — the residue of real work.
- The conversation. Ask the student to walk through their argument and choices. Someone who wrote the essay can discuss it; this single step resolves most cases in both directions.
If, after all four, the evidence still points one way — you have a case built on process, not on a percentage. If it doesn't, you almost accused someone on a coin flip dressed as a certainty.
The bias you are ethically required to know about
Mainstream detectors flag non-native English writing at dramatically elevated rates — a Stanford-led study measured over 60% of TOEFL essays falsely flagged across seven detectors [Liang et al., 2023]. Formal, careful, grammar-polished prose — the writing of your most conscientious students — triggers the same failure mode. We measure and publish our own non-native (ESL) false-positive rate and hold it to the same gate as every other register, because a bias you don't measure is a bias you deploy. Whatever tool you use, ask it for this number; if it doesn't publish one, weight its flags accordingly.
A fair-process checklist for your syllabus
- State up front what AI use is allowed, and that flagged work triggers a conversation, not a penalty.
- Never act on a score alone, from any tool — ours included.
- Ask for process evidence before forming a view, and give the student the chance to show it.
- Apply extra caution with non-native speakers; the false-positive skew is documented.
- Document what the tool reported — the probability and its published error rate, not just “flagged.”
- Know what your students will read if accused: their side of this page is the false-accusation playbook — a fair process survives contact with it.
What Cobalynx can and can't do here
No detector output — ours included — is proof of misconduct. We publish exactly how often each confidence band is wrong so a score can inform a conversation, never replace one. We can't tell you who typed the words, and we won't pretend otherwise with a certainty stamp: text too short to judge is refused, borderline evidence is reported as inconclusive, and every number traces to a public measurement. Why detectors disagree — and what an honest accuracy claim looks like — is covered in Can AI detectors be trusted?; the checker itself is free, with no signup.
Sources
- Liang, Yuksekgonul, Mao, Wu, Zou, “GPT detectors are biased against non-native English writers,” Patterns (2023), arxiv.org/abs/2304.02819.
- Weber-Wulff et al., “Testing of detection tools for AI-generated text,” International Journal for Educational Integrity (2023), arxiv.org/abs/2306.15666.
- Public university decisions to disable AI-detection features after false-accusation concerns (UCLA, UC San Diego, Vanderbilt, among others; accessed Aug 2026).
Common questions
Is an AI-detection score enough to fail a student?
No — no detector output, ours included, is proof of misconduct. A score is one probabilistic signal with a measured error rate; institutions that treated it as a verdict generated documented false accusations. Pair any score with version history, prior work, and a conversation before drawing conclusions.
What does calibrated confidence actually tell me?
It tells you the probability the text is AI-generated and how often that confidence band has been wrong on a public benchmark — so “uncertain” genuinely means uncertain rather than a dramatized alarm. The measured band-level error rates live on the evidence page.
Why did several universities stop using AI detectors?
Because scores were being used as verdicts, and the false positives — disproportionately hitting non-native speakers and conscientious, formal writers — caused real harm. The lesson is not that detection is worthless; it is that an uncalibrated score in an unfair process is worse than no score at all.
How do I handle a flagged non-native speaker fairly?
With extra caution: research (including a widely cited Stanford study) shows mainstream detectors flag non-native English at dramatically elevated rates. Compare the flagged work against the student’s prior writing, ask for drafts, and weight process evidence over the score — we measure and publish our own non-native error rate on the evidence page for exactly this reason.