Detection basics
AI detector #
An AI detector is a tool that estimates how likely a text is to have been generated by a language model, by measuring statistical patterns in the writing. It produces evidence, not proof: every detector has a measured error rate, and a score alone can never establish authorship. Cobalynx publishes its own error rates on the evidence page.
AI checker #
“AI checker” is a synonym for AI detector — a tool that checks text for statistical signs of machine generation. The name does not change the mathematics: a checker returns a probability with an error rate, and any checker that reports certainty is overclaiming.
Perplexity #
Perplexity measures how predictable a text is to a language model — low perplexity means the model would likely have chosen the same words. A worked example: after “The sky was a brilliant shade of”, a model finds “blue” highly predictable (low perplexity), “cerulean” a little surprising (medium), and “accountant” very surprising (high perplexity) — a text that stays relentlessly in the low-perplexity lane reads as machine-smooth. AI-generated text tends to sit at lower perplexity than human writing, which is why the signal is useful; but formal, conventional human prose is also highly predictable, which is why perplexity alone produces false positives.
Burstiness #
Burstiness is the variation in rhythm and predictability across a text — how much sentence length, structure, and word surprise fluctuate. Human writing tends to be bursty (short sentences beside long ones, plain phrases beside odd ones); model output tends toward uniformity. Like every single signal, it is weak alone and meaningful only in aggregate.
AI-generated vs. AI-refined #
AI-generated text was produced by a model from a prompt; AI-refined text started as human writing and was polished by a model or grammar tool. Detectors measure statistical origin, so heavy refinement moves human text toward machine-typical patterns — one of the main honest reasons false positives happen to real writers.
The statistics of an honest verdict
Calibrated probability #
A calibrated probability is a score that means what it says on the population it was measured on: among texts in our evaluation set given 80%, about 80% were actually AI-generated. Calibration is what separates a measurement from a vibe — an uncalibrated “83% AI” is just a dial with numbers painted on it — but what a score means in your context also depends on the base rate. How ours is produced is on the methodology page.
Calibration #
Calibration is the process of mapping a detector's raw internal scores onto real-world probabilities using measured outcomes. Without it, two detectors can score the same text 12% and 88% and both be “working as designed.” It is why cross-tool disagreement is normal — detectors are not calibrated to each other.
False positive #
A false positive is human writing that a detector wrongly labels as likely AI — the most damaging error this category of tool can make. Formal structure, careful grammar, and non-native English are all documented triggers. Any detector that will not publish its false-positive rate is asking you to take this risk on faith.
False negative #
A false negative is AI-generated text that a detector fails to flag — it passes as human-written. Short texts, edited output, and paraphrase attacks all raise the false-negative rate. A tool that never misses does not exist; a tool that claims so has simply not published its misses.
False-positive rate (FPR) #
The false-positive rate is the fraction of genuinely human texts that a detector wrongly flags, measured on a known test set. It is the single most important number a detector can publish, because it is the probability of harming an innocent writer. Ours is published with raw counts and intervals on the evidence page.
Confidence interval #
A confidence interval is the honest range around a measured rate — “4% (95% CI 2–7%)” says the measurement is an estimate from a finite sample, not a constant of nature. A published accuracy figure without an interval and raw counts cannot be checked and should not be trusted.
Abstention (inconclusive verdict) #
Abstention is a detector declining to give a verdict because the evidence is too thin — the honest answer for short, mixed, or genuinely borderline text. A detector that never says “inconclusive” is not more capable; it is converting uncertainty into confident-sounding errors. Cobalynx treats the abstention band as a feature, not a failure.
Operating point (decision threshold) #
The operating point is the probability threshold at which a detector turns a score into a verdict label — where “likely AI” begins and “likely human” ends. Error rates only mean something at a stated operating point, which is why moving the threshold after measuring is a classic way to make a bad number disappear.
Base rate #
The base rate is how common AI text actually is in the population you are checking — and it changes what a flag means. If few texts in a stack are AI, even a low false-positive rate means many flags are wrong. This is the arithmetic behind our rule that a score should open a conversation, never close one — our base-rate explainer does it with our own published rates.
Non-native (ESL) false-positive bias #
AI detectors as a class flag writing by non-native English speakers at elevated rates, because learned, careful, conventional prose resembles machine output statistically. The bias is documented in peer-reviewed research and is a fairness problem, not a detail. We measure and publish our own ESL error rate separately on the evidence page.
Watermarks & provenance
AI text watermark #
An AI text watermark is a statistical pattern hidden in a model's word choices at generation time — no invisible ink, no secret metadata, just biased picks that the generator's operator can later test for. Only the key holder can check it, and editing or paraphrasing degrades it. How text watermarks work has the full story.
SynthID-class keyed watermark #
A keyed watermark (SynthID-Text is the best-known example) shapes a model's word choices with a secret key, so only whoever holds the key can verify the mark. That is by design: a publicly checkable text watermark would also be publicly removable. It is why no third-party tool — ours included — can honestly claim to read another company's text watermark.
C2PA / Content Credentials #
C2PA Content Credentials are a cryptographically signed manifest inside a media file recording who or what created and edited it — an open standard backed by camera makers and the major AI labs. A valid AI-generator manifest is near-conclusive evidence of origin; a missing one proves nothing, because most files never had credentials. You can read a file's credentials here.
Provenance #
Provenance is evidence about where a file came from, read from the file itself — signed manifests, tool records, edit chains — rather than guessed from its content. Provenance is deterministic where detection is statistical: when it is present it is strong, and when it is absent an honest tool says “unknown,” not “human.”
Machine-readable marking #
Machine-readable marking is labeling AI output so software — not just people — can recognize it: embedded credentials, watermarks, or structured metadata. It is the technique EU law now requires of AI providers for synthetic content, and the reason provenance standards are spreading beyond photography into text and video tooling.
The law
EU AI Act — Article 50 #
Article 50 of the EU AI Act is the transparency rule: since August 2, 2026, providers must mark AI-generated content machine-readably, chatbots must disclose they are not human, and deepfakes must be labeled. The duty sits on providers and deployers, not on readers. What the law actually says covers deadlines, penalties, and who is bound.
AI disclosure #
AI disclosure is telling people that content or a conversation partner is AI — voluntarily, by platform policy, or by law. In the EU it is now a legal right in chatbot interactions; elsewhere it is mostly an emerging norm. Disclosure duties bind the party deploying the AI; a detector can inform your suspicion but cannot replace the duty.
Evasion & its economy
Humanizer #
A “humanizer” is a rewriting tool sold to make AI text evade AI detectors. Independent testing shows they degrade every detector's recall — including ours, and we publish those adversarial numbers instead of claiming immunity. We do not build or sell evasion tools: a vendor selling both the alarm and the lockpick has told you what its detector is worth.
Paraphrase attack #
A paraphrase attack rewords AI-generated text — by another model or by hand — to blur the statistical fingerprint detectors measure. It is the single most effective known evasion, and every honest evaluation includes it. Our measured recall under paraphrase attack is published on the evidence page.
AI slop #
AI slop is low-effort machine-generated content published at scale for clicks or commerce — flooded reviews, spun articles, fake product photos. Slop is a volume problem: any single item may pass, but sameness across a batch is exactly what statistical checking can see, which is why batch checking works.
General AI terms, honestly
Large language model (LLM) #
A large language model is a neural network trained on vast text corpora to predict the next token, which turns out to be enough to write, summarize, translate, and converse. LLMs generate the text that AI detectors try to recognize — and detectors themselves rely on the same statistical regularities that make LLM output fluent.
Token #
A token is the unit a language model reads and writes — usually a word piece a few letters long, not a whole word. Models choose one token at a time by probability, which is why statistical detection is possible at all: generation leaves a measurable trail of typical choices.
Training data #
Training data is the text corpus a model learns from. For detectors it cuts twice: what a detector saw in training bounds what it can recognize, and what it never saw is where it fails quietly. Cobalynx never uses text you submit as training data — the promise and its enforcement are on the privacy page.
Hallucination #
A hallucination is a language model stating something false with fluent confidence — invented citations, fabricated facts, plausible nonsense. It matters here for a subtler reason: detection verdicts estimate statistical origin, not truth. A human can lie and a model can be right; no text checker measures honesty.
Zero retention #
Zero retention means a service processes your input in memory and keeps nothing after answering — no stored copies, no training reuse, nothing to breach later. The claim is only worth what enforces it: ours is release-tested by an automated no-persistence check, described on the privacy page.