That last part, about the key, changes how you should think about AI detection, and it is what most explainers leave out.

How a watermark gets into the words

When a language model wrote something for you, it was not picking each next word the way a lookup table would. At every position it had a whole spread of words that could have come next, each with its own probability, and it drew one of them. If the text so far was “The sky was”, then “The sky was overcast” and “The sky was grey” were both fine continuations, and each had a reasonable chance of being picked.

A watermark takes advantage of that freedom. The model still draws, but it draws from a keyed source of randomness, and that source keeps leaning it toward certain words in certain spots, so it might go for “overcast” over “grey” whenever the words just before it hash a particular way. One nudged choice on its own tells you nothing, but hundreds of them spread over a document add up to a statistical fingerprint that is very hard to argue with, provided you know which words were the favoured ones in the first place. When Anthropic described its own scheme it said the choices are still random but “the source of the randomness is different” [Anthropic, Aug 2026], and that is a good way of putting it, since the text is only being steered a little at a time, thousands of times over.

How it works under the hood (the SynthID-Text scheme)

In October 2024 a group of researchers at Google DeepMind published a paper in Nature that described a scheme they called SynthID-Text [Dathathri et al., Nature, Oct 2024], and that is the scheme Google runs in production today and the one Anthropic picked up in August 2026. It comes down to three things the model does every time it is about to write a word.

  1. It looks back at the last few words it has written and hashes them together with a secret watermarking key that only the company holds, and the number that comes out of that hash seeds a pseudorandom function for this one position. Anybody with the key could do the same sum, and that is the only way to do it.
  2. That pseudorandom function hands out a hidden score to each word that could come next, and the paper calls that score a g-value. The real scheme hands out several layers of these at each position, but one score per word is close enough to follow it by.
  3. A few candidate words get drawn from the model's normal probability distribution, the way they would have been anyway, and then they play off against each other in a little bracket, and at each layer of the bracket the word with the higher g-value wins. Whichever word wins is the one the model writes down, and then it moves on to the next position and does the whole thing again.

What comes out of all this is text that still reads naturally and still follows the model's own distribution. Google said it had run the scheme on roughly 20 million live Gemini responses without any measurable loss of quality [Dathathri et al., Nature, Oct 2024], and that surprised us when we first read it, since you would think that pushing a model toward particular words would make it write worse. Underneath, though, the text is leaning over and over again toward the words that happened to have high g-values, and that is what the other side of the scheme goes looking for.

Detection runs the same arithmetic again. With the text in front of you and the same secret key, you can work out the g-values for every position and then ask one question, which is whether the words that were actually used keep coming out with suspiciously high scores, position after position. Ordinary human writing scores about the way a coin flip would, and watermarked text scores the way a coin would if it had come up heads 800 times out of 1,000.

A tiny worked example

Suppose the text so far is “By afternoon the sky had turned” and the model's top candidates are grey (35%), overcast (30%) and dark (20%). The key and the context hash together and hand out the g-values, and say they come out as overcast 0.91, dark 0.44 and grey 0.12. The candidates get drawn by their normal probabilities, and then the tournament favours the higher g-value, which means overcast wins more often than its 30% would lead you to expect. Reading the page you see an ordinary sentence, and holding the key you see one more place where the “hot” word won. Do that at every position in a 1,000-word document and the pattern is impossible for the key holder to miss.

Why can't anyone check for the watermark?

Because if anyone could check it, anyone could remove it. The g-values come out of a keyed pseudorandom function, and without the key they look like noise, which is how the scheme was designed [Dathathri et al., Nature, Oct 2024; Google AI dev docs, accessed Aug 2026]. Imagine there had been a public checker that anybody could use. Then cheating would have had a very easy recipe, which is that you keep paraphrasing the text and running it through the checker until the checker says “clean”, and at that point the watermark is gone, and you never had to understand how any of it worked.

So the company keeps the key to itself, and that is what makes the watermark so hard to get rid of. There is a cost to that, and the cost is that only the provider can ever verify its own watermark. When Anthropic announced its scheme it said the same thing in its own words, which were that the pattern “is detectable to anyone who has a key that encodes it” [TechCrunch, Aug 11, 2026], and the key it was talking about was its own, so there is no back door and no third party holding a copy.

Google did put the SynthID-Text algorithm out as open source, and you can find it in Hugging Face Transformers if you want to try it for yourself. People get this wrong. What the open-source release gives you is a way to put a watermark on your own model's output with your own key, and then to find it again with that same key. It does not give you a way to see what Google or Anthropic have put into their production text, since their keys and their detector settings are secret [Google AI dev docs, accessed Aug 2026], and we do not expect that to change, since the moment they were shared you would be back to the paraphrasing problem above.

What breaks a watermark, and what leaves it alone

The watermark lives in the words the model picked, so anything you do that leaves those words where they are leaves the mark there with them, and anything that changes a lot of them starts to wear it away. Anthropic said as much, and so did the Nature paper [Anthropic, Aug 2026; Dathathri et al., Nature, Oct 2024], and we have put what they said in order, from the things that do nothing to the things that wipe it out.

  • Copying and pasting the text somewhere, changing the formatting, making a few light edits, or trimming a few sentences off the end all leave the watermark where it was, and it is still found.
  • Editing it moderately, mixing the AI text in with your own writing, or having only a short excerpt to work with all leave fewer positions to count, so the statistics get weaker and whoever runs the test comes away less sure of what they are looking at.
  • Paraphrasing it heavily or rewriting it from top to bottom takes it away, and Anthropic's own words were that “comprehensive rewrites will eliminate” it. Having it translated into another language does the same, since the word choices that carried the mark are gone once the text is in another language.
  • And then there is text where the watermark was barely there to begin with. Ask a model for something where there is only one right answer, like a piece of code that has to compile or a sum, and it never had a free choice to make in the first place, so hardly any watermark went in. That is what the researchers mean when they talk about low-entropy output. Anthropic has said that Claude's code output is only minimally watermarked, and that most of what is there sits in the comments [Anthropic, Aug 2026], which makes sense to us, since the comments are the only place in a piece of code where the model gets to choose its words.

Who watermarks text today (as of August 16, 2026)

Google was first, and it has been watermarking the text in the Gemini app and on the Gemini web experience with SynthID-Text since 2024, and it said so publicly at the time [Google DeepMind, accessed Aug 2026], so if you have been using Gemini for a while then everything it wrote for you carried a watermark and you would not have noticed. Anthropic came next, over a few days from August 11 to 14, 2026, and any Claude model launched after August 2 of that year now marks the text it writes, wherever you are and whatever surface you are on, with the older models to follow “over the coming months.” Anthropic also said a detection API was “coming soon” [Anthropic, Aug 2026], and as we write this it has not appeared.

OpenAI, Meta, Mistral and Microsoft have not shipped a text watermark, as far as we can tell, but every one of them signed the EU's Transparency Code of Practice in July 2026, and the EU law says that AI text from systems that were already on the market has to be marked in a machine-readable way by December 2, 2026 [European Commission, Jul 2026; ReedSmith, accessed Aug 2026]. xAI, which makes Grok, did not sign, and we could not find anything it has said about watermarking either way, though it is bound by the AI Act the same as everyone else who serves people in Europe [European Commission signatory list, Aug 2026; CityAM, Aug 12, 2026].

For what each company is doing on the day you are reading this, rather than the day we wrote it, we keep a page that reads straight from the same configuration our own detection backend uses: Who watermarks AI text today?

The awkward case: watermarked text that looks human

Say somebody hands in a long essay that was written by one of the new Claude models and was never edited. That text carries a keyed watermark, and over a document that long it is strong enough that Anthropic could show it was AI generated with a great deal of confidence, bearing in mind, as we said above, that editing and shortness both weaken the signal. But no tool other than Anthropic's can see that evidence, and Anthropic's detection API is not public.

Meanwhile a statistical scanner, and ours is one of them, looks at the same essay and may well come back with “likely human”, since the watermark does not change the statistics of the text in any way that a classifier can be relied on to catch, and the newer models are getting harder and harder to tell apart from people. So you can end up with the strongest evidence that exists, which is the keyed watermark, and the only evidence available to you, which is a statistical estimate, pointing in opposite directions on the same document. A scanner that gave you a confident “human-written” verdict would be promising more than it could deliver, since the statistics coming back empty does not mean a person wrote the text.

What this site can and cannot tell you

Cobalynx does two things, and you should know what each of them does before you trust it. The first is the text scan, where you paste some text and get back one of three statistical verdicts, which are Likely AI, Inconclusive and Likely human, and it comes to you as a calibrated probability with published error rates behind it (our methodology page spells out what each of the three means). The second is the file check, where you give us a PNG or a JPEG and we verify its C2PA Content Credentials cryptographically, or you give us a DOCX or a PDF and we read the metadata for the fingerprints that tools leave behind, and we tell you plainly that metadata is a weak signal (what a credentials check can and cannot prove is spelled out in the Content Credentials guide). Pasted text carries no file-level provenance, since pasting strips it, so on the text scan the only thing of that kind we look for is invisible Unicode artifacts, and we label that as the weak hint it is.

When there is a provenance signal to look for at all, you get one of three states.

  1. Confirmed AI is the one we are most careful with, and you only ever see it when there is a deterministic signal we can point to. Today that means a valid provenance manifest (C2PA Content Credentials) inside an image you uploaded, where the credentials themselves say that an AI made it AND the signature traces back to an authority we recognise. A manifest that traces back to nobody is a claim attached to the file, and we show you which of the two you have. Once the providers open up their detection APIs it will also mean a watermark hit that the provider itself has verified. There is no text-watermark API open yet, for the reason we gave above about the key, so there is no text scan anywhere that can reach this state.
  2. Indicators present means the statistical analysis found patterns that are typical of AI generation, and we report that as a probability with its error rate next to it, never as a verdict. That is what the live scan does, so a “likely AI” or an “inconclusive” result is this kind of evidence and not any kind of confirmation, and we would want you to treat it that way.
  3. Unknown is what you get when there is no deterministic signal and the statistics cannot settle it one way or the other. We will never label this “human-verified”, since there is nothing available that could prove a person wrote something, and we do not see that changing.

We cannot detect the text watermarks that Google and Anthropic put in, and neither can anybody else who does not have their keys. The EU's Code of Practice says that the providers have to open up interoperable detection access by February 2, 2027 [TechPolicy.Press; Stibbe, accessed Aug 2026]. When that access opens up we will connect to it, and we will update this page.

Sources

  1. Anthropic, “A text watermark for Claude,” anthropic.com/news/claude-text-watermark (announced Aug 11–14, 2026; accessed Aug 16, 2026)
  2. Dathathri et al., “Scalable watermarking for identifying large language model outputs,” Nature (Oct 2024), nature.com/articles/s41586-024-08025-4
  3. Google DeepMind, SynthID overview, deepmind.google/models/synthid/ (accessed Aug 16, 2026)
  4. Google AI developer docs, SynthID safeguards, ai.google.dev/responsible/docs/safeguards/synthid (accessed Aug 16, 2026)
  5. TechCrunch, “Anthropic says it will watermark text generated by its AI models” (Aug 11, 2026) and follow-up detail piece (Aug 15, 2026)
  6. Kirchenbauer et al., “A Watermark for Large Language Models,” ICML 2023, arxiv.org/abs/2301.10226
  7. European Commission, Code of Practice on Transparency of AI-generated Content — policy page and signatory announcement (Jul 2026; accessed Aug 16, 2026)
  8. CityAM, “ChatGPT might follow Claude's watermark pledge — but Grok to swerve it” (Aug 12, 2026)
  9. ReedSmith, analysis of the Art. 50 Commission Guidelines and Code adequacy (accessed Aug 16, 2026)
  10. TechPolicy.Press, “The EU's AI Transparency Code of Practice, Explained” (accessed Aug 16, 2026)
  11. Stibbe, “(Water-)marking the machine” (accessed Aug 16, 2026)