On 11 August 2026, Anthropic announced that text produced by Claude would carry an invisible mark. It sits inside the prose itself, as a statistical signal woven through the sentences, which a reader's eye passes straight over and a machine knows how to recover. Generated files (SVG, PNG, JPG) get a second mechanism on top: signed provenance metadata following the C2PA standard.
The reaction was immediate, and largely hostile. People saw a tracker, a betrayal, a way of branding anyone who uses the tool.
Taken from the other end, though, the problem looks different, and for literature this mechanism supplies precisely what was missing.
How a text watermark actually works
As the model picks each word, a secret key biases the draw among several near-equivalent candidates instead of simply keeping the most probable one, and that bias, spread across hundreds of consecutive choices, forms a signature anyone holding the key can recover.
The mark therefore lives entirely in the choice of words, and the text is otherwise unchanged. An invisible character, metadata bolted on the side or an exotic non-breaking space would all have vanished at the first copy-paste into a bare editor.
The reference method is Google DeepMind's: SynthID-Text, described by Sumanth Dathathri and colleagues in Nature in October 2024 ("Scalable watermarking for identifying large language model outputs", Nature 634, 818-823). The mechanism has a telling name, tournament sampling: at each position, candidate words compete in pairs according to a function derived from the key, and the winner is kept. Deployed inside Gemini, it was compared across nearly 20 million responses with no significant difference in user satisfaction between watermarked and unwatermarked answers. The code was later open-sourced.
Two opposite consequences follow from that design. The watermark survives copy-paste and light edits, because it is spread across the whole text, and it dissolves as soon as you genuinely rewrite, since rewriting means redoing the word choices.
What the watermark repairs
This is precisely where it matters for an author. For three years, a writer who publishes clean, controlled, even prose has been exposed to an accusation they cannot refute. In May 2026, a shortlisted story for the Commonwealth Short Story Prize was declared "100% AI-generated" by a commercial detector; we told that story elsewhere, along with what it reveals about how reliable these tools really are.
The underlying problem is an asymmetry. Statistical detectors measure how regular a text is, which says nothing about where it came from. Smooth, predictable prose with expected turns of phrase scores high, whether it came from a machine or from a diligent human. A Stanford study published in Patterns in 2023 put brutal numbers on it: across 91 TOEFL exam essays written by non-native English speakers, seven off-the-shelf detectors flagged 61% as AI-generated, and 19% were misclassified by all seven tools at once. OpenAI had pulled its own classifier that same year with barely kinder figures: 26% of machine text correctly identified, and 9% of human text wrongly accused.
The result is that an author can be accused without ever being cleared. Their defence has no object, since they are handed a percentage and the percentage looks like proof.
The watermark reverses the burden. It concerns the machine rather than the human, since instead of guessing whether prose is too clean to be someone's, it registers the passage of a specific model, or else stays silent. We leave presumption behind for physical evidence, which changes the very nature of the question being asked.
What the watermark does not prove
Anthropic is explicit on this, and credit where it's due: the mark indicates only that content passed through Claude.
It leaves open who had the idea for the text, because a text written entirely by hand, then handed to the model for a proofread, a translation or a format conversion, can come back marked. Conversely, the absence of a mark proves nothing, given how many reasons there are to find nothing: an older model, a passage too short, heavy rewriting, an uncovered platform.
That is the hard limit of the system, and it covers precisely the zone where the argument is fiercest. Between the author who has their punctuation fixed and the one who has their chapter written, the watermark does not adjudicate, and both texts carry the same mark. Yet that is exactly where the line that matters sits, and we've proposed a word for naming it.
The technical fragility, and the asymmetry it creates
It has to be said that the system is fragile, and its designers do not hide it.
Robustness evaluations published since 2024 converge on the same finding: a single paraphrasing pass by another model is enough to push detection rates below 0.3 across every method tested. Synonym substitution, rewriting targeted at the most information-bearing words, round-trip translation: the attacks are known, documented, and some of them trivial. On the file side, C2PA metadata can be stripped with open-source tools.
There is an uncomfortable asymmetry here, and it would be dishonest to skate over it. Someone who wants to cheat removes the mark in one operation, while someone who isn't cheating keeps it. The watermark therefore catches the careless and the honest far more reliably than the determined.
OpenAI drew its own conclusions: the company has had a watermarking tool for several years that it describes as highly reliable on long enough texts, and has never shipped it. The stated reasons are instructive: the risk of stigmatising non-native speakers, for whom these tools are a legitimate linguistic crutch, and the fear of watching users migrate to a competitor that doesn't mark.
The other side of the coin
The objection now being raised deserves to be taken seriously.
Many uses of AI are entirely ordinary. A writer producing a product description, technical documentation or an email sequence never claimed to be making art. Marking risks penalising them twice over: with readers who will read the label as a confession, and with search engines and aggregators, which will have a machine-readable signal at scale for the first time. On whether they will feed it into their ranking, every one of them stays quiet.
That is probably where the anger comes from. It rises less from novelists than from the trades where AI has become an ordinary production tool, and where the mark lands like a retroactive penalty on an openly held practice.
One difference separates the two situations, and the debate mostly talks past it. In commercial writing, the value is in the result: a text that informs and converts is worth what it's worth, whatever hand produced it. In literature, value is inseparable from provenance. In a novel, someone speaks to the reader, so removing the author from the book deletes the very object it claimed to offer.
So the same mechanism is a constraint on one side and a repair on the other. Both camps are right at the same time, which is why they get along so badly.
The legal frame, briefly
Anthropic acted mainly under the pressure of the law. Article 50 of the European AI Act has applied since 2 August 2026: providers of generative systems must mark their outputs (audio, image, video and text) in a machine-readable format. On 10 June 2026 the European Commission published a voluntary code of practice on marking and labelling AI-generated content, signed by around 190 companies and organisations as of late July. Among other things, it calls for clear labelling of generated content on matters of public interest.
The scope Anthropic announced reaches beyond Europe: marking applies everywhere, on the API as well as the apps, and across cloud partners. Models launched from 2 August 2026 mark at release, while earlier ones follow progressively.
One point is worth holding onto: the public detection tool does not exist yet. The company says it is working on it and points to forthcoming technical documentation. Until it ships, the watermark remains a promise more than an instrument.
What a writer can do today
In the meantime, the best evidence is still the one an author builds themselves, and there is nothing technological about it: the trace of the work.
A human manuscript has a history, made of successive versions, chapters reworked four times, abandoned threads and dates. That history is very hard to fabricate after the fact, and far more convincing than a detector score, because it shows the whole road that led to the final text. That is exactly what snapshots are for in Extypis: a timeline of named versions, comparable side by side, restorable in one click. The day a doubt arises, an author who can scroll through six months of versions has little to fear from a percentage.
A second, more prosaic reflex: declare what needs declaring. Amazon KDP has required since late 2023 that AI-generated content be disclosed when publishing or updating a book, and explicitly distinguishes generated from assisted content: brainstorming, grammar checking and refining text the author wrote do not need disclosing. The declaration stays internal to the platform and does not appear on the book's page.
An imperfect clue beats a permanent suspicion
The watermark will not settle the question of AI in literature. It comes off in one operation, it conflates proofreading with ghostwriting, and the tool that would let anyone read it is still to come.
It does, however, move the question away from the human and onto the machine, which no detector had managed until now. For three years the only available move was to suspect an author of writing too well, a dead end that hurt honest people first.
For those who actually write, the mark is therefore good news, and the best reason for that is simply that it doesn't concern them.
Hubert Giorgi
Author
The writing studio novelists were missing
Outline, characters, places and objects, the links between them, mind map, scene board, export: everything lives in one place. Plus an AI that knows your story, whenever you want one.
Free to write and to export. No credit card asked.
