Aller au contenu principal
General

AI watermarking on text: why it's good news for writers

Claude now marks its text with an invisible watermark. For authors, that is finally physical evidence instead of permanent suspicion.

9 days ago7 min read
AI watermarking on text: why it's good news for writers
Woman Writing a Letter, with her Maid (c. 1670) by Johannes Vermeer — rawpixel

On 11 August 2026, Anthropic announced that text produced by Claude would carry an invisible mark. It sits inside the prose itself, as a statistical signal woven through the sentences, which a reader's eye passes straight over and a machine knows how to recover. Generated files (SVG, PNG, JPG) get a second mechanism on top: signed provenance metadata following the C2PA standard.

The reaction was immediate, and largely hostile. People saw a tracker, a betrayal, a way of branding anyone who uses the tool.

Taken from the other end, though, the problem looks different. For literature, this is precisely what was missing.

How a text watermark actually works

In one sentence: as the model picks each word, it doesn't simply keep the most probable one, because a secret key biases the draw among several near-equivalent candidates, and that bias, spread across hundreds of consecutive choices, forms a signature anyone holding the key can recover.

Nothing is added to the text. You might have imagined an invisible character, metadata bolted on the side, an exotic non-breaking space, three tricks a plain copy-paste into a bare editor would strip out. The mark sits elsewhere, in the word choices themselves.

The reference method is Google DeepMind's: SynthID-Text, described by Sumanth Dathathri and colleagues in Nature in October 2024 ("Scalable watermarking for identifying large language model outputs", Nature 634, 818-823). The mechanism has a telling name, tournament sampling: at each position, candidate words compete in pairs according to a function derived from the key, and the winner is kept. Deployed inside Gemini, it was compared across nearly 20 million responses with no significant difference in user satisfaction between watermarked and unwatermarked answers. The code was later open-sourced.

Two opposite consequences follow from that design. The watermark survives copy-paste and light edits, because it is spread across the whole text, and it dissolves as soon as you genuinely rewrite, since rewriting means redoing the word choices.

What the watermark repairs

This is precisely where it matters for an author. For three years, a writer who publishes clean, controlled, even prose has been exposed to an accusation they cannot refute. In May 2026, a shortlisted story for the Commonwealth Short Story Prize was declared "100% AI-generated" by a commercial detector; we told that story elsewhere, along with what it reveals about how reliable these tools really are.

The underlying problem is an asymmetry. Statistical detectors do not measure where a text came from. They measure how regular it is. Smooth, predictable prose with expected turns of phrase scores high, whether it came from a machine or from a diligent human. A Stanford study published in Patterns in 2023 put brutal numbers on it: across 91 TOEFL exam essays written by non-native English speakers, seven off-the-shelf detectors flagged 61% as AI-generated, and 19% were misclassified by all seven tools at once. OpenAI had pulled its own classifier that same year with barely kinder figures: 26% of machine text correctly identified, and 9% of human text wrongly accused.

The result is that an author can be accused without ever being cleared. Their defence has no object, since they are handed a percentage and the percentage looks like proof.

The watermark reverses the burden. It speaks about the machine rather than the human: instead of guessing whether prose is too clean to be someone's, it registers the passage of a specific model, or else stays silent. We leave presumption behind for physical evidence, which changes the very nature of the question being asked.

What the watermark does not prove

Anthropic is explicit on this, and credit where it's due: the mark indicates only that content passed through Claude.

It does not say the model had the idea. A text written entirely by hand, then handed to the model for a proofread, a translation or a format conversion, can come back marked. Conversely, the absence of a mark proves nothing, given how many reasons there are to find nothing: an older model, a passage too short, heavy rewriting, an uncovered platform.

That is the hard limit of the system, and it covers precisely the zone where the argument is fiercest. Between the author who has their punctuation fixed and the one who has their chapter written, the watermark does not adjudicate, and both texts carry the same mark. Yet that is exactly where the line that matters sits, and we've proposed a word for naming it.

The technical fragility, and the asymmetry it creates

It has to be said that the system is fragile, and its designers do not hide it.

Robustness evaluations published since 2024 converge on the same finding: a single paraphrasing pass by another model is enough to push detection rates below 0.3 across every method tested. Synonym substitution, rewriting targeted at the most information-bearing words, round-trip translation: the attacks are known, documented, and some of them trivial. On the file side, C2PA metadata can be stripped with open-source tools.

There is an uncomfortable asymmetry here, and it would be dishonest to skate over it. Someone who wants to cheat removes the mark in one operation, while someone who isn't cheating keeps it. The watermark therefore catches the careless and the honest far more reliably than the determined.

OpenAI drew its own conclusions: the company has had a watermarking tool for several years that it describes as highly reliable on long enough texts, and has never shipped it. The stated reasons are instructive: the risk of stigmatising non-native speakers, for whom these tools are a legitimate linguistic crutch, and the fear of watching users migrate to a competitor that doesn't mark.

The other side of the coin

The objection now being raised deserves to be taken seriously.

Not every use of AI is suspect. A writer producing a product description, technical documentation or an email sequence never claimed to be making art. Marking risks penalising them twice over: with readers who will read the label as a confession, and with search engines and aggregators, which will have a machine-readable signal at scale for the first time. On whether they will feed it into their ranking, every one of them stays quiet.

That is probably where the anger comes from. It rises less from novelists than from the trades where AI has become an ordinary production tool, and where the mark lands like a retroactive penalty on an openly held practice.

One difference separates the two situations, and the debate mostly talks past it. In commercial writing, the value is in the result: a text that informs and converts is worth what it's worth, whatever hand produced it. In literature, value is inseparable from provenance. A novel is not a service rendered to the reader, it is someone speaking, and removing the author from the book does not degrade the product but deletes the object.

So the same mechanism is a constraint on one side and a repair on the other. Both camps are right at the same time, which is why they get along so badly.

The legal frame, briefly

Anthropic did not act out of philanthropy. Article 50 of the European AI Act has applied since 2 August 2026: providers of generative systems must mark their outputs (audio, image, video and text) in a machine-readable format. On 10 June 2026 the European Commission published a voluntary code of practice on marking and labelling AI-generated content, signed by around 190 companies and organisations as of late July. Among other things, it calls for clear labelling of generated content on matters of public interest.

The scope Anthropic announced reaches beyond Europe: marking applies everywhere, on the API as well as the apps, and across cloud partners. Models launched from 2 August 2026 mark at release, while earlier ones follow progressively.

One point is worth holding onto: the public detection tool does not exist yet. The company says it is working on it and points to forthcoming technical documentation. Until it ships, the watermark remains a promise more than an instrument.

What a writer can do today

In the meantime, the best evidence is still the one an author builds themselves, and there is nothing technological about it: the trace of the work.

A human manuscript has a history, made of successive versions, chapters reworked four times, abandoned threads and dates. That history is very hard to fabricate after the fact, and far more convincing than a detector score, because it shows the road and not just the arrival. That is exactly what snapshots are for in Extypis: a timeline of named versions, comparable side by side, restorable in one click. The day a doubt arises, an author who can scroll through six months of versions has little to fear from a percentage.

A second, more prosaic reflex: declare what needs declaring. Amazon KDP has required since late 2023 that AI-generated content be disclosed when publishing or updating a book, and explicitly distinguishes generated from assisted content: brainstorming, grammar checking and refining text the author wrote do not need disclosing. The declaration stays internal to the platform and does not appear on the book's page.

An imperfect clue beats a permanent suspicion

The watermark will not settle the question of AI in literature. It comes off in one operation, it conflates proofreading with ghostwriting, and the tool that would let anyone read it is still to come.

It does, however, move the question away from the human and onto the machine, which no detector had managed until now. For three years the only available move was to suspect an author of writing too well, a dead end that hurt honest people first.

For those who actually write, the mark is therefore good news, and the best reason for that is simply that it doesn't concern them.

HU

Hubert Giorgi

Author

Write with Extypis

Writing a novel?

Extypis holds the outline, the sheets and the six hundred pages while the text keeps changing. Free to begin, free to finish.

No credit card. You can write and export a whole book without paying.