Evidence note · AI attribution

Claude Left a Mark. What Can Investigators Prove?

What Claude's text watermark can establish, where provider records may extend the trail, and why attribution to a person still requires corroboration.

Published
August 20, 2026
Updated
August 26, 2026
Author
Kevin V. Watson
Reading time
12 min
Format
Evidence note

Share this analysis

No social platform widgets are loaded.
LinkedInEmail

In August 2026, Anthropic announced machine-readable marking for supported Claude-generated content, beginning with models launched in the European Union on or after August 2. The marking applies worldwide wherever those supported models are offered, while support for earlier models remains in progress. For text, the company uses a version of Google DeepMind's SynthID-Text approach. The method uses low-stakes word choices to create a pattern that a reader cannot see, but that a holder of the detection key can recognize. The mark travels with the text when it is copied and may survive some editing. It carries no information identifying a person, an organization, or a chat.

For anyone who works in investigations, the implication arrives fast. If a document can be linked to an AI model, the provider may hold relevant records, and those records may be obtainable through appropriate legal process. The distance from an anonymous document to a named person looks shorter than it did a month earlier.

It is worth slowing down at that exact point. The inference that feels strongest here is the one most likely to fail. To see why, it helps to return to what attribution was before AI entered the picture, and to be honest about what that older method actually delivered.

What the traditional chain established

For years, one familiar form of digital attribution followed a chain most investigators could recite from memory. Where the available message and provider records exposed a source address, that address could resolve to an internet service provider. The provider, under legal process, could resolve it to a subscriber account. From there the investigation narrowed toward a person.

That chain could support investigations and prosecutions when the remaining links were corroborated. It was also narrower than the way it was often described. It attributed a connection, not a person. It established that a particular account, on a particular network, at a particular time, was associated with the activity in question.

An account is not a person. A subscriber line is not a pair of hands. Households share connections. Credentials get stolen. A single address can sit behind a router used by five people or a VPN used by five thousand. The chain narrowed the population of possible actors, sometimes to a single household. It did not, by itself, place anyone at the keyboard.

This is the part the public tends to forget. Attribution was never a single link that reached the human. It was a set of links that reached an account, followed by corroboration that closed the remaining distance. Device forensics. Behavioral correlation. Physical evidence. Testimony. The chain pointed. Convergence closed. Hold that structure in mind, because AI does not remove it. It moves the starting point.

The new artifact, and the mark it now carries

Traditional attribution had a useful property. The artifact carried its own routing information. An email header is self-describing. It states where it claims to have come from, and that claim can be tested against provider records. The artifact and the trail were part of the same object.

Generated text did not work this way. For most of AI's short history, the output of a model was a poor self-attributing artifact. There was no durable marker in the text that identified the model, the account, or the session that produced it. The watermark changes part of this. The artifact can now indicate that a particular provider's model was involved. That is a real shift, and it is why the question deserves fresh attention.

But the mark has to be read for what it says, not for what an investigator hopes it says. A detected mark supports an inference that Claude was involved in producing or processing the text. It is not presented as proof that Claude authored the document. And it carries no information identifying a person, an organization, or a chat. That single design fact controls everything downstream.

Provider provenance is not generation-instance attribution

The tempting chain runs like this. Detect the watermark. Compel the provider's records. Identify the account. Each step feels like the previous one continued. It is not.

The watermark can support an inference that Claude was involved in producing or processing the text. It does not identify the generation event that produced it. It carries no account identifier, no session identifier, no address, no organization, and no user identity. So the step that would locate a specific session has nothing to act on. To reach a session record through appropriate legal process, an investigator needs an identifier, a time window, or an account already in view. The mark supplies none of these. It tells you Claude was likely involved. It does not tell you which generation, among an enormous number, produced this document.

This is the same asymmetry that has always shaped attribution, now in a new form. Working forward from a known account to its outputs is tractable. Working backward from an orphan document to the account that produced it is the hard direction. The email header pointed backward on its own. The watermark does not. It names the vendor and stops there.

So the watermark is not the key that unlocks the account. It is the reason to approach the provider at all. Any actual link has to come from somewhere else.

Where attribution actually lives now

If the mark does not carry the trail, the trail sits with the provider. That is where the comparison to traditional forensics both holds and breaks.

It holds because the mechanism is familiar. Compel the provider's records, within the retention window and through appropriate legal process, and you may connect a generation to an account, and an account to an address and a signup identity. From there the classic chain resumes: address to provider to subscriber, closed by corroboration.

It breaks because the record is more conditional than the old one was. Traditional attribution often benefited from transactional records generated as communications crossed provider infrastructure. Those records were never guaranteed, and retention was always a variable. But where they existed and survived the retention period, they gave investigators identifiers to work backward from. AI content records shift even further toward discretion. They are governed by policy, not physics.

The retention periods define the practical window. For consumer accounts, conversations may persist until the user deletes them, after which backend deletion generally occurs within thirty days. Data a user allows to be used for model improvement can remain for up to five years. On the commercial API, inputs and outputs generally follow a thirty-day deletion window, and approved zero-data-retention arrangements can narrow it further, although limited safety records may still be retained. Inputs and outputs flagged for policy violations may remain for up to two years, with trust and safety classification scores retained for up to seven years.

If an interaction is flagged by a provider's safety systems, the longer retention period may preserve records that would otherwise have been deleted. An investigator cannot assume that an illicit use was detected or flagged, however. Retention is therefore an evidentiary variable, not an assumption.

Retention is not the same as retrievability

There is a further limit, and it deserves to be stated as a limit rather than assumed past. Retention establishes availability, not retrievability by arbitrary content. Stored does not mean indexed. Indexed does not mean searchable by an orphan text fragment. And searchable does not mean uniquely attributable to a single generation event.

Whether a provider can perform reverse content matching across retained generations is a separate technical and legal question. The watermark does not answer it. It would be a mistake to infer the capability merely because it would complete the chain an investigator wants to draw.

Editing splits the signal in two

A document used in a crime is rarely the raw output of a model. It gets trimmed, personalized, and folded into other text. That editing acts on two different problems, at two different rates, and the difference is easy to miss.

Detection asks whether the statistical watermark remains. The signal is spread across many token choices, so a partially edited document can still carry enough marked tokens to register. This is the durable half. It survives moderate editing.

Provider-side correlation asks a different question: can any retained record be connected to this particular artifact? That may require the document to resemble text the provider retained, depending on how its records are stored, indexed, and searched. Editing widens the gap between the examined document and any stored output. Provider-side correlation may therefore fail before the watermark becomes undetectable, but the provider's actual retrieval capabilities would have to be established rather than assumed.

So the same edit can leave a surviving detection and a broken path to the account. The watermark may survive while the route to a generation instance disappears. It is tempting to assume that if the mark held, the traceability held with it. Under editing, the two come apart.

The mark implies an invocation. It does not identify the human behind it.

A generation event requires an initiating input or process. A detected mark can therefore support the existence of model involvement, subject to the detector's operating characteristics and the adequacy of the sample. Something invoked the model. That much is a defensible inference.

But invocation is not the same as a person at a keyboard. A human may have prompted the model directly. An application may have called it through an API. An autonomous agent may have produced the instruction as one step in a longer workflow. Somewhere upstream there may be human agency. Upstream agency is not the same thing as immediate action.

The evidentiary chain therefore cannot jump from "the model was involved in processing this text" to "this person instructed the model to produce it." The break sits between attributable system involvement and attributable human responsibility. Closing that break requires evidence of control, authorization, knowledge, or direction. Evidence that the system acted is not evidence of who directed it, or why.

A lead, not a finding

At this point, method governs what the signal can support. The distinction between an artifact and a finding does the work.

A watermark detection is an artifact. It is an observable fact about a document. Promoting an artifact to a finding requires corroboration, and on edited text the distance between the two is wide. Recording the detection honestly means sorting what it actually supports.

Known: When examined through an authorized and technically validated detection mechanism, the document produced a positive watermark result. The result indicates statistical evidence consistent with Claude having generated or processed the text, subject to the detector's operating characteristics and the adequacy of the sample.

Assumed: Enough relationship remains between the examined document and any provider-retained generation for account-level correlation to be technically possible.

Undetermined: Which generation event produced the text, what account or system initiated it, who controlled that account or system, what the initiating instruction contained, and whether that person is connected to the criminal use of the document.

Notice the structure. Known describes the observation. Assumed describes the bridge. Undetermined holds the attribution questions that remain open. Written this way, the detection can be a strong artifact while supporting only a narrow finding. The narrow finding is that Claude likely processed the text. It does not yet support the claim that a named person authored a threat using Claude. The evidential hierarchy stays intact.

This is the discontinuity that has to be held open rather than closed. The break sits between evidence of system involvement and human responsibility, which is not yet established. The failure mode is collapsing the two, letting "Claude was involved" slide into "this person did it." The risk is higher here than it was with an address, because an AI-origin signal names a recognizable system. That recognizability is seductive. It invites the collapse. The discipline exists to resist it.

The answer is the one forensics has always given. No single node stands alone. The watermark converges with an account or system record showing a matching generation, with a device showing the relevant session, with timing that aligns, or it remains a lead. This is not a weakness in the signal. It is the method working as intended. Corroboration among independent sources is what turns a probable origin into a defensible attribution.

The chain moved. Its end did not.

The temptation is to say AI has broken attribution. That claim is imprecise, and imprecision is its own risk in this work.

The terminal gap, the distance between a proximal account or system identity and a human actor, was present in the traditional chain from the beginning. AI does not introduce it. What AI changes is the front of the chain. The watermark gives an orphan document something an email header could provide: a hint of origin. But it hints at the vendor, not the person. Retention and provider-side correlation are mechanisms that might carry that hint toward an account or system identity, and both are conditional. The person remains beyond that point, to be established through corroboration rather than inferred from proximity.

So the watermark is a trigger, not a locator. It tells you the record is worth pursuing. It does not tell you whom to charge. AI-origin signals may give investigators a new way to generate leads from artifacts that once carried almost no provenance information. What they do not provide is permission to skip the convergence that turns a lead into attribution. That permission never existed. The chain moved. Its evidentiary discipline did not.


Author's note. The reasoning in this article applies three working frameworks the author maintains and has deposited on Zenodo. The Artifact-to-Finding Promotion framework addresses when an observed artifact may be promoted to a finding, and what corroboration that promotion requires. The Human-Agent Attribution Discontinuity methodology addresses the break between an attributable system action and attributable human responsibility. The Zemi Method, Version 1.2 addresses how an artifact is classified, and how far attribution may travel before a finding is defensible. Readers who recognized the underlying discipline will find it stated formally there.

Sources