Send email Copy Email Address
2026-10-09
Felix Koltermann

Do Watermarks Survive the Training of New AI Models? A CISPA Study Provides Answers

Watermarks are increasingly being used to make AI-generated images recognizable and to ensure their origin can be traced. Previous research has focused on whether watermarks can withstand image manipulation. CISPA researcher Michel Meintz from the SprintML Lab has investigated whether watermarks can withstand the training of a new generative model. The result: Not all watermarks are equally robust. Their ability to persist across multiple model generations depends heavily on their design and the model that is used. The paper “Watermark Degradation Across Model Iterations” was presented at the ACM Workshop on Information Hiding and Multimedia Security (IH&MMSec 26) in Florence.

AI-generated images are ubiquitous today and are widely distributed, especially via the internet. This increases the risk that AI-generated images will be used to train new image generation models. “When companies train on their own synthetic data, it can lead to model collapse,” explains Michel Meintz. “The more you train on your own synthetic data, the more likely the model’s quality is to deteriorate.” One way to detect synthetic data and, if necessary, filter it out of training data is through watermarks. “They embed imperceptible signals in AI-generated images that enable the detection of generated content,” explains the researcher. Until now, however, it was unclear whether watermarks could still be detected even after a new generative model had been trained. For this reason, Meintz decided to investigate this issue. “We took a closer look at whether watermarks are even suitable for proving, after model training, that the model was trained on the marked data,” says Meintz.

A Test of Watermark Durability

The CISPA researcher took a two-step approach: “First, we used a model to generate images containing watermarks,” explains Meintz. “We then used these AI-generated images as a dataset to train a second model. Our focus was on the training process of this second model. After training, we use the second model to generate new data and check whether the watermark is still detectable or not.” Specifically, Meintz and his colleagues examined four watermarking methods—TrustMark, StableSignature, TreeRing, and BitMark—and two different model architectures: diffusion models and autoregressive image models. The researchers wanted to determine whether the origin of a dataset could still be traced even after model training. The watermark serves as a signal that indicates whether a model was trained on the watermarked data. To do this, the researchers did not examine individual images but instead performed a joint statistical analysis of the watermark signals from various image datasets.

Watermarks Vary in Robustness: The Results

The key finding of the study is that watermarks persist to varying degrees across multiple training and generation cycles. While some methods lose their signal after just one training cycle, others remain detectable for significantly longer. BitMark, developed by SprintML, performed particularly well. In the Infinity-2B model examined, the watermark was already reliably detectable when only 1 percent of the training data was marked. “The transferability of watermarks also depends heavily on the model structure,” adds Meintz. “Depending on which generation method and model architecture are used, watermarks behave very differently.” Meintz concludes that watermarks must be adapted to the respective models. “For a watermark to survive a diffusion process or model training and still be learnable or detectable afterward, it must be more robust. It may require adjustments that aren’t necessary for other models,” says the researcher.

Challenges in Practice 

In addition to technical durability, there is the practical question of how synthetic data can be detected when different providers use different watermarks. “If companies use different watermarks, one provider cannot easily determine whether an image was synthetically generated,” explains Meintz. “Companies can then only detect their own watermarks and might overlook an AI-generated image if it bears another provider’s watermark.” A common watermark that could be recognized and verified by different providers could simplify the filtering of synthetic data in the future. “In current practice, however, reaching an agreement on a common watermark seems rather difficult because companies have different requirements and interests, and watermark structures are likely to continue changing as technology evolves,” explains Meintz.

Goal of Further Research: A Robust Watermark

The CISPA researcher is fascinated by the topic of watermarks. The results of his study motivate him to stay on the ball. “My goal is to investigate how watermarks themselves can be better designed and whether it’s possible to develop a watermark that is more robust,” he says. In addition, he wants to shift the focus from the dataset to the individual image: “Our current process requires a dataset to determine whether a watermark is present. Ideally, it would be possible to detect from a single image whether it was generated by a model trained on our watermarked data. To achieve this, the watermark signal would need to be significantly stronger.” To do so, he intends to build further on BitMark. “We would further develop the ideas—or rather, the principle—behind BitMark because it already works well. So far, however, it has been developed specifically for Infinity.” He is eager to take on the challenge of extending it to other architectures and thus making it more widely applicable.