CoverLock

Neural Image Watermarking · 2026

Residual Transferability
in Neural Image Watermarking

Why can watermark residuals be transplanted to unrelated images—and how can we bind them back to image content?

Ziping Dong1 · Qi Li1 · Xinchao Wang1

1National University of Singapore

TL;DR

Residual-transfer forgery moves watermark evidence from released images to unrelated content. We show that architecture determines whether this evidence depends on the cover image, and introduce CoverLock to protect existing watermarking systems without architectural redesign.

01 · Attack mechanism

A watermark residual can leave its image behind.

Let a released watermarked image be y = x + r. A residual-based forger estimates r and adds it to an unrelated target image x′. If the decoder recognizes the residual without checking the carrier content, x′ + r inherits the source watermark.

1
y = x + r

Released image

A valid watermarked image exposes both its visible content and a hidden message-bearing residual.

estimater̂ = y − x̂→
2
r̂

Forgery residual

The attacker isolates reusable watermark evidence; no access to the original encoder is required.

transfery′ = x′ + r̂→
3
✓ decoded

False attribution

An unrelated image is accepted as carrying the source identity.

Residual transfer attack. The attack succeeds when watermark evidence is portable across carrier images.

02 · Why it happens

The model may take a shortcut.

A watermark encoder is asked to preserve a message, but is not necessarily required to make that message depend on the image. The easiest solution can be a consistent, cover-independent residual: different images receive similar patterns for the same payload.

A

Same message, different images
The learned residuals remain strongly aligned.

B

Average away the image content
Multi-image estimation exposes the shared component.

C

Transfer the shared component
The decoder accepts it on a new carrier because it never learned a strong content dependency.

Comparison of residual consistency and decoder feature variation across high- and low-transferability watermark models
Residual consistency reveals the shortcut. High-RT models produce nearly identical patterns across covers, and their decoder features are dominated by the message rather than the cover. Low-RT models show the opposite behavior.
03 · What determines RT?

Not the training recipe. The architecture.

We first control for common training-side explanations. Changing the recipe does not reproduce the large gap in residual transferability, directing attention to how image and message information are combined inside the network.

ControlledTraining-side configurations
Training datanot predictive of RT
Distortion layernot predictive of RT
Loss configurationnot predictive of RT
Training progressnot predictive of RT
≠Primary cause
of the RT gap
Architectural intervention

Broadcasting the message changes residual transferability.

Adding broadcast fusion to MBRS sharply lowers RT, while removing it from HiDDeN raises RT. What changes is whether message information is forced to interact spatially with the image.

Left · RT@100With broadcast, HiDDeN, MBRS, and CIN drop from near-perfect transferability to 0.21, 0.10, and 0.09.

Right · Training BitAccBroadcast slows learning, but the models still converge to high decoding accuracy as they see more images. The RT reduction is therefore not a failure to learn the watermark.

Paper analysis comparing residual transferability with and without message broadcasting, alongside training bit accuracy
Controlled architectural intervention. Left: broadcast sharply reduces RT. Right: broadcast models learn more slowly but still reach high training BitAcc, ruling out under-training as the explanation.
Conclusion

Content dependence is induced by architecture. Preserving local image–message interaction makes the learned residual less reusable across unrelated covers.

04 · Protecting existing models

CoverLock binds the payload to the carrier.

Redesigning an encoder helps future models, but deployed watermark systems already exist. CoverLock is a plug-in wrapper: it derives a stable image-conditioned bit mask and XORs it with the user message before watermark encoding.

Embeddingp = m ⊕ b(x)
Detectionm̂ = D(ỹ) ⊕ b(ỹ)
Original paper figure showing CoverLock training, code regularization, and plug-in watermark encoding and decoding
CoverLock overview from the paper. A frozen DINO backbone and lightweight trainable adaptor learn a distortion-stable, diverse content code. At inference time, the pretrained watermark encoder and decoder remain frozen and unmodified.
Watermark-agnostictrained without a watermark codec
Plug-and-playworks with existing encoders and decoders
Content-boundthe same message maps to carrier-specific payloads
05 · Results

A better security–robustness frontier.

CoverLock does not obtain security by simply rejecting more distorted images. It achieves a substantially more favorable trade-off: residual-transfer attacks are suppressed while genuine watermarks remain detectable after common image distortions.

Security and robustness Pareto frontier comparing CoverLock with Vanilla, ConvNeXt, and MHDW binding defenses
Security–robustness trade-off on COCO. Security is measured by GT@1 attack success rate (ASR; lower is better), and robustness by true-positive rate under distortions (TPR; higher is better). View the vector PDF.

Pareto frontier

More secure and more robust.

Points toward the upper left are better: lower forgery success and higher detection of genuine watermarks after distortion. All CoverLock variants occupy this favorable region.

Security ↓
0.111% GT@1 ASR
Robustness ↑
92.50% distortion TPR

CoverLock-G versus MHDW: ASR decreases from 11.367% to 0.111%, while robustness TPR increases from 86.38% to 92.50%.

Citation

@misc{dong2026residual,
  title         = {Residual Transferability in Neural Image Watermarking},
  author        = {Dong, Ziping and Li, Qi and Wang, Xinchao},
  year          = {2026},
  eprint        = {2609.32241},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CR}
}