FrançaisEnglishEspañolالعربيةहिन्दीবাংলা中文

← Back to the blog

Published on 2026-07-25

Do a watermark or a 'verified' badge actually prove where a photo came from?


Content labels can be stripped by a screenshot, and researchers have already shown invisible watermarks fail against modern editing tools. A badge is a clue, not a guarantee.

It sounds like the tidy solution to all of this: bake a signature into every image at the moment it's made, real or AI, so that anyone downstream can check it and know for certain. That system exists. It's called C2PA, short for the Coalition for Content Provenance and Authenticity, and it's backed by companies including Adobe, Microsoft and camera makers who've agreed on a shared standard for what they call Content Credentials. The idea is genuinely useful. It is also, right now, far more fragile than the label suggests.

What the label is built to survive, and what it isn't

A Content Credential is a small signed manifest attached to a file: who or what made it, when, with which tools, whether it was edited afterward. The Coalition's own FAQ is careful about what this manifest actually promises. It's a record of origin and editing history, cryptographically signed so that tampering with the record itself is detectable. It was never designed to be a truth stamp on the image's content. A credential can tell you a photo was made by a camera and later cropped in an editor. It cannot tell you whether the scene in front of that camera was staged.

The bigger practical problem is survival. That manifest travels attached to the file, and plenty of ordinary things detach it: a screenshot, a re-upload to a platform that recompresses images, a simple format conversion. The credential doesn't get proven false in these cases. It just isn't there anymore by the time the image reaches you, three shares and one screenshot removed from wherever it started.

Picture the ordinary path a photo actually takes: someone takes it, posts it on one platform, a second account screenshots that post and reposts it, a messaging app compresses it further on the way to your phone, and a friend forwards it to you from there. Every one of those steps is a completely normal, non-malicious thing people do dozens of times a day, and every one of them is also a place where an attached credential can quietly fall off. By the time an image is popular enough for you to be asking whether it's trustworthy, it has usually travelled through several of those steps already.

The backup plan has its own limits

The people behind the standard know this, which is why they've built a second layer: an invisible watermark meant to survive even after the visible metadata is gone, letting a checker rediscover the credential later. It's a reasonable answer to the screenshot problem. It's a weaker answer to a different one. A 2026 research paper testing watermark robustness found that several established watermarking methods, built specifically to survive resizing, compression and everyday edits, broke down against a newer category of threat: diffusion-based editing tools that can quietly alter an image while leaving the watermark degraded or gone entirely. The watermarks held up fine against the attacks they were designed for. They weren't designed for what came next.

None of this is a story about one bad piece of engineering. It's the ordinary shape of any security measure built against a moving target: as soon as a watermarking method becomes widely used, removing it becomes worth someone's time to figure out, and eventually someone does. That doesn't make the effort pointless. It makes 'this image has a badge' a weaker claim than a badge tends to sound like, and it means the badge you're relying on today is being tested by tools that didn't exist when it was designed.

A badge tells you a system tried to remember. It doesn't tell you the system succeeded.

So what does that leave you with

Treat a Content Credential, when you find one, as a genuinely good sign: whoever published this cared enough about provenance to attach a verifiable record, and that record hasn't been tampered with. That's worth something. But treat its absence with more nuance than 'no badge, must be fake', because the badge falls off constantly for entirely innocent reasons, a screenshot, an old photo from before the standard existed, a platform that strips it on upload. And treat 'it has a watermark' the same way you'd treat a lock on a door: real deterrent, not proof against everyone who might want in. The label is one input. Where the image has actually been seen before, and who else is showing the same scene, still carry more weight than any badge sitting inside the file itself, and they don't disappear the moment someone takes a screenshot.

Topics : AI images watermarks verification

Sources

Keyboard shortcuts

?Open this help
EscClose the panel or this help
1-3Pick an answer during the quiz
EnterNext case, once you've read the answer
TabMove from link to link: the whole site works by keyboard