LTP 155: Provenance not AI Detectors
In this solo show Bart argues in favour digital provenance over AI detectors, and specifically, in favour of the Content Credentials standard from the C2PA. Rather than trying to detect AI fakes, we should learn from the antiques and collectibles world, and reverse our starting assumption and the burden of proof.
Introduction
It’s been some time since I dedicated an instalment to the thorny issue of Artificial Intelligence, but it obviously hasn’t gone away! (LTP 137 “Copilots not Replacements” in February 2025).
This episode was inspired by a pair of new AI laws which come into force in this month — Article 50 of the EU AI Act in Europe, and the California AI Transparency Act in the US. These laws both aim to address the same problem — transparency. On other words, helping people tell what’s real from what’s generated.
Article 50 of the EU AI Act mostly focuses on labelling, which is nice, but I find the California AI Transparency act is much more interesting because it mandates support for provenance metadata. I dedicated LTP 125 “Image Provenance with Content Credentials” to the concept, and explained how one such system, Content Credentials developed by the C2PA, or Coalition for Content Provenance and Authenticity.
The California law doesn’t explicitly mention the C2PA standard, but that’s how many are interpreting it. What the law requires is support for “widely accepted industry standards”, and as things stand, C2PA Content Credentials are the de facto standard. This is not surprising given some of the organisations that are members — Adobe, Canon, Nikon, Microsoft, Google, Meta, TikTok, the BBC, AP, Bloomberg, and even OpenAI!
As a reminder, here’s a very short summary of what Content Credentials are from my favourite privacy-protecting AI chatbot:
Content Credentials are “an open technical standard for attaching tamper-evident provenance metadata to digital media — essentially a cryptographically signed “birth certificate” for content that records who created it, what tools were used, and any subsequent edits, so that downstream viewers can verify whether a piece of content is authentic or has been manipulated.” — Lumo
The key point is that Content Credentials use cryptography to prove the provenance of a piece of digital media. While you can remove a content credential (like you can any piece of metadata), you can’t fake one. If an image, video, or audio file has a valid Content Credential embedded, then you can be confident that it really was generated in the manner described by the content credential. Note that Content Credentials simply capture the history of a piece of media, they’re not about detecting AI or proving an image is real, but about transparency. A Content Credential can verify that an image is entirely artificial, entirely free from any AI manipulation, or, a mixture of the two. What matters is that an image with a Content Credential is a known quantity.
What California are mandating is that, starting now (August 2026), large AI providers with over 1 million users need to embed provenance information into the content they generate, and provide free tools for reading provenance information from media files. Large online platforms that distribute media have to preserve any provenance information in media files uploaded by users. Large platforms are also encouraged to add support for displaying provenance information into their sites and apps, but unfortunately, that’s not a legal requirement, just a request or suggestion.
If this was all the law did it would be nice, but there’s more to come! Starting in January 2028 “capture devices”, including cameras, will need to give users the option to embed provenance information into media as it’s created.
There’s no guarantee all this is going to happen, because this is Californian law, not a national law, let alone an international one, but California has a long history of successfully driving change in large industries, so I’m optimistic.
If the law has it’s desired effect, then it will revolutionise our media landscape!
Reputable media sources will be able to cryptographically prove an image really was taken where and when they say, and has not been manipulated since!
AI Detection is a Fool’s Errand — We Should Learn from Antiques!
When you talk about restoring trust in the media, I find that most people immediately focus on the need for some kind of AI Detector — a tool that can reliably tell you if any image is real or generated.
As a computer scientist I’m pretty sure this is impossible. At best this will be another cat-and-mouse game where both the detectors and the generators are constantly improving, and neither will ever gain the upper hand. My favourite analogy is anti-virus software. Our AV tools have been evolving for decades, and yet, they still miss stuff, because the attackers have kept evolving too!
If AV is an example of the kind of approach I don’t think will work, is there an example of an approach that I think can work?
Yes — the problem of determining whether or not antiques and other collectibles are real.
There’s no magical test for determining of some random Fender Stratocaster guitar was in fact one played by Jimmy Hendrix. Instead, if you want to sell a guitar as having been played by the great, you need to provide an un-broken chain of custody all the way back to the exact concert the guitar was played at. This is referred to as the guitar’s provenance. No provenance, no credibility!
It’s About Starting Assumptions, and the Burden of Proof
Auction houses don’t assume an item is real and then try to prove it’s fake, they assume it’s fake and make you prove it’s real before they’ll stake their reputation by adding it to their catalog!
It’s not Christies’ responsibility to prove an item is fake, it’s the seller’s responsibility to prove it’s genuine!
This is where we’re getting things backwards with media — we assume all media is real unless we can prove it’s fake! This is of course what leads to the quest for a reliable AI detector. We may as well quest for a unicorn!
We need to flip our assumptions and the burden of proof. We need to assume that all journalistic images, videos, and recordings are untrustworthy, and make media organisations prove they’re genuine.
When some one shows me something supposedly shocking but true on social media, I really don’t think the burden of proof is on me! The burden of proof is on the new outlet, not anyone else!
In a world without digital provenance like Content Credentials, that’s hard to do. You end up falling back on general concepts like earned trust — this is from the BBC, you can see the original on their website here, and they have a long history of being trustworthy, therefore, you should trust this too.
A Content Credential Future
If the California law lives up to it’s promise, we’ll be in a very different media landscape by the end of the decade.
The cameras used by photojournalists will embed Content Credentials tying the images to a specific person, place, and time right as the image is captured. Every edit made to image between the shutter firing and the image being posted online will be made using editors that support Content Credentials, so every tweak will be captured in the image’s provenance metadata. Finally, social media sites and apps will expose the provenance information to us, using clear verification badges, making it possible for media to be reliably authenticated.
In this world, a sensational photo or video without a Content Credential will look immediately suspicious, and people will simply assume it’s some kind of fake, be that a traditional fake using Photoshop, or some kind of generative AI.
Side Note — Why Bother Adding Content Credentials to Generated Media?
As well as mandating camera support from 2028, the California law mandates that large generative AI companies embed provenance information into generated content. If our default assumption has to be that everything unknown is fake, why bother?
Explicitly and provably marking media as fully or partially generated minimises the need for assumptions. If a sensational video is provably generated, then there’s no need to argue about where the burden of proof lies or anything like that. It just makes things simpler!
Of course, there will still be lots and lots of media without provenance information, but the fewer unknowns, the better!
Why will there always be media without provenance? All sorts of reasons, but here are just some of the most obvious ones:
- Metadata can be stripped — remember, provenance impossible to fake, but easy to break!
- Only large platforms with over a million users are covered by this law, so smaller AI companies can continue to generate media without provenance metadata
- People can run their own models on their own devices, and do what ever they want!
- The baddies don’t care about laws, so there will always be an underworld maliciously generating fakes
So, generated images that pro-actively disclose their origins are nice to have — nothing more, but also, nothing less!
Final Thoughts
Only time will tell how much of this comes to pass, or how quickly, but I like the future these laws are trying to build, especially the Californian one.
These bills can’t make things any worse, so even if they don’t deliver on all their promise, there’s still value in encouraging the wider use of provenance metadata!
All in all though, I’m quietly optimistic that things are moving in the right direction 🤞
This episode is free for you to enjoy, free from ads, and free from sponsors, but it's not free to produce. This podcast is 100% listener supported, without the support of people like you, it would not exist. Please consider supporting our work.
More Options …