ProvBench · EU AI Act Article 50 / C2PA manifest survival · signed run of 13 August 2026
Does provenance survive the real world?
We marked assets with C2PA manifests (Ed25519, embedded + sidecar), applied real-world transforms (JPEG re-encode, crop, resize, metadata strip, format convert, screenshot), and checked whether the marking still validated. The measured result: 0 of 12 marked assets kept an intact embedded C2PA manifest (0 of 108 measured cells). A marking present but whose binding no longer validates is scored DESTROYED, not SURVIVES. Every per-cell outcome is in the signed result file.
The headline
One number, one scoring rule: present-but-invalid does not count as survival.
Across 12 marked assets and 9 measured transforms
0 of 12 assets survived
0 of 108 measured cells; clustered 95% interval 0 to 24.25%, computed at n=12 assets; one-sided 95% Clopper–Pearson upper bound 22.1%
A marking is scored SURVIVED only when its binding still validates against the asset it is attached to. Present markers with a broken binding are scored DESTROYED. The identity control passes; ordinary saves do not.
artefact: /packs/eu-article-50/provbench.json
Embedded binding does not survive an ordinary save
A C2PA manifest baked into the asset bytes — what Article 50 imagines — does not survive JPEG re-encode, resize, crop, metadata strip or a format change. The binding is a cryptographic hash over the bytes; change the bytes and the binding no longer validates. Scored DESTROYED, not SURVIVES.
A sidecar recovers the disclosure, never the binding
A detached manifest handed straight to the verifier preserves the disclosure markers (manifest present, signature valid) but cannot re-bind to bytes a transform has already changed. A manifest lifted from a completely different asset still reports signature_valid — only binding_intact catches the transplant.
Present-but-invalid is not survival
The scoring rule behind the headline: a marking counts as SURVIVED only if its binding still validates against the asset it is attached to. Present markers with a broken binding are scored DESTROYED. Under that rule, 0 of 12 marked assets kept an intact embedded C2PA manifest (0 of 108 measured cells).
The transforms
Real-world transforms applied to each signed asset. Every cell measured, never assumed.
The statistics
The scoring rule
A marking is SURVIVED only if its binding still validates against the asset it is attached to. Present-but-invalid markings are scored DESTROYED. Under that rule, in the signed run of 13 August 2026: 0 of 12 marked assets kept an intact embedded C2PA manifest (0 of 108 measured cells); clustered 95% interval 0 to 24.25%, computed at n=12 assets; one-sided 95% Clopper–Pearson upper bound 22.1%.
Clustered, not per-cell
Cells cluster by asset and transform — they are not independent trials. Nine transforms of the same signed asset are one deterministic fact restated nine times, not nine observations. Intervals are computed at the asset level.
Pre-registered predictions
Predictions were locked before first measurement. Embedded binding: predicted 'destroyed' — confirmed. Sidecar disclosure: predicted to recover the marker but not the binding — confirmed.
Reproducibility
A hard binding is a cryptographic hash over asset bytes. Re-running any cell reproduces the identical outcome with probability 1 — run-to-run uncertainty is structurally zero. The residual is generalisation to unseen assets, not sampling noise.
The harness
Open, deterministic, reproducible. C2PA marking, Ed25519 signing, SHA-256 hash-chained (OpenTimestamps anchoring is roadmap, not yet wired). The harness lives at csoai-static-deploy2/provbench.py. Same code, same seed, same outcome on every run — the binding is a cryptographic hash over asset bytes, so re-running any cell reproduces the identical result with probability 1.
The audio wing (ProvBench-Audio) measures open anti-spoofing detectors against modern TTS synthesis — running on Kaggle T4, same discipline.
Earlier runs, not the headline
- Earlier unsigned 20-asset run, 29–30 July 2026. 20 marked assets, the same C2PA embedded and sidecar configurations, 180 measured cells: 0 of 20 assets survived (0 of 180 measured cells); one-sided 95% Clopper–Pearson upper bound 13.9%. Status: unsigned; its result file is not published on councilof.ai.
- Earlier unsigned preprint run, 30 July 2026. 15 assets × 7 transforms = 105 cells, mixed binding types including a soft watermark: the preprint reports 18 of 105 cells surviving (17.14%). Status: unsigned; its per-cell data is not published on councilof.ai, so the figure cannot be re-derived here.
Until 15 September 2026 this page headlined the preprint figure and described it as the durability of a watermark. The signed run tested no watermark.
What this benchmark does not claim
- A surviving marking proves PROVENANCE, NOT CORRECTNESS. It states that these bytes carry a claim signed by this key with this declared history. It says nothing about whether the content is accurate, safe, or lawful.
- Our certificate chains to a PRIVATE ROOT CA not on the C2PA trust list. issuer_resolvable therefore fails by construction everywhere — a property of the credential, not damage from a transform. The survival figure measures binding integrity, not trust-list membership.
- A verifier reporting 'signature valid' without reporting the binding is telling you almost nothing: a manifest transplanted from another asset still reports its signature as valid. Only binding_intact catches it.
- Transforms are applied by Pillow. A different re-encoder (libvips, ImageMagick, a phone ISP, a CDN) may behave differently. This measures the common case, not every case.
- Every figure here is recomputable from the published artefact. Where an encoder is unavailable in the harness (HEIC) the cell is reported UNMEASURED, never scored as a pass or a fail.