ProvBench · EU AI Act Article 50 / C2PA manifest survival · signed run of 13 August 2026

Does provenance survive the real world?

We marked assets with C2PA manifests (Ed25519, embedded + sidecar), applied real-world transforms (JPEG re-encode, crop, resize, metadata strip, format convert, screenshot), and checked whether the marking still validated. The measured result: 0 of 12 marked assets kept an intact embedded C2PA manifest (0 of 108 measured cells). A marking present but whose binding no longer validates is scored DESTROYED, not SURVIVES. Every per-cell outcome is in the signed result file.

The headline

One number, one scoring rule: present-but-invalid does not count as survival.

Across 12 marked assets and 9 measured transforms

0 of 12 assets survived

0 of 108 measured cells; clustered 95% interval 0 to 24.25%, computed at n=12 assets; one-sided 95% Clopper–Pearson upper bound 22.1%

A marking is scored SURVIVED only when its binding still validates against the asset it is attached to. Present markers with a broken binding are scored DESTROYED. The identity control passes; ordinary saves do not.

artefact: /packs/eu-article-50/provbench.json

Embedded binding does not survive an ordinary save

A C2PA manifest baked into the asset bytes — what Article 50 imagines — does not survive JPEG re-encode, resize, crop, metadata strip or a format change. The binding is a cryptographic hash over the bytes; change the bytes and the binding no longer validates. Scored DESTROYED, not SURVIVES.

A sidecar recovers the disclosure, never the binding

A detached manifest handed straight to the verifier preserves the disclosure markers (manifest present, signature valid) but cannot re-bind to bytes a transform has already changed. A manifest lifted from a completely different asset still reports signature_valid — only binding_intact catches the transplant.

Present-but-invalid is not survival

The scoring rule behind the headline: a marking counts as SURVIVED only if its binding still validates against the asset it is attached to. Present markers with a broken binding are scored DESTROYED. Under that rule, 0 of 12 marked assets kept an intact embedded C2PA manifest (0 of 108 measured cells).

The transforms

Real-world transforms applied to each signed asset. Every cell measured, never assumed.

Identity (control)JPEG re-encode q90JPEG re-encode q70JPEG re-encode q50Resize 50%Crop 10%Strip metadataScreenshot equiv.Convert → PNGConvert → WebPConvert → HEIC

The statistics

The scoring rule

A marking is SURVIVED only if its binding still validates against the asset it is attached to. Present-but-invalid markings are scored DESTROYED. Under that rule, in the signed run of 13 August 2026: 0 of 12 marked assets kept an intact embedded C2PA manifest (0 of 108 measured cells); clustered 95% interval 0 to 24.25%, computed at n=12 assets; one-sided 95% Clopper–Pearson upper bound 22.1%.

Clustered, not per-cell

Cells cluster by asset and transform — they are not independent trials. Nine transforms of the same signed asset are one deterministic fact restated nine times, not nine observations. Intervals are computed at the asset level.

Pre-registered predictions

Predictions were locked before first measurement. Embedded binding: predicted 'destroyed' — confirmed. Sidecar disclosure: predicted to recover the marker but not the binding — confirmed.

Reproducibility

A hard binding is a cryptographic hash over asset bytes. Re-running any cell reproduces the identical outcome with probability 1 — run-to-run uncertainty is structurally zero. The residual is generalisation to unseen assets, not sampling noise.

The harness

Open, deterministic, reproducible. C2PA marking, Ed25519 signing, SHA-256 hash-chained (OpenTimestamps anchoring is roadmap, not yet wired). The harness lives at csoai-static-deploy2/provbench.py. Same code, same seed, same outcome on every run — the binding is a cryptographic hash over asset bytes, so re-running any cell reproduces the identical result with probability 1.

The audio wing (ProvBench-Audio) measures open anti-spoofing detectors against modern TTS synthesis — running on Kaggle T4, same discipline.

Earlier runs, not the headline

  • Earlier unsigned 20-asset run, 29–30 July 2026. 20 marked assets, the same C2PA embedded and sidecar configurations, 180 measured cells: 0 of 20 assets survived (0 of 180 measured cells); one-sided 95% Clopper–Pearson upper bound 13.9%. Status: unsigned; its result file is not published on councilof.ai.
  • Earlier unsigned preprint run, 30 July 2026. 15 assets × 7 transforms = 105 cells, mixed binding types including a soft watermark: the preprint reports 18 of 105 cells surviving (17.14%). Status: unsigned; its per-cell data is not published on councilof.ai, so the figure cannot be re-derived here.

Until 15 September 2026 this page headlined the preprint figure and described it as the durability of a watermark. The signed run tested no watermark.

What this benchmark does not claim

  • A surviving marking proves PROVENANCE, NOT CORRECTNESS. It states that these bytes carry a claim signed by this key with this declared history. It says nothing about whether the content is accurate, safe, or lawful.
  • Our certificate chains to a PRIVATE ROOT CA not on the C2PA trust list. issuer_resolvable therefore fails by construction everywhere — a property of the credential, not damage from a transform. The survival figure measures binding integrity, not trust-list membership.
  • A verifier reporting 'signature valid' without reporting the binding is telling you almost nothing: a manifest transplanted from another asset still reports its signature as valid. Only binding_intact catches it.
  • Transforms are applied by Pillow. A different re-encoder (libvips, ImageMagick, a phone ISP, a CDN) may behave differently. This measures the common case, not every case.
  • Every figure here is recomputable from the published artefact. Where an encoder is unavailable in the harness (HEIC) the cell is reported UNMEASURED, never scored as a pass or a fail.