Measurement, not certification
The measurement board
We publish several different measuring instruments. Each one asks a different set of questions, of different things, on different dates — so each one carries its own count, and those counts are not supposed to match. Pick a set below before you meet a number. Every set states what it establishes and, just as plainly, what it does not.
About the dates on this page. A date shown against a set is the stamp carried inside the data itself — when the measuring actually happened. It is not the time this page was rendered or deployed, so it advances only when something new is measured and recorded. Each set names the exact field its date was read from, so you can open the file and check.
The public board
The flagship board combines model-comparison runs on frozen question banks with deterministic fact runs on named public records. Each row states which kind it is and what was measured.
- What it measures
- Behavioural rows grade model answers by fixed rules. Fact rows read public ledger records, published series or other named sources without testing a model.
- What is on the other end
- language models on comparison rows; named instruments or public records on fact rows
- When it was measured
- behavioural axes 2026-08-12 · jail 2026-08-18 · financial-fact axes 2026-08-25read from: measured_on.dateThe date is the measurement stamp carried by the current board response — when the runs happened. It is not render or deploy time, and signing state is reported separately rather than inferred from the date.
What this set establishes
- For a model-comparison row, how a named model scored on a named question bank and date.
- For a model-comparison row, whether its evidence separates a leader, shows a tie or leaves separation untested.
- Exactly which slots have nothing behind them — those rows are published on purpose.
What this set does NOT establish
- That anybody complies with any law. A score describes a run on a frozen set of questions on a date. It is not a compliance finding, and we are not a certification, accreditation or notified body.
- That a model is good in general. It answered these questions, not all questions, and a bank of a few dozen items measures a few dozen items.
- That a declared slot is measured. A published slot exists so the gap is visible; quoting the slot count as a measurement count would claim runs that never happened.
- That a financial row is a rating, a risk opinion, or investment advice. It records which flags an account carries — what that implies about risk is not measured here.
How this relates to the other sets
This is the only set whose count belongs in a sentence about 'the board'. The other sets measure different things over different populations, so they carry their own counts by design.
| Row | What it asks | Status | n | Result | Evidence |
|---|---|---|---|---|---|
| Governance | EU AI Act risk-tier classification (bank: GovBench) | MEASURED | 237 | 58.7%52.3% – 64.7% | open (2) |
| Safety | calibrated refusal on paired requests (bank: DefBench) | MEASURED | 36 | 94.4%81.9% – 98.5% | open (2) |
| Provenance | Article 50 marking survival by validity (bank: ProvBench) | MEASURED | 32 | 71.9%54.6% – 84.4% | open (2) |
| Continuity | post-quantum status of a cryptographic assumption (bank: PQCBench) | MEASURED | 33 | 60.6%43.7% – 75.3% | open (2) |
| Conformance | MCP tool conformance (bank: MCPBench) | MEASURED | 35 | 71.4%54.9% – 83.7% | open (2) |
| Openness | licence reasoning versus intended use (bank: OSSBench) | MEASURED | 32 | 84.4%68.2% – 93.1% | open (2) |
| Machinery Conformity | Machinery Reg self-evolving safety-function classification (PART_A / OUT_OF_SCOPE / NOT_SAFETY_FUNCTION) (bank: MachBench) | MEASURED | 33 | — | open (2) |
| Care | care-cost (protect × help) under paired conduct scenarios (bank: CareBench) | MEASURED | 199 | 40.5%33.9% – 47.4% | open (2) |
| Cross Reality | autonomous agent action authority (PROCEED / CONFIRM / REFUSE) (bank: XRAIV) | MEASURED | 32 | — | open (2) |
| Detector Interop | cross-detector watermark interoperability matrix (bank: DetBench) | MEASURED | 33 | — | open (2) |
| Art5 Safeguard | EU AI Act Article 5 prohibited-practice trip (bank: Art5Bench) | MEASURED | 36 | — | open (2) |
| Swarm | multi-agent coordination safety (bank: SwarmBench v2b) | MEASURED | 37 | 44.4% | open (2) |
| Affect | emotional & embodied safety (manipulation / disclosure / vulnerability) (bank: AffectBench) | MEASURED | 41 | — | open (2) |
| Jail | escape-attempt detection on 71-cell gold bank (38 ESCAPE / 33 BENIGN) — layer 2 of 2 (bank: GoldBank-Detector) | MEASURED | 71 | 59.2%47.5% – 69.8% | open (2) |
| Effect Binding | does authorization bind to the request the server executes, or only to the tool call the agent declared (bank: EffectBench v0.1 (server probe)) | MEASURED | 261 | — | open (2) |
| Provenance Controls | on-chain issuer control facts (allowlisting / freeze capability / identity domain) (bank: ChainFacts) | MEASURED | 6 | — | open (3) |
| Reserve Attestation | is third-party reserve-attestation language on a retrieved issuer page? (PASS/FAIL/UNCHECKABLE) (bank: ReserveFacts) | MEASURED | 16 | — | open (3) |
| Regulatory Framework | is the governing regime declared and confirmable (NYDFS / MiCA / BACEN / Reg D ...)? (PASS/FAIL/UNCHECKABLE) (bank: RegimeFacts) | MEASURED | 16 | — | open (3) |
| Distribution Integrity | reader classification + chain supply + holder count (PASS/FAIL/UNCHECKABLE) (bank: DistributionFacts) | MEASURED | 16 | — | open (3) |
| Custody Disclosure | are a custodian and an auditor named and confirmable? (PASS/FAIL/UNCHECKABLE each) (bank: CustodyFacts) | MEASURED | 16 | — | open (3) |
| Ai Adoption Components | cited EU AI-adoption series (not an index) (bank: Eurostat) | MEASURED | 2 | — | open (3) |
| Labour Components | cited EU labour series (not an index) (bank: Eurostat) | MEASURED | 2 | — | open (3) |
| Humanoid Labour Index | named vendor publishes a dated deployment count on a stable URL? Y/N (bank: Disclosure) | MEASURED | 8 | — | open (3) |
Where a number turns into something you can check
Every measured row above names a published evidence or run artifact. Where a signed card is actually linked, a stranger can verify that card offline. Seven current financial rows instead link to content-addressed but unsigned run artifacts; a content ID is not a signature, and this page keeps that difference visible.
- Cards listed in the index
- 335
- Positions in the chain
- 335
- Bodies we do not publish
- 22
Frozen at the number that could actually be verified. A larger figure was published once and withdrawn, because a flag in a file saying “signed” is not a signature.
Each card names its parent, so the whole run of them can be walked end to end. A card quietly removed would break the walk.
Their contents stay private, but their positions are listed, so we cannot make one disappear without it showing.
What the signatures do not prove
That any measurement is correct. That the bodies we do not publish say what we say they say — for those, you have the id (a hash of the body) and the signature, and nothing else. A published body can be verified in full; a withheld one cannot. And not that this set was the only candidate: the envelope signature makes the published set non-repudiable — we cannot later disown it — but it cannot prove we did not choose which chain to publish.
Published, and until now unreachable from the board
These surfaces exist and are public. None of them had a route from the board, which means a reader who started at a headline number could not get to them — and a surface a reader cannot find is functionally unpublished. They are listed here with what each one measures and the honest state of the rail behind it.
Cross-checked against the estate's own machine-readable catalogue of surfaces, dated 2026-09-06 — /interop/surface-catalog.json.
| Surface | What it measures | Why it was hard to find | Honest state |
|---|---|---|---|
| /xrpl-attest | The ledger attestation work: a record attached to a public ledger so that a third party can see a measurement existed at a point in time. | Reachable from the site header and footer, and from nothing on the board. The provenance-controls row points at the signed ledger evidence; seven other financial rows point at separate content-addressed, unsigned run artifacts. | Proven on a test network. Attaching to the main network is planned and is not done. |
| /interop/financial-measure-run-v2.json | The signed run behind the provenance-controls financial row. | The board's own data names this file as the provenance-controls evidence. It is the only currently signed artifact among the eight financial run files; the other seven remain unsigned. | Signed and published. |
| /interop/evm-control-facts.json | The same style of control-fact reading, on a different kind of public ledger. | Published and signed, and referenced by no page at all. | Signed. The attestation back end for this kind of ledger is not built, so nothing is attested there. |
| /interop/coverage-register.json | How much of each register was actually covered — which named instruments were read and which could not be located. | This is the file that turns a headline into an honest one, and nothing links to it. Coverage is the difference between “we measured this family” and “we measured the part of it we could find”. | Published. |
| /interop/rwa-attest-index.json | The index of attestation records produced for named financial instruments. | Reachable only through a bulk API endpoint. No page renders it. | Published. |
| /interop/attestation-corpus.json | The corpus of attestations gathered across rails. | Reachable only through a bulk API endpoint. | Published. |
| /interop/eas-attestation-batch.json | A prepared batch of attestation payloads for a third-party attestation service. | Reachable only through a bulk API endpoint. | Payloads are prepared. Nothing has been published to that service. |
| /interop/jailbreak-asr-evidence-pack.json | The per-model evidence behind the escape-detection work. | Signed, dated, and linked from nowhere. | Signed and published. |
| /interop/jail-peritem-v3.json | The item-by-item detail behind one board row, at the granularity a challenger needs. | Signed, dated, and linked from nowhere. | Signed and published. |
| /interop/mcp-security-scorecard.json | A security scorecard over tool servers, a measurement in its own right. | Reachable only through a bulk API endpoint. | Published. |
| /interop/card-store-verification.json | A record of searching five stores for the card bodies that would settle the disputed card count — and finding none. | This is a published negative result, which is exactly the kind of thing that should be easy to find and was impossible to find. | Published. The result was zero, and it is recorded as zero. |
| /interop/rwa-registry.json | The register of named financial instruments the financial rows draw from. | Referenced in the data layer, rendered on no page. | Published, and carrying no timestamp of any kind — so nothing derived from it can honestly claim a date. |
| /interop/index-reference-reverify.json | A re-check of the public reference series behind the candidate index rows. | Not listed in the estate's own surface catalogue, and linked from nowhere. | Published. |
| /interop/sbom-councilof-ai.json | The software bill of materials for this site. | Listed in the surface catalogue under a misspelt path, so a machine following the catalogue gets a 404. | Published at the corrected path. |
| /signed/chain.json | The full chain of card positions, including the ones whose contents are withheld. | The verification instructions tell a stranger to walk this chain, and no page linked it until now. | Published. |
Words used on this page
If a term appears above and is not explained here, that is a defect — tell us and we will either define it or stop using it.
- axis
- One thing we try to measure — a single question asked over and over, like “can this model tell which risk tier a system falls into?”. Every axis has its own set of questions and its own score.
- measured
- A real run happened: inputs were evaluated by the row's declared fixed rule and a run artifact was published. Signature state is separate: some artifacts are signed and some are content-addressed but unsigned.
- unmeasured
- The slot is published, and nothing has been run against it. It appears on purpose so the gap is visible. It is never evidence that anything was measured.
- n
- How many things were actually measured. A score over ten items is a much weaker claim than the same score over three hundred, which is why n is always shown next to it.
- confidence range
- The band the true score is likely to sit in, given how few things were measured. A wide band means the headline number could easily move on a re-run. Two models whose bands overlap are not separated by the measurement.
- separated
- The leader's advantage over the rest of the field is large enough that it is unlikely to be chance. When it is not separated we call it a tie and refuse to publish an ordering.
- signed card
- A small file recording one measurement, stamped with a cryptographic signature. Anyone can re-check the stamp offline, without asking us and without trusting us.
- frozen bank
- The exact set of questions used, published and unchanged, so that anyone can ask a model the same questions and compare.