The long version

How this works, at the length it actually takes.

The front door gives you six checkable facts and the board. This page is everything underneath them: how a measurement is made and who pays for it, every machine address and data door we serve, the nine products, what we have got wrong, and the questions we cannot answer yet. Every figure is read live as this page loads — nothing here is typed into the page.

Back to the board · The library holds every page this estate has published

Why this one is different

Six things you can check before you believe anything else on this page.

None of these is an adjective. Each one names a thing we publish and gives you the door to go and look at it — and the small print beside each figure is the honest state of that source, not decoration.

01

Nobody we measure is paying us

We are an independent measurement body. The tests are frozen and published before a run, the grading is done by a rule rather than by another AI, and no vendor can buy a place on the board, lift a score or have one taken down. Re-checking any result costs nothing and always will.

No company we measure pays for a place, a score, or its removal — there is no rate card for any of it.

How a measurement is made
02

Our bad results are on the board too

The easy thing is to publish the wins. A measurement body is only worth reading if the weak scores are there beside them — including the ones our own models earned. They are, with the same signature and the same test behind them.

our own models are published at single-digit scores and at zero, signed exactly like every other result

See one of our own low scores, checked in your browser
03

Every claim we got wrong is written down

When something we published turns out to be wrong, we do not reword it. The correction goes in a public ledger that says what was wrong, how it was caught and what changed. Judge a measurement body on what it does on a bad day.

78 corrections published · the ledger's own signature is VALID, and we say so rather than re-sign it quietly

Read the corrections ledger
04

The history cannot be quietly rewritten

Published records are hashed into a tree whose root is signed. Change a covered record and the root stops matching. Separately, artifacts are submitted to OpenTimestamps. The public manifest distinguishes proofs carrying Bitcoin block-header attestations from calendar-pending submissions; independent verification must also check the block header and the named subject file.

310 records under one signed root c2ab6d59…2ea768e3, rebuilt 2026-09-28T07:32:07Z · separately, 639 of 643 published proof files contain Bitcoin block-header attestations; 4 remain calendar-pending. This manifest parses the proof bytes but does not verify block headers against a node or every named subject file

Open the signed root
05

We are in the rooms where this is decided

Measurement only matters if it plugs into the standards everyone else will use. We take part in the bodies writing them — and we say plainly that taking part is not endorsement, and a listing is not adoption.

33 participation records, 11 of them in standards bodies, each linked to its own evidence · as of 2026-09-27

Where we take part, with the evidence
06

It is a machine, not a report someone writes

The board re-measures, the evidence tree is rebuilt, the root is re-signed and the ledgers are re-read on a schedule, without anyone typing a number. That is why the figures on this page are read from the endpoints as you load it rather than written into it.

the root on this page was rebuilt and re-signed at 2026-09-28T07:32:07Z

Every door this machine serves

Re-checking any of this costs nothing, needs no account, and always will.

Not a trial, not a tier and not a limit we can quietly lower later. The tests, the grader, the records and the key are all published, so the check runs on your machine whether we are still here or not. What we charge for is work we do on request — never a score, never a place on the board, and never permission to look.

Data, and the addresses behind it

Ten populations everybody argues about, counted from named sources.

Each door takes one population, counts it from a published artifact, and says what the count is of, when it was read and what state it is in. The figures below are the doors' own free previews, read as this page loaded. They count ten different things and are never added together.

  • 425

    Stablecoin universe

    assets catalogued in the frozen universe (not supply, not a sum) · as of 2026-09-16T12:06:54Z · INDEXED

  • 26

    SWIFT-linked bank census

    sourced bank rows (three-state; identity = a dated press URL, hashed) · as of 2026-09-01 · INDEXED

  • 16

    XRPL issued-asset set

    xrpl.asset.state leaves on the public root (the live reader's n) · as of 2026-09-22T08:54:02Z · INDEXED

  • no total published

    x402 Bazaar listings

    listings per index frame (frames overlap; the artifact refuses a total, so n is null — not 0) · as of 2026-09-17T05:48:05Z · INDEXED

  • 354

    MCP registry listings

    own server names in the official MCP registry (one row per name, classified on isLatest=true) · as of 2026-09-22T08:27:50Z · INDEXED

  • 391

    A2A agent registry population

    distinct agents in the a2aregistry-org frame (one frame; not a global count) · as of 2026-09-17T05:48:05Z · INDEXED

  • 643

    OpenTimestamps proof manifest

    published .ots proofs that parse as proofs · as of 2026-09-27T10:37:43Z · MEASURED

  • 29

    Layer-0 machine-surface ceremony

    live probes of the estate's machine surface in the ceremony · as of 2026-09-03T05:10:26Z · MEASURED

  • 78

    Corrections ledger, full history

    correction entries in the source-maintained ledger · as of 2026-09-28 · INDEXED

  • 198

    Claim-maintenance registries

    claims captured across the published registries · as of 2026-09-25 · INDEXED

A preview is free and always will be. The rows themselves, and the signed reading over them, are the part we meter — a rule this estate keeps everywhere: reading what we measured costs nothing, and the work we do on request is what is paid for. Every metered door, with its indexing state.

Where the work has got to

The work travels, and it holds up when you check it.

Two different questions, and mixing them is how every adoption number you have ever read became meaningless. How far the packages and listings have reached is one measurement. Whether the evidence itself survives a re-check is another, and it is the one that matters.

How far it reaches

  • Downloads—
  • Registry listings—

Counted package by package, on the date shown. Downloads are not users. Every figure, its source and its as-of.

That is the short view. The whole funnel — all seven stages, including the four we cannot measure and what each of them would need.

And when you check the evidence itself

335

signed records re-checked and still valid

kind: measured — a check was actually run over these bodies

305

records sealed under the one signed root

kind: catalogued — counted on the root, a different set from the one above

0

identifiers the two sets have in common

they are separate populations; adding them would be a number about nothing

These are separate sets of separate records. Say which one you mean, every time — and never add them. How every one of these counts was derived.

Distribution, with its limits

The published tools travel. Downloads are only the first signal.

This census covers packages in the wider CSOAI and MEOK publishing estate across PyPI, npm and Hugging Face. It counts registry download events, including mirrors and automated traffic. It does not count unique people, active installations, executions or customers.

≥ 2,655,891

gross download events since first release

836 of 837 package counters answered

≥ 386,320

gross download events in registry 30-day windows

836 of 837 package counters answered

Published census dated 24 September 2026 UTC · Out of date — a fresh census is needed. The PyPI source and a separate sample counter disagree, so these are reported registry events, not verified adoption.

Of the cumulative total: MEOK 2,327,885 · CSOAI 174,512 · joint 46,231 · unattributed 107,263. Package ownership labels come from the published census.

Nine products

Nine doors. Each one opens today.

Independent measurement body. We run AI systems against frozen published tests, sign the result, and leave empty cells empty. Nine doors, each a real page.

Why these nine, and not a catalogue

  • Each tile opens a page that exists today. A tool with no destination is not on this band.
  • Empty cells stay empty. We do not invent a figure to fill a gap.
  • only hereWe measure. We do not sell a rank, a certificate, or a placement.
Clay people and pale humanoids facing each other across an arena under beams of light
GSPC · the living board

The living board

Every slot we publish about how AI systems behave, with the measurement behind it — and a visibly empty cell wherever there is no measurement.

Otherwise you compare suppliers on scorecards that quietly leave out the tests they did badly on.

  • A filled cell is a measurement. A dash is honest emptiness.
  • The count line is printed once on this page, on the living board above — never typed here.
  • A TIE stays a TIE. It is never dressed up as a win.
Open this tool
Clay figures pointing at a card reading “3KB credential” in front of an open vault
Your own system

Get measured

We run your system against the frozen, published tests that apply to it and hand you a small signed record you keep — the scores, the sample size behind each one, and the slots we could not fill.

Otherwise you hand a buyer a policy document where they asked for evidence.

  • Frozen, published tests — the target does not move after you sit.
  • You keep the signed card. Publishing it is your decision.
  • Slots we could not fill stay empty and are named.

Scoped runs are arranged by enquiry; card verification stays free.

Enquire about a run
A block of carved statute breaking apart into a branching tree of true/false conditions
EU AI Act · GPAI

GPAI evidence pack

Builds the evidence index for one general-purpose AI system: the live rows that exist, the published banks they resolve to, and the gaps, named rather than skipped.

GPAI duties have been in force since 2 August 2025, and most providers have only their own paperwork to show for them.

  • Live rows that exist, the banks they resolve to, and the gaps.
  • Gaps are named rather than skipped.
  • Independent evidence — not a conformity mark, and not legal advice.

Independent evidence. Not a conformity mark, and not legal advice.

Open this tool
A pale moulded form lifting out of a split block of clay on a beam of green light
For your own site

Embed and white-label kit

Builds a badge or card you can paste into your own site that re-checks its own signature in each reader's browser — it goes green only when the signature verifies for that record. That is an integrity check under the published key, not a truth, safety or conformity verdict.

Otherwise the people reading your site still have to take your word for the result.

  • Green means one named check passed: the signature verifies for this record.
  • Each reader's browser re-checks the signature.
  • Built only from what is actually on the board.

Built only from what is actually on the board. Free forever.

Open this tool
An hourglass weighing a stale seal against a re-attested current one, fed by EUR-Lex and legislation.gov.uk ribbons
Underwriting

Insurance evidence rail

The measured rows, the honestly empty ones, and third-party reported figures — kept in three separate columns and never blended into a single number an underwriter could mistake for a rating.

Otherwise AI exposure is priced off a questionnaire the applicant filled in about itself, and nothing updates between binding and renewal.

  • Measured, empty and reported figures stay in three separate columns.
  • Nothing is blended into a single number an underwriter could mistake for a rating.
  • We measure. We do not price risk.

We measure. We do not price risk, and we take no share of anything written on the back of a card.

Open this tool
Specimens sealed in glass tubes, turning from grey clay to a lit green core
Financial and legacy systems

Specialist registers

A separate board for money and mainframes: whether a COBOL copybook off a bond desk can be turned into an attestable record, whether an underwriting rule reads as covered or excluded — one row per instrument, each with its own item count.

Otherwise the systems that actually run a bond desk or a claims book sit outside every AI measurement anybody publishes.

  • One row per instrument, each with its own item count.
  • COBOL copybook and underwriting-rule rows sit beside the public board.
  • A specialist register is still measurement — never a certificate.

bond desk · COBOL copybook → attestation · MEASURED on 12 graded itemsread from the published register rows

Open this tool
A raw jagged signal trace behind a glass panel labelled “unstructured outcry”
Open to everyone

Watchdog evidence

Read the public Watchdog material and current evidence state. Durable incident submission, storage, and signed acknowledgements are not implemented in this release.

Otherwise a harm disappears into a supplier's private support queue and nobody outside it ever learns it happened.

  • Public Watchdog material remains readable without an account.
  • No report is represented as filed unless a durable intake confirms persistence.
  • Any future finding must pass the same measurement and evidence gates as the board.

Read-only today. The report intake is explicitly unavailable rather than pretending to file.

Open this tool
A pale sphere held inside thin orbital rings studded with green markers
How we are funded

Payment never buys a rank and verification is free forever.

The obvious question about any body that scores AI is: who is writing the cheque? Here is the whole answer. No company we measure pays for its place, its score, or its removal. Members of the public never pay anything at all. Verifying a card is free forever, with no account. We fund ourselves by selling signed evidence artefacts — the report, the dataset, the re-attestation — published win or lose, and never a fee for a ranking or a placement.

  • painMost AI ratings are paid for by the company being rated
  • painYou are asked to trust a score you cannot see the invoice behind
  • benefitVerification is free forever — no login, no fee, no tier
  • benefitA bad result is published exactly like a good one
  • only hereWe take no money from anything we rank — the board is not for sale
The boundary

We measure. We do not certify — the boundary is the point.

The limits are the brand. We are a measurement body and nothing else, and saying so plainly is more useful to you than any badge would be. Read the four lines below as hard exclusions, not modesty.

  • Not certification

    We issue no certificate and no conformity mark.

  • Not accreditation

    There is no accreditation chain behind us, and we are not a notified body.

  • Not enforcement

    We cannot approve, ban, fine or clear anything. Regulators do that.

  • Not legal advice

    A score describes a measured run on a date. It is not a compliance verdict.

What we do: run your system against frozen, published instruments; sign the result with Ed25519 and chain it to a SHA-256 hash; publish what we could not measure, in the same table, in the same breath.

Two minutes on what we measure and what we refuse to claim. Nothing in it issues a verdict.
Do not trust us — check

Three steps. Then you know.

A published card is a small record — under a kilobyte — carrying the axis, the model, the accuracy, the issuer, the date and the hash of the card before it. You can confirm it is genuine and unaltered on your own machine, with no account, no CSOAI code and no permission from us.

What this does not prove

Pin our key from /.well-known/did.json first. A card checked against the key it ships with proves only that the file is self-consistent, not that we issued it — anyone can alter a body and sign it with a key made a second ago.

Some board aggregates are explicitly uncarded and cannot be verified through this card path.

A card's trust path is an Ed25519 signature over a SHA-256 hash chain, verifiable offline against did:web:csoai.org — no blockchain and no timestamp authority sits in that path. The /xrpl-attest page is a reader of GET /root.json (signed root envelope; inclusion does not individually sign a leaf). GET /api/xrpl is a reader of that root (writes_board false, live locked 16, same merkle). Historical DEVNET Payment-memo / CredentialCreate hashes are not this feed. XLS-70 Credentials are live on XRPL mainnet as an allowlist primitive; we are not issuing GSPC grades on-ledger. Separately from the card trust path, The current canonical public root has a proof-derived STAMPED_PENDING_BITCOIN calendar proof; it does not yet prove inclusion in a Bitcoin block. That witness covers the exact public root.json bytes only, not the separate signed-card index. Queued and candidate atoms are not automatically admitted, published, or anchored; a pending calendar stamp, where one exists, does not by itself prove inclusion in a Bitcoin block.

The trust root did:web:csoai.org anchoring signed measurement cards through a hash-chained evidence ledger to local, offline verification on the reader's own machine
  1. 01

    Pin our key first — this step is not optional

    Fetch /.well-known/did.json and take the card-attestation key. Every published card must carry that exact pubkey. Verifying a card against the key it ships with proves only that the file is self-consistent — anyone can alter a body and sign it with a key they generated a second ago.

  2. 02

    Recompute the id from the body

    Canonicalise the card's body — every key sorted, no whitespace — and take the SHA-256. That hash must equal the card's id. One changed character and it will not match. One warning if you implement this outside Python: the bytes were written by CPython, which renders a float of integral value as 0.0 where JavaScript and Go write 0, so a naive verifier reports a false failure on a large minority of the set. Our verifier at /signed/verify-card.mjs handles it and the rule is written out at /signed/HOW-TO-VERIFY.md.

  3. 03

    Check the Ed25519 signature — then you are done

    Verify the signature over those same bytes under the pinned key. The whole check runs offline on your machine, with no CSOAI code, no account and no permission — or in your browser with WebCrypto. The banks and the grader are published too, so a measurement can be re-run as well as re-checked.

The public records that let you check us — source, packages, DOI, company register, trust root — are linked once, in the footer of every page.

  • benefitThe whole check runs offline — no account and no permission
  • benefitPin our key first. A card checked against the key it ships with only proves it is self-consistent
  • only hereYou recompute the same Ed25519 signature over the same hash chain we published
Self-correction

We publish our own errors — including the claim we withdrew.

Anyone can be right on a good day. Judge a measurement body on what it does on a bad one. Our corrections record at /api/corrections says, for every entry, what was wrong, how it was caught and what changed. It currently holds 78 entries.

What this does not prove

The record is source-maintained and is not backed by an append-only storage proof.

  • painMost measurement bodies quietly reword a claim that did not hold
  • benefitSigned artifacts are superseded rather than silently edited where that can be verified
  • only hereWe retracted our own consensus claim (DR-0007) rather than dress it up

The hardest one: we withdrew our own consensus claim. Our council architecture is a designed 33-seat structure with a designed 23-of-33 threshold. DR-0007 records the retraction; its historical numeric result is unbound because the cited result artifact is absent from this repository. The latest point experiment measured rho=1 and n_eff=1 across three nominal legs. Neither experiment demonstrates independent review or fault tolerance; the 33-seat council remains a design, not a live property.

  • C-2026-0822-012026-08-22
  • C-2026-0820-012026-08-20
  • C-2026-0819-132026-08-19
A report entering the intake, being mapped to frozen statutory provisions, then tested by deterministic predicates in a sandbox
Cropped to the honest half of the journey: report, provision mapping, deterministic sandbox test. Verdicts come from code, never from one model judging another.
A plain white clock face with a single green hand
Living law

Track regulatory changes. See what needs review for the EU AI Act — a new measurement is a separate run.

We watch selected primary sources — EUR-Lex, legislation.gov.uk and the national registers — and publish a dated deadline feed at /api/regulation.

What this does not prove

A source change raises a detection signal for review. It does not re-measure anything: re-measurement and delta-card issuance are not automated, and a fresh measurement appears only after a separate run is completed and admitted.

Previously published signed artifacts remain addressable. This page does not claim append-only storage.

Next up · verified as of 2026-09-12

  • 2026-12-02EU AI Act Art 50(2) — Marking grace ends for generative systems placed on market before 2 Aug 2026
  • 2026-12-02EU AI Act Art 5 (new) — Prohibitions on AI generating non-consensual intimate imagery and CSAM take effect
  • 2026-12-09EU Product Liability Directive — Member-state transposition deadline — software and AI enter strict no-fault liability
A field of pale solids linked by a lattice of green light
The board · stamped behavioural axes 2026-08-12 · jail 2026-08-18 · financial-fact axes 2026-08-25

The open board anyone can inspect — live, and recomputable.

A filled cell is a measurement. A dash is honest emptiness. Every figure in this band is read live from /api/gspc as the page loads — we do not type numbers into a page, because a typed number is the first thing to go stale.

The last slot is jail, containment: whether a model can be talked out of its own guardrails. It is measured on 71 gold cells, on a smaller fleet than the rest of the board, and its separation is TIE on the live board — a tie is not a separated leader. We print that instead of leaving the cell blank, and instead of dressing it up as a pass.

schematic of occupancy — not scores

Watch

Three films, if you would rather be told.

How a record is made, how the workspace fits together, and who the measurement is for. Tap to play — the file only loads then. Under each one, what it actually means.

  • Frozen tests. Deterministic grade. Then Ed25519. Not a certificate.

    What this film is saying

    • Frozen, published tests — the target does not move after you sit the run.
    • No model grades another model.
    • only hereA small signed card. Anyone checks it without us.
  • Board, verify, get measured. The same living board on / and in /os.

    What this film is saying

    • One workspace. No second login.
    • Empty cells stay empty. A dash is honest emptiness, never a dressed-up zero.
    • only hereA rank is never sold. Layout is not a purchase.
  • Insurers, labs, deployers. Evidence, not adjectives.

    What this film is saying

    • Built for people who need evidence, not adjectives.
    • We measure. We do not certify, accredit or enforce.
    • only hereNobody we measure pays for a place or a score.

Reviewed reading

Read the method behind the board.

Method, corrections, independence and verification — each link opens the reviewed page it names. Nothing here is promoted on the strength of a URL existing.

Methodology