
You cannot underwrite what nobody has measured
Actuaries price tails from evidence. AI assurance mostly produces adjectives. We produce signed, deterministic rows — per-axis results with their sample sizes and intervals, recomputable offline — so the exposure can be modelled in your own book rather than taken on a vendor's word.
Measurement, not certification. 23 axis · 23 measured — live from /api/gspc
Human error is independent. Model failure is correlated.
Traditional indemnity works because a thousand people have a thousand separate bad days, and averages behave. One model deployed a thousand times has one bad day a thousand times at once. The law compounds this: statute does not grade your average, it names the events that must not happen. That is a tail problem wearing an averages costume.
- PainClaims history from human-error lines does not transfer to model failure
- PainA single weight update can move every deployed instance at once
- You getPer-axis failure mass, not just a headline accuracy
- Only hereWe publish severity-weighted harm alongside accuracy, because the mean hides the tail
From reactive indemnity to evidence you can price against
The gap is not appetite, it is inputs. Underwriters have questionnaires and incident anecdotes; what they need is a continuous, deterministic feed of how the system actually behaves against the provisions that carry the penalties. That is the only thing we make.
- PainSubjective questionnaires that the applicant fills in about themselves
- PainForensic claims investigations that take longer than the policy period
- You getDeterministic, signed evidence with n and interval on every row
- You getRe-measured when the model or the statute moves, not annually
- Only hereSame rows, same grader, same number — reproducible by you, without us

GSPC — governance, safety, provenance, continuity
Four axis families, anchored to a frozen corpus of 417 statutory provisions. The board publishes a slot count and a measured count and they are different numbers — cite totals.public_count on GET /api/gspc rather than either alone, because a slot with no run behind it is published UNMEASURED and is never counted as a measurement. The behavioural axis carry separation on the full fleet; jail is measured on a smaller fleet at n=71 and prints separation TIE — a tie is not a separated leader, and we never rank on it or compare it against the rest. Coverage on this site is read from the live board, never typed by hand.
- PainAssurance scores with no provision behind them and no n on the row
- You getEvery result names its provision, its sample size and its status
- You getWilson 95% intervals wherever n is honestly independent — and none where it is not
- Only hereA point-estimate lead with no statistical separation is printed as a tie, not a win
Behaviour in, signed rows out — and the model stays yours
An agent acts. We measure that behaviour against the frozen instrument and bound it statistically. We sign the result and hash-chain the record. Then we stop. We do not compute your tail measure, we do not hold a trigger, we do not settle anything. The distribution, the threshold and the payout logic stay inside your book, built on rows you can recompute.
- PainVendors who want to own the trigger as well as the data
- You getPer-item rows and aggregation functions published, so your actuaries rebuild the number
- You getEvidence you can attach to a file and defend later
- Only hereWe take no position in the risk we measure — no capacity, no broking, no settlement

One card pins the gold, the rows and the aggregator
Silent edits are the enemy of claims evidence, so the card fixes three things at once: the provision and threshold it was measured against, the per-item execution rows behind the number, and the exact aggregation function and version used to reduce them. Change any one and the Ed25519 signature over the SHA-256 hash chain stops verifying — offline, against the key published at did:web:csoai.org.
- PainEvidence that lives on the counterparty's server and can be quietly restated
- You getA ~3KB file on your side of the wall, verifiable without contacting us
- You getProvision, rows and aggregator all bound by one signature
- Only hereA surviving signature proves provenance — we say exactly that, and never that it proves correctness

Nobody we rank pays us. The relying party does.
This is the whole reason the evidence is worth anything. A grader funded by the graded is a marketing department with a methodology page. So the developer never pays and never can: verification is free forever, the rail is open, and the money comes from the observers who need the truth to be neutral — insurers, procurement, auditors.
- PainAssurance firms that grade a model and sell the remediation for the grade
- PainUnderwriters buying scores from a party that also carries the risk
- You getA grader with no commercial exposure to the outcome of any grade
- Only hereWe take no money from anything we rank — verification is free, forever, for everyone
One rail, three parties, no conflict of interest
The regulator sets the provisions and the penalty bands — under the EU AI Act those run to €35 million or 7% of worldwide turnover for prohibited practices, which is the statute's number and not a measurement of ours. The developer needs continuous evidence against those provisions. The insurer needs rows it can model. We are the piece in the middle that is paid by none of the first two.
- PainEvery party re-running its own half-trusted assessment of the same system
- You getOne signed measurement all three sides can read and check
- You getThird-party figures we cite are served separately, dated and attributed
- Only hereIndependence is structural here, not a promise in a policy document

For insurers & underwriters — verify everything, free
Evidence an underwriter can verify
We measure AI systems against the rules that govern them, sign the result, and publish what we cannot measure. Every number on this page is either fetched live from a signed endpoint or labelled with its third-party source and capture date — nothing is blended.
What this is not: not a certification, not a conformity mark, not a legal determination — it is a measurement record you can recompute yourself.
MEASURED
Our own signed, deterministic runs. Live at GET /api/gspc.
UNMEASURED
Honestly withheld, with the reason stated — insufficient n, or no separation test yet. Empty cells stay empty.
REPORTED
Third-party figures, cited and dated — reported by the source, not measured here. Live at GET /api/reported.
What a signed measurement card gives you
A measurement card is a small (~3KB) signed record. Each part exists so an actuary can check it without trusting the issuer:
Deterministic per-axis results
Every axis result carries its n and, where the n is honestly independent, a Wilson 95% interval. Same rows, same grader, rerun gives the same number — no model judging another model.
Ed25519 signature
The board is signed. The signature covers the canonical board content (minus the signature fields themselves), so any edit after signing breaks verification.
sha256 hash chain
Signed record sets chain their hashes: sha256 of the canonical content, sorted keys. Recompute it locally — if a record was edited after signing, your hash will not match the stored one.
SHA-256 hash chain
Record hashes are sha256-linked and Ed25519-signed, so “this content is unaltered since signing” is checkable offline against the published key. There is no independent time-stamping authority behind these cards: the anchor is the signature over the hash chain, and nothing more.
did:web:csoai.org published key
The Ed25519 public keys are published at GET /.well-known/did.json under did:web:csoai.org. You fetch the key from the domain itself — no key exchange with us required.
Verify one yourself, offline
A stranger with a terminal can check us. No account, no key exchange, no permission:
- Fetch the signed board —
curl https://councilof.ai/api/gspc
- Fetch the published verification key —
curl https://councilof.ai/.well-known/did.json
- Recompute the hash chain in your own browser at /gspc-verify — client-side, nothing leaves your machine.
Verification is free, requires no account, and always will be.
Loss context
Why a measurement record maps onto an underwriting file:
Frequency
The board shows which axis a system fails, with the n behind each result. A failure rate with a sample size and a Wilson interval is a frequency input, not a marketing claim.
Severity tails
Where n≥100, the board publishes mean_harm and cvar05_harm per axis. CVaR@5% is the average harm across the worst 5% of items — the tail an underwriter prices, not the average day. Where n<100 the field is honestly null.
Drift
Regulation changes; measurements go stale. A daily reg-watch detector watches the governing corpus, and state changes are published to /api/feed.xml so re-measurement is observable, not promised.
The honesty gate
At /honesty we publish our own models losing our own arena. An instrument that catches its owner is the one an underwriter can verify; we do not sell a rating, and the underwriter still prices.
The live board — MEASURED
Fetched live from GET /api/gspc (the live count lives there, not here). A TIE means the leader's edge is statistically indistinguishable — ties are never counted as wins.
| Axis | n | Leader accuracy | Separation |
|---|---|---|---|
| governance | 237 | 58.7% | TIE — indistinguishable |
| safety | 36 | 94.4% | TIE — indistinguishable |
| provenance | 32 | 71.9% | TIE — indistinguishable |
| continuity | 33 | 60.6% | TIE — indistinguishable |
| conformance | 35 | 71.4% | TIE — indistinguishable |
| openness | 32 | 84.4% | TIE — indistinguishable |
| machinery-conformity | 33 | no public leader score | UNTESTED |
| care | 199 | 40.5% | TIE — indistinguishable |
| cross-reality | 32 | no public leader score | UNTESTED |
| detector-interop | 33 | no public leader score | UNTESTED |
| art5-safeguard | 36 | no public leader score | UNTESTED |
| swarm | 37 | ≥44.4% | UNTESTED |
| affect | 41 | no public leader score | UNTESTED |
| jail | 71 | 59.2% | TIE — indistinguishable |
| effect-binding | 261tool-call servers probed | no leader accuracy261 of 600 servers tried, from a population of 20,992 third-party servers | MEASURED — deterministic factsnot applicable — no fleet, no leader |
| provenance-controls | 6issuer accounts (not bank items) | no leader accuracy6 of the 16 instruments named in the registry | MEASURED — deterministic factsnot applicable — no fleet, no leader |
| reserve-attestation | 16issuer accounts (not bank items) | no leader accuracy | MEASURED — deterministic factsnot applicable — no fleet, no leader |
| regulatory-framework | 16issuer accounts (not bank items) | no leader accuracy | MEASURED — deterministic factsnot applicable — no fleet, no leader |
| distribution-integrity | 16issuer accounts (not bank items) | no leader accuracy | MEASURED — deterministic factsnot applicable — no fleet, no leader |
| custody-disclosure | 16issuer accounts (not bank items) | no leader accuracy | MEASURED — deterministic factsnot applicable — no fleet, no leader |
| ai-adoption-components | 2public series | no leader accuracy | MEASURED — deterministic factsnot applicable — no fleet, no leader |
| labour-components | 2public series | no leader accuracy | MEASURED — deterministic factsnot applicable — no fleet, no leader |
| humanoid-labour-index | 8frozen vendor URLs | no leader accuracy | MEASURED — deterministic factsnot applicable — no fleet, no leader |
Full per-axis detail — Wilson intervals, fleet means, harm tails, the signature: /gspc-scoreboard and GET /api/gspc.
REPORTED — third-party context, never blended
Figures published by others, cited with their capture date. Each entry is reported by the source, not measured here — unsigned, and never enters the board.
No third-party figures are currently carried. An empty REPORTED set is the honest answer, not a missing section.
Measurement, not certification. CSOAI Ltd · UK Companies House 16939677 · nicholas@csoai.org. Every MEASURED number on this page is recomputable from GET /api/gspc; every REPORTED figure carries its source and capture date; what we cannot measure is withheld and says so.
What this page does not claim
We publish the limits with the results. Everything below is something a reader could reasonably assume from a page like this one — and each is something we cannot presently evidence, so we say so rather than let the assumption stand.
- We do not publish a market size for AI insurance. No premium projection, no CAGR, no carrier capacity figure appears on this page — none of it is ours to evidence. Third-party figures we do cite are served from /api/reported, dated and attributed, as reported by the source and not measured here.
- We do not operate an oracle, a trigger or a settlement mechanism. We publish signed rows; the tail measure, the threshold and the payout logic are computed inside the insurer's own model.
- We do not claim a signature proves a system is safe or lawful. It proves provenance — these bytes, unaltered, from this key. The measurement is what speaks to behaviour, and it comes with its limits attached.
- We do not claim slot 14 is comparable to the canonical axes. Jail is measured at n=71 on a smaller fleet with separation TIE on the live board — a tie is not a separated leader, and we never rank on it.
- We do not certify, accredit or approve. There is no conformity mark here and no accreditation chain behind it.
Coverage on this page is never typed by hand. 23 axis · 23 measured — live from /api/gspc Corrections to anything we have published live in the refutation ledger — append-only, never a silent edit.