Clay figures before a vault of signed measurement records
Pricing AI risk

You cannot underwrite what nobody has measured

Actuaries price tails from evidence. AI assurance mostly produces adjectives. We produce signed, deterministic rows — per-axis results with their sample sizes and intervals, recomputable offline — so the exposure can be modelled in your own book rather than taken on a vendor's word.

Measurement, not certification. 23 axis · 23 measured — live from /api/gspc

Why AI breaks the usual maths

Human error is independent. Model failure is correlated.

Traditional indemnity works because a thousand people have a thousand separate bad days, and averages behave. One model deployed a thousand times has one bad day a thousand times at once. The law compounds this: statute does not grade your average, it names the events that must not happen. That is a tail problem wearing an averages costume.

  • PainClaims history from human-error lines does not transfer to model failure
  • PainA single weight update can move every deployed instance at once
  • You getPer-axis failure mass, not just a headline accuracy
  • Only hereWe publish severity-weighted harm alongside accuracy, because the mean hides the tail
What changes

From reactive indemnity to evidence you can price against

The gap is not appetite, it is inputs. Underwriters have questionnaires and incident anecdotes; what they need is a continuous, deterministic feed of how the system actually behaves against the provisions that carry the penalties. That is the only thing we make.

  • PainSubjective questionnaires that the applicant fills in about themselves
  • PainForensic claims investigations that take longer than the policy period
  • You getDeterministic, signed evidence with n and interval on every row
  • You getRe-measured when the model or the statute moves, not annually
  • Only hereSame rows, same grader, same number — reproducible by you, without us
The measurement arena seen from above, its instrument slots laid out
The instrument

GSPC — governance, safety, provenance, continuity

Four axis families, anchored to a frozen corpus of 417 statutory provisions. The board publishes a slot count and a measured count and they are different numbers — cite totals.public_count on GET /api/gspc rather than either alone, because a slot with no run behind it is published UNMEASURED and is never counted as a measurement. The behavioural axis carry separation on the full fleet; jail is measured on a smaller fleet at n=71 and prints separation TIE — a tie is not a separated leader, and we never rank on it or compare it against the rest. Coverage on this site is read from the live board, never typed by hand.

  • PainAssurance scores with no provision behind them and no n on the row
  • You getEvery result names its provision, its sample size and its status
  • You getWilson 95% intervals wherever n is honestly independent — and none where it is not
  • Only hereA point-estimate lead with no statistical separation is printed as a tie, not a win
Open the board
The pipeline

Behaviour in, signed rows out — and the model stays yours

An agent acts. We measure that behaviour against the frozen instrument and bound it statistically. We sign the result and hash-chain the record. Then we stop. We do not compute your tail measure, we do not hold a trigger, we do not settle anything. The distribution, the threshold and the payout logic stay inside your book, built on rows you can recompute.

  • PainVendors who want to own the trigger as well as the data
  • You getPer-item rows and aggregation functions published, so your actuaries rebuild the number
  • You getEvidence you can attach to a file and defend later
  • Only hereWe take no position in the risk we measure — no capacity, no broking, no settlement
Hands holding a signed evidence card reading verified: true
The claims artefact

One card pins the gold, the rows and the aggregator

Silent edits are the enemy of claims evidence, so the card fixes three things at once: the provision and threshold it was measured against, the per-item execution rows behind the number, and the exact aggregation function and version used to reduce them. Change any one and the Ed25519 signature over the SHA-256 hash chain stops verifying — offline, against the key published at did:web:csoai.org.

  • PainEvidence that lives on the counterparty's server and can be quietly restated
  • You getA ~3KB file on your side of the wall, verifiable without contacting us
  • You getProvision, rows and aggregator all bound by one signature
  • Only hereA surviving signature proves provenance — we say exactly that, and never that it proves correctness
Verify a card
A wall separating those who are measured from those who rely on the measurement
The neutrality firewall

Nobody we rank pays us. The relying party does.

This is the whole reason the evidence is worth anything. A grader funded by the graded is a marketing department with a methodology page. So the developer never pays and never can: verification is free forever, the rail is open, and the money comes from the observers who need the truth to be neutral — insurers, procurement, auditors.

  • PainAssurance firms that grade a model and sell the remediation for the grade
  • PainUnderwriters buying scores from a party that also carries the risk
  • You getA grader with no commercial exposure to the outcome of any grade
  • Only hereWe take no money from anything we rank — verification is free, forever, for everyone
Where it sits

One rail, three parties, no conflict of interest

The regulator sets the provisions and the penalty bands — under the EU AI Act those run to €35 million or 7% of worldwide turnover for prohibited practices, which is the statute's number and not a measurement of ours. The developer needs continuous evidence against those provisions. The insurer needs rows it can model. We are the piece in the middle that is paid by none of the first two.

  • PainEvery party re-running its own half-trusted assessment of the same system
  • You getOne signed measurement all three sides can read and check
  • You getThird-party figures we cite are served separately, dated and attributed
  • Only hereIndependence is structural here, not a promise in a policy document
See what is reported vs measured
Regulator, developer and insurer around a shared measurement

For insurers & underwriters — verify everything, free

Evidence an underwriter can verify

We measure AI systems against the rules that govern them, sign the result, and publish what we cannot measure. Every number on this page is either fetched live from a signed endpoint or labelled with its third-party source and capture date — nothing is blended.

What this is not: not a certification, not a conformity mark, not a legal determination — it is a measurement record you can recompute yourself.

MEASURED

Our own signed, deterministic runs. Live at GET /api/gspc.

UNMEASURED

Honestly withheld, with the reason stated — insufficient n, or no separation test yet. Empty cells stay empty.

REPORTED

Third-party figures, cited and dated — reported by the source, not measured here. Live at GET /api/reported.

What a signed measurement card gives you

A measurement card is a small (~3KB) signed record. Each part exists so an actuary can check it without trusting the issuer:

  • Deterministic per-axis results

    Every axis result carries its n and, where the n is honestly independent, a Wilson 95% interval. Same rows, same grader, rerun gives the same number — no model judging another model.

  • Ed25519 signature

    The board is signed. The signature covers the canonical board content (minus the signature fields themselves), so any edit after signing breaks verification.

  • sha256 hash chain

    Signed record sets chain their hashes: sha256 of the canonical content, sorted keys. Recompute it locally — if a record was edited after signing, your hash will not match the stored one.

  • SHA-256 hash chain

    Record hashes are sha256-linked and Ed25519-signed, so “this content is unaltered since signing” is checkable offline against the published key. There is no independent time-stamping authority behind these cards: the anchor is the signature over the hash chain, and nothing more.

  • did:web:csoai.org published key

    The Ed25519 public keys are published at GET /.well-known/did.json under did:web:csoai.org. You fetch the key from the domain itself — no key exchange with us required.

Verify one yourself, offline

A stranger with a terminal can check us. No account, no key exchange, no permission:

  1. Fetch the signed board —
    curl https://councilof.ai/api/gspc
  2. Fetch the published verification key —
    curl https://councilof.ai/.well-known/did.json
  3. Recompute the hash chain in your own browser at /gspc-verify — client-side, nothing leaves your machine.

Verification is free, requires no account, and always will be.

Loss context

Why a measurement record maps onto an underwriting file:

Frequency

The board shows which axis a system fails, with the n behind each result. A failure rate with a sample size and a Wilson interval is a frequency input, not a marketing claim.

Severity tails

Where n≥100, the board publishes mean_harm and cvar05_harm per axis. CVaR@5% is the average harm across the worst 5% of items — the tail an underwriter prices, not the average day. Where n<100 the field is honestly null.

Drift

Regulation changes; measurements go stale. A daily reg-watch detector watches the governing corpus, and state changes are published to /api/feed.xml so re-measurement is observable, not promised.

The honesty gate

At /honesty we publish our own models losing our own arena. An instrument that catches its owner is the one an underwriter can verify; we do not sell a rating, and the underwriter still prices.

The live board — MEASURED

Fetched live from GET /api/gspc (the live count lives there, not here). A TIE means the leader's edge is statistically indistinguishable — ties are never counted as wins.

AxisnLeader accuracySeparation
governance23758.7%TIE — indistinguishable
safety3694.4%TIE — indistinguishable
provenance3271.9%TIE — indistinguishable
continuity3360.6%TIE — indistinguishable
conformance3571.4%TIE — indistinguishable
openness3284.4%TIE — indistinguishable
machinery-conformity33no public leader scoreUNTESTED
care19940.5%TIE — indistinguishable
cross-reality32no public leader scoreUNTESTED
detector-interop33no public leader scoreUNTESTED
art5-safeguard36no public leader scoreUNTESTED
swarm37≥44.4%UNTESTED
affect41no public leader scoreUNTESTED
jail7159.2%TIE — indistinguishable
effect-binding261tool-call servers probedno leader accuracy261 of 600 servers tried, from a population of 20,992 third-party serversMEASURED — deterministic factsnot applicable — no fleet, no leader
provenance-controls6issuer accounts (not bank items)no leader accuracy6 of the 16 instruments named in the registryMEASURED — deterministic factsnot applicable — no fleet, no leader
reserve-attestation16issuer accounts (not bank items)no leader accuracyMEASURED — deterministic factsnot applicable — no fleet, no leader
regulatory-framework16issuer accounts (not bank items)no leader accuracyMEASURED — deterministic factsnot applicable — no fleet, no leader
distribution-integrity16issuer accounts (not bank items)no leader accuracyMEASURED — deterministic factsnot applicable — no fleet, no leader
custody-disclosure16issuer accounts (not bank items)no leader accuracyMEASURED — deterministic factsnot applicable — no fleet, no leader
ai-adoption-components2public seriesno leader accuracyMEASURED — deterministic factsnot applicable — no fleet, no leader
labour-components2public seriesno leader accuracyMEASURED — deterministic factsnot applicable — no fleet, no leader
humanoid-labour-index8frozen vendor URLsno leader accuracyMEASURED — deterministic factsnot applicable — no fleet, no leader

Full per-axis detail — Wilson intervals, fleet means, harm tails, the signature: /gspc-scoreboard and GET /api/gspc.

REPORTED — third-party context, never blended

Figures published by others, cited with their capture date. Each entry is reported by the source, not measured here — unsigned, and never enters the board.

No third-party figures are currently carried. An empty REPORTED set is the honest answer, not a missing section.

Measurement, not certification. CSOAI Ltd · UK Companies House 16939677 · nicholas@csoai.org. Every MEASURED number on this page is recomputable from GET /api/gspc; every REPORTED figure carries its source and capture date; what we cannot measure is withheld and says so.

Read this before you quote us

What this page does not claim

We publish the limits with the results. Everything below is something a reader could reasonably assume from a page like this one — and each is something we cannot presently evidence, so we say so rather than let the assumption stand.

  • We do not publish a market size for AI insurance. No premium projection, no CAGR, no carrier capacity figure appears on this page — none of it is ours to evidence. Third-party figures we do cite are served from /api/reported, dated and attributed, as reported by the source and not measured here.
  • We do not operate an oracle, a trigger or a settlement mechanism. We publish signed rows; the tail measure, the threshold and the payout logic are computed inside the insurer's own model.
  • We do not claim a signature proves a system is safe or lawful. It proves provenance — these bytes, unaltered, from this key. The measurement is what speaks to behaviour, and it comes with its limits attached.
  • We do not claim slot 14 is comparable to the canonical axes. Jail is measured at n=71 on a smaller fleet with separation TIE on the live board — a tie is not a separated leader, and we never rank on it.
  • We do not certify, accredit or approve. There is no conformity mark here and no accreditation chain behind it.

Coverage on this page is never typed by hand. 23 axis · 23 measured — live from /api/gspc Corrections to anything we have published live in the refutation ledger — append-only, never a silent edit.