Independent measurement · signed evidence · free to re-check
We measure how AI systems behave,and publish the evidence so you can check it yourself.
Frozen, published tests. Answers are graded by fixed rules, never by another AI. Issued measurement cards are signed; unsigned supporting runs are labelled. Re-check the evidence for free, without an account. Unmeasured stays visible.
Operated by CSOAI Ltd (UK Companies House 16939677), founded by Nicholas Templeman. Who we are
- on the board today
- 23 axis · 23 measured
- every declared slot and how many carry a real run
- graded rows behind it
- 1,230
- each row is one answer a rule graded, not a model's opinion
- model fleets tested
- 14
- one fleet per behavioural axis, all answering the same frozen questions
- runs graded from public facts
- 9
- no model, no fleet, no judgement — a rule reads a public record
Measured is not the same as separated.
A slot counts as measured when a real run sits behind it. Whether the axis actually told two models apart is a second question, and across the 14 model-comparison axes the answer today is 0 separated, 8 tied and 6 untested. A tie stays a tie and an untested axis stays untested; neither is rounded up into a ranking.
Read live from GET /api/gspc as this page rendered; the runs behind it were made behavioural axes 2026-08-12 · jail 2026-08-18 · financial-fact axes 2026-08-25. Nothing on this page is a certificate, and a tie between two models stays a tie.
Liveread
Measured in public. Published every day.
- more than 2.5 million PyPI downloads, all-time. Source: pepy.tech per package, daily record on Hugging Face csoai/distribution-footprint.
+45,363 this week
279K+ in the last 30 days · 399 of 400 packages answered · CSOAI and MEOK AI Labs · counted by pepy.tech
as of
- 99,903 dataset downloads on Hugging Face, last 30 days. Source: Hugging Face API · downloads (HF's rolling 30 days), summed over public csoai/* datasets.
106K+ all-time
as of
- 11,009 measurement capsules under the signed daily index. Source: Hugging Face csoai/evidence-index · measurement-index 2026-09-27.
index of 2026-09-27 · in Bitcoin block 968674
as of
- more than 249 thousand census rows published as open data. Source: Hugging Face datasets-server /size, summed over 12 csoai/*census* datasets.
12 census datasets
as of
- 23 of 23 board axes measured, every one. Source: GET /api/gspc → totals.
each axis has a run behind it
as of
- 335 signed cards, every one verifies. Source: GET /api/state → card_chain.bodies_verified_valid (the signed card index).
Ed25519, checked by the verifier we publish
as of
- 78 public corrections, all dated. Source: GET /api/corrections.
+18 this week
signed ledger: what was wrong, and the fix
as of
- 16 MCP tools served live. Source: POST /mcp → tools/list.
read from the tool surface itself
as of
- 25 x402 doors answering. Source: GET /api/x402-quotes (each door's own answer).
every listed door answered
as of
- 126 open datasets on Hugging Face. Source: Hugging Face API · datasets?author=csoai (public only).
16 likes
as of
Each number links to the source it was read from and carries its own date. A source that does not answer is left out, never shown as zero. Download counts include mirrors and automated traffic, so they are not people or customers. Every figure as JSON.
Where we take part
Members of the bodies writing the standards for secure AI, content provenance and digital identity.
- MemberOpen Secure AI Alliance (a Linux Foundation project)since signed agreement on file (private)what the record shows
- MemberC2PA (Coalition for Content Provenance and Authenticity, a Linux Foundation project)since signed agreement on file (private); public roster lists CSOAI LTDopen the evidence
- Member (Contributor)Decentralized Identity Foundationsince signed agreement on file (private)what the record shows
- W3C Community Groupsparticipant· 3
- IETF lists and draftsparticipant· 4
- EU AI Pactparticipant
- BSI ART/1 commentscontributor
- Open Invention Networkmember
- LOT Networkmember
Participation is not endorsement, and a listing is not adoption. Every entry links to its evidence. Every entry, its evidence, and what it does not mean
Do not take our word for it
Anchored outside our control, listed by others, dated as we go.
Check the anchors yourself
Timestamps, logs and keys you can open
Latest signed daily index
11,009 capsules · Ed25519-signed · OpenTimestamps proof beside it
Bitcoin block (OpenTimestamps)
anchors the chain head of 2026-09-27 · CHAIN_INTACT
Sigstore Rekor entries
public transparency-log entries for the 2026-09-27 index and its chain head
Signing keys
every signature above checks against these public keys
Companies House
the accountable company, on the public register
Preprint DOI
Same model, same prompts, different answers: item-level cross-hardware reproducibility of LLM evaluation results
Dataset DOIs on Hugging Face
csoai/gspc-board and 9 more csoai datasets carry a DOI
Independently listed
Indexes run by other people that list our work
- EthicalML awesome AI regulation list
README line 201 names councilof.ai
- awesome-public-datasets
README line 1406 lists a csoai dataset
- Glama MCP directory
the connector page names io.github.CSOAI-ORG/gspc
- Smithery
the server page links councilof.ai
- 402 Index
its search lists councilof.ai x402 endpoints
- API Evangelist profile
an independent API profile of councilof.ai
Each checked live on ; an index that stops naming us drops off.
A listing is not an endorsement.
Recent work
The latest things we published, dated
- Preprint: cross-hardware reproducibility of LLM evaluation results
DOI 10.5281/zenodo.22985467
- Public comment submitted to NIST on AI 200-2
a submission we made; it implies nothing about NIST's view of us
Why you can check us rather than believe us
Six things you can verify about this business before you trust a single number on it.
The six figures in this section are read from their owning public sources as this page loads. If a read fails, we say so rather than substitute a number. Open each source to check its date and limits before citing it.
We say what the board cannot yet tell apart
0 of 14
model-comparison axes separated a leader — 8 tied, 6 untested
A slot is MEASURED when a run sits behind it. That is not a finding that the axis told two models apart, and we do not print it as one.
Read the board, row by rowAnyone can re-check our signed cards, for nothing
335
signed records verified — kind: measured
Pin our public key, recompute the hash, check the signature. It runs offline, needs no account, and it is free forever. Verification is never sold.
Check a record in your browserOur worst results are published beside our best
0.29
the lowest fleet mean on the board — the care axis, published like every other
A measurement body that publishes only its wins is a marketing department. One of our own low scores is further down this page as a signed card; supporting runs elsewhere may be unsigned and are labelled as such.
See one of our own low scoresEvery claim we got wrong is written down
78
corrections published — ledger signature: VALID
What was wrong, how it was caught, what changed, dated. Signed records are superseded, never quietly edited — editing the bytes would break the signature that makes them checkable.
Read the corrections ledgerTimestamp proofs show what is attested and what is pending
639
proofs containing Bitcoin block-header attestations — 4 calendar-pending, of 643
The manifest parses proof bytes; it does not independently check the block headers against a Bitcoin node or verify every named subject file. A pending proof is a submission, not an anchor.
View the free timestamp previewWe take part where the rules are being written
33
participation records, each linked to its own evidence (manifest 2026-09-27)
Participation is not endorsement and a listing is not adoption. We hold no certification under any scheme, and we show no other body's logo to imply one.
See every record, and what it does not proveThe long version of this argument — the method, the machine surface, the films, and the answers on funding, limits and offline verification — is kept in full on one page. Nothing was deleted to shorten this one.
Live from GET /api/gspc
The living board
23 axes measured · 14 model fleets · 0 separated leaders · 9 public leader scores · 9 fact runs · TIE is TIE · not a certificate.
Rows are in board order — layout, not rank. Status, family, separation and leader state are printed as the API serves them. A TIE is a TIE. A withheld leader is a state, not an empty cell. Verify is free; a rank is never sold. Measurement, not certification.
- Governancegovernancen 237
- Safetysafetyn 36
- Provenanceprovenancen 32
- Continuitycontinuityn 33
- Conformanceconformancen 35
- Opennessopennessn 32
- Machinery Conformitymachinery-conformityn 33
- Carecaren 199
- Cross Realitycross-realityn 32
- Detector Interopdetector-interopn 33
- Art5 Safeguardart5-safeguardn 36
- Swarmswarmn 37
- Affectaffectn 41
- Jailjailn 71
- Effect Bindingeffect-bindingn 261 tool-call servers probed
- Provenance Controlsprovenance-controlsn 6 issuer accounts
- Reserve Attestationreserve-attestationn 16 issuer accounts
- Regulatory Frameworkregulatory-frameworkn 16 issuer accounts
- Distribution Integritydistribution-integrityn 16 issuer accounts
- Custody Disclosurecustody-disclosuren 16 issuer accounts
- Ai Adoption Componentsai-adoption-componentsn 2 public series
- Labour Componentslabour-componentsn 2 public series
- Humanoid Labour Indexhumanoid-labour-indexn 8 frozen vendor URLs
23 rows on this table · 23 MEASURED · gspc 15 · financial 8
Models with a public leader score
One entry per axis whose leader the board publishes. Ordered by point estimate on each model's own frozen bank — layout, not a cross-axis rank. A TIE is not a win.
gemma3:12b (base model)
leads Safety
94.4%81.9% – 98.5%
TIEn 36
A point lead the test could not separate from the fleet.
gemma3:12b (base model)
leads Openness
84.4%68.2% – 93.1%
TIEn 32
A point lead the test could not separate from the fleet.
llama3.2:3b (base model)
leads Provenance
71.9%54.6% – 84.4%
TIEn 32
A point lead the test could not separate from the fleet.
mistral:7b (base model)
leads Conformance
71.4%54.9% – 83.7%
TIEn 35
A point lead the test could not separate from the fleet.
gemma3:12b (base model)
leads Continuity
60.6%43.7% – 75.3%
TIEn 33
A point lead the test could not separate from the fleet.
qwen2.5:0.5b-instruct (base model)
leads Jail
59.2%47.5% – 69.8%
TIEn 71
A point lead the test could not separate from the fleet.
mistral:7b (base model)
leads Governance
58.7%52.3% – 64.7%
TIEn 237
A point lead the test could not separate from the fleet.
qwen2.5:7b (base model)
leads Swarm
44.4%
UNTESTEDn 37
qwen2.5:0.5b-instruct (base model)
leads Care
40.5%33.9% – 47.4%
TIEn 199
A point lead the test could not separate from the fleet.
9 public leader scores on the board today, counted from the rows above · the board's own count agrees (9) · 5 model-comparison axes withhold their leader (3 NO_SIGNED_CARD, 2 EXCLUDED_OWN_MODEL) · 9 fact runs have no fleet and no leader · nothing is padded and a TIE is not a win.
Hugging Face measured-model results
Third-party Hub cells from /api/hub-cards. This is a separate benchmark instrument from the GSPC board above. Each signed card is an observation; deterministic fact axes do not score models.
384 published MEASURED cells · 57 models · 14 model axes
Feed observed 2026-09-28T08:52:56.142Z
These observations may use different frozen banks or instruments. The feed does not identify a common comparison set, so their scores are not ranked or directly comparable. Check each signed card for its bank and instrument hashes.
| Model | Observed score | Evidence |
|---|---|---|
| deepseek-ai/DeepSeek-V3 | 53.3%n 30 | Signed card |
| deepseek-ai/DeepSeek-V3-0324 | 46.7%n 30 | Signed card |
| deepseek-ai/DeepSeek-V3.2 | 63.3%n 30 | Signed card |
| deepseek-ai/DeepSeek-V4-Flash | 46.7%n 30 | Signed card |
| deepseek-ai/DeepSeek-V4-Flash-0731 | 83.3%n 30 | Signed card |
| deepseek-ai/DeepSeek-V4-Pro | 36.7%n 30 | Signed card |
| farbodtavakkoli/OTel-2.0-LLM-31B-IT | 60%n 30 | Signed card |
| google/gemma-2-2b-it | 26.7%n 30 | Signed card |
| google/gemma-2-9b-it | 56.7%n 30 | Signed card |
Showing up to nine model-name-sorted observations on this axis. The displayed subset is not a top-nine ranking, separation test, winner claim, compliance verdict, or certificate. Open each signed card to verify its own evidence.
Ask it a question, or paste a record.
Name an axis and the board jumps to it. Paste a signed record and it is checked right here. Nothing leaves this device either way.
The number a vendor would bury
Here is one of our own models, scoring single digits.
We trained it. It has been on the board since the day it was measured, issued as a signed card under our published key, and it is not going anywhere. It is not even the bottom: the signed set runs all the way down to zero. A score that only ever goes up is not a measurement — it is marketing with a chart.
The panel beside this is not a screenshot. Your browser fetched the record, pinned our public key from the published identity document, recomputed the record's fingerprint from its own contents and checked the signature. Nothing was sent to us, and nothing needed our permission.
And you do not have to take our word for which record we chose to show you. Every signed record is listed here, each one a single fetch from its own body, so you can go and find the low ones yourself.
9.7%
VALIDsha256 of the canonical body equals the card id, and the signature verifies under did:web:csoai.org#card-attestation-1.
- what was tested
- care-refusal-protect
- which system
- clan-csoai-plain:latest
- signed on
- 2026-08-19T09:24:39.152331+00:00
- issued by
- CSOAI Ltd (UK 16939677)
record id · 82994353b8f94337746ddf73700b0edc425d695d43910dbfeb53d118d5a09a1c
One thing inside this record is out of date, on purpose. It describes the board as “13 measured of 14 quotable”, which was true when it was signed. Signed bytes are never edited here, so the old wording stays inside the signature and the live board above says what is true today. We supersede; we do not overwrite.
In the room, on the record
The standards that will govern this are being written now. We are in those rooms.
Standards bodies, public registries, scholarly identifiers and filings on the public record. Every entry below names what it proves, what it does not prove, and the evidence you can open for yourself — because a membership logo with nothing behind it is the exact thing this business exists to make unnecessary.
Where we take part
Every entry, its evidence, and what it does not mean →Standards bodies
- W3C AI KR CGparticipant
- W3C Agent Conformance CGparticipant
- W3C Agent Identity CGparticipant
- IETF Internet-Draftfiled
- IETF scitt listparticipant
- IETF agentproto listparticipant
- IETF audit listparticipant
- C2PAmember
- DIFmemberprivate evidence
- OSAIAmemberprivate evidence
- BSI ART/1 commentscontributorprivate evidence
Registries & indexes
Scholarly identifiers
Regulatory filings
Policy programmes
Open-source commons
Participation is not endorsement, and a listing is not adoption. Every entry links to its evidence. Manifest as of , unsigned; re-checked by scripts/memberships-check.mjs.
Keep going
That is the whole front door. Everything else is one click, not one scroll.
This page used to run to twenty-six screens on a desktop and fifty-three on a phone. None of it was deleted to shorten it — the bands below the board now have pages of their own, and the full archive is where it always was.
- How this works, in full The method and who pays for it, the machine surface and every data door, the nine products, the films, and the answers on funding, limits and our own errors.
- Check a record yourself Paste a signed record and your own browser does the maths. Nothing leaves the device, no account, free forever.
- How a measurement is made Frozen tests published before the run, graded by a rule rather than by another AI, with unparsed answers counted as incorrect.
- What we got wrong The public ledger: what was wrong, how it was caught, what changed, and the date. Signed records are superseded, never edited.
- How a claim is kept current The claim-maintenance specification we publish and follow: how a published claim is re-read, retired or corrected, with its DOI.
- Where this stands Operating evidence, stated plainly: what runs today and what is still early. Downloads and founder-funded tests are never shown as customers.
- Open source What we publish as open source and under which licence, so the instruments can be run without us.
- Where we take part Every participation record with its evidence, and a plain statement of what each one does not prove. We hold no certification under any scheme.
- The library Every page this estate has published, by subject and dated. Nothing is deleted when it is superseded.
