
Board source: GET /api/gspc — recompute published results, free
The GSPC board
23 axes measured · 14 model fleets · 0 separated leaders · 9 public leader scores · 9 fact runs · TIE is TIE · not a certificate. (14 model-comparison + 9 fact runs — deterministic fact checks have no model fleet, leader or accuracy) · deterministic grading on frozen, published splits · a TIE means the leader's edge is statistically indistinguishable (McNemar p≥0.05) — ties are never counted as wins. Empty cells stay empty. Measured axes are not the same as public leader scores.
23 axes are on the board and every one carries a measurement — no declared slot is empty. Both counts are DERIVED from the axis array, never typed; if a future slot is added with no run behind it, this line separates the two again on its own. A measurement is not a separated leader: 0 of 14 model-comparison axes separated a leader · 8 TIE · 6 UNTESTED. A point-estimate lead is not a measured advantage, and UNTESTED is not a tie.
GSPC terminal · one board
23 axis · 23 measured
npx -y csoai-gspc-mcp| Axis | Bench | Figure | n | Status | |
|---|---|---|---|---|---|
| governance | GovBench | 58.7% | 237 | TIE | ▸ |
| safety | DefBench | 94.4% | 36 | TIE | ▸ |
| provenance | ProvBench | 71.9% | 32 | TIE | ▸ |
| continuity | PQCBench | 60.6% | 33 | TIE | ▸ |
| conformance | MCPBench | 71.4% | 35 | TIE | ▸ |
| openness | OSSBench | 84.4% | 32 | TIE | ▸ |
| machinery-conformity | MachBench | UNMEASURED | 33 | MEASURED | ▸ |
| care | CareBench | 40.5% | 199 | TIE | ▸ |
| cross-reality | XRAIV | UNMEASURED | 32 | MEASURED | ▸ |
| detector-interop | DetBench | UNMEASURED | 33 | MEASURED | ▸ |
| art5-safeguard | Art5Bench | UNMEASURED | 36 | MEASURED | ▸ |
| swarm | SwarmBench v2b | 44.4% | 37 | UNTESTED | ▸ |
| affect | AffectBench | UNMEASURED | 41 | MEASURED | ▸ |
| jail | GoldBank-Detector | 59.2% | 71 | TIE | ▸ |
| effect-binding | EffectBench v0.1 (server probe) | facts | 261 | FACTS | ▸ |
| provenance-controls | ChainFacts | facts | 6 | FACTS | ▸ |
| reserve-attestation | ReserveFacts | facts | 16 | FACTS | ▸ |
| regulatory-framework | RegimeFacts | facts | 16 | FACTS | ▸ |
| distribution-integrity | DistributionFacts | facts | 16 | FACTS | ▸ |
| custody-disclosure | CustodyFacts | facts | 16 | FACTS | ▸ |
| ai-adoption-components | Eurostat | facts | 2 | FACTS | ▸ |
| labour-components | Eurostat | facts | 2 | FACTS | ▸ |
| humanoid-labour-index | Disclosure | facts | 8 | FACTS | ▸ |
The full board — with intervals and harm tails
Axis deep-dive
care
Bench CareBench · n=199 · leader accuracy 40.5% · 95% CI 33.9–47.4% · separation TIE
No model separated from the next best on this axis (exact McNemar, p≥0.05, n=200; p=0.427).
Leader qwen2.5:0.5b-instruct 81/200 40.5% [33.9–47.4%] vs next best qwen2.5:3b 74/200 37.0% [30.6–43.9%] · 199 paired items · exact McNemar p=0.427
The test counts 200 rows per model over 199 distinct paired items; the board's n (199) counts unique scored texts. The rows are published as frozen, duplicates included.
leader shown from per-item rows; no signed per-model card yet for this run — the signed card for qwen2.5:0.5b-instruct on this axis records a different measurement (accuracy 0.0), so it does not back the number shown · the card on record
Our own council specialist held the point lead in the full 19-model fleet and is excluded: a neutral measurement body does not rank its own models against the vendors it measures. The leader shown is the external-only re-rank of the same published rows (6 base models).
Frozen gold bank (Hugging Face)Verify the signed chainRaw JSON (GET /api/gspc)Full board
Bank exposure. The 106 live signed cards on this axis pin banks with different exposure: 72 UNPINNED, 34 PUBLIC_BANK. Measured on a public benchmark. Every item in the bank behind this number is publicly readable, so a model could have been trained on it. Read it as a score on a public test, not on unseen items; no held-out slice backs it.
106 live signed cards on this axis · per-card labels · signed record · the label changes no score.
| Axis | Bench | n | Leader accuracy | 95% CI | Separation | Open traces / graphs |
|---|---|---|---|---|---|---|
| governance | GovBench | 237 | 58.7% | 52.3–64.7% | TIE — indistinguishablep=0.3787 | |
| safety | DefBench | 36 | 94.4% | 81.9–98.5% | TIE — indistinguishablep=0.6875 | |
| provenance | ProvBench | 32 | 71.9% | 54.6–84.4% | TIE — indistinguishablep=1 | |
| continuity | PQCBench | 33 | 60.6% | 43.7–75.3% | TIE — indistinguishablep=0.6875 | |
| conformance | MCPBench | 35 | 71.4% | 54.9–83.7% | TIE — indistinguishablep=0.625 | |
| openness | OSSBench | 32 | 84.4% | 68.2–93.1% | TIE — indistinguishablep=0.424 | |
| machinery-conformity | MachBench | 33 | no public leader score | not published — no public leader score | UNTESTED | |
| care | CareBench | 199 | 40.5% | 33.9–47.4% | TIE — indistinguishablep=0.427 | |
| cross-reality | XRAIV | 32 | no public leader score | not published — no public leader score | UNTESTED | |
| detector-interop | DetBench | 33 | no public leader score | not published — no public leader score | UNTESTED | |
| art5-safeguard | Art5Bench | 36 | no public leader score | not published — no public leader score | UNTESTED | |
| swarm | SwarmBench v2b | 37 | ≥44.4%lower bound | withheld (n not independent) | UNTESTED | |
| affect | AffectBench | 41 | no public leader score | not published — no public leader score | UNTESTED | |
| jail | GoldBank-Detector | 71 | 59.2% | 47.5–69.8% | TIE — indistinguishable | |
| effect-binding | EffectBench v0.1 (server probe) | 261tool-call servers probed | no leader accuracy261 of 600 servers tried, from a population of 20,992 third-party servers | not applicable — no accuracy to bound | MEASURED — deterministic factsnot applicable — no fleet, no leader | |
| provenance-controls | ChainFacts | 6issuer accounts (not bank items) | no leader accuracy6 of the 16 instruments named in the registry | not applicable — no accuracy to bound | MEASURED — deterministic factsnot applicable — no fleet, no leader | |
| reserve-attestation | ReserveFacts | 16issuer accounts (not bank items) | no leader accuracy | not applicable — no accuracy to bound | MEASURED — deterministic factsnot applicable — no fleet, no leader | |
| regulatory-framework | RegimeFacts | 16issuer accounts (not bank items) | no leader accuracy | not applicable — no accuracy to bound | MEASURED — deterministic factsnot applicable — no fleet, no leader | |
| distribution-integrity | DistributionFacts | 16issuer accounts (not bank items) | no leader accuracy | not applicable — no accuracy to bound | MEASURED — deterministic factsnot applicable — no fleet, no leader | |
| custody-disclosure | CustodyFacts | 16issuer accounts (not bank items) | no leader accuracy | not applicable — no accuracy to bound | MEASURED — deterministic factsnot applicable — no fleet, no leader | |
| ai-adoption-components | Eurostat | 2public series | no leader accuracy | not applicable — no accuracy to bound | MEASURED — deterministic factsnot applicable — no fleet, no leader | |
| labour-components | Eurostat | 2public series | no leader accuracy | not applicable — no accuracy to bound | MEASURED — deterministic factsnot applicable — no fleet, no leader | |
| humanoid-labour-index | Disclosure | 8frozen vendor URLs | no leader accuracy | not applicable — no accuracy to bound | MEASURED — deterministic factsnot applicable — no fleet, no leader |
Board dataset on Hugging Faceaxis,bench,status,n,accuracy,interval_lo,interval_hi,separation,fleet_mean,dataset,as_of · empty cells stay empty — never zeroed, never interpolated. The dataset is a versioned mirror; compare its timestamp with live GET /api/gspc.
Separation from the published per-item rows
15,580 rows, published byte-identical at csoai/gspc-peritem-rows-2026-08-12 · peritem_sha256 0d8dacfbe7384a5d2f6a83e7455482ab75f935366bbb6e18da61cfdac8890dec · signed record. The test was fixed on 2026-08-13: exact McNemar on the discordant items, leader vs the best base model, and p<0.05 is required to separate. Our own models are removed before ranking. A TIE is not a win.
governance · TIE
No model separated from the next best on this axis (exact McNemar, p≥0.05, n=237; p=0.3787).
Leader mistral:7b 139/237 58.7% [52.3–64.7%] vs next best deepseek-r1:8b 128/237 54.0% [47.6–60.2%] · 237 paired items · exact McNemar p=0.3787
leader shown from per-item rows; no signed per-model card yet for this run — the signed card for mistral:7b on this axis records a different measurement (accuracy 0.0), so it does not back the number shown · the card on record
Our own council specialist held the point lead in the full 19-model fleet and is excluded: a neutral measurement body does not rank its own models against the vendors it measures. The leader shown is the external-only re-rank of the same published rows (6 base models).
safety · TIE
No model separated from the next best on this axis (exact McNemar, p≥0.05, n=36; p=0.6875).
Leader gemma3:12b 34/36 94.4% [81.9–98.5%] vs next best qwen2.5:3b 32/36 88.9% [74.7–95.6%] · 36 paired items · exact McNemar p=0.6875
leader shown from per-item rows; no signed per-model card yet
provenance · TIE
No model separated from the next best on this axis (exact McNemar, p≥0.05, n=32; p=1.0).
Leader llama3.2:3b 23/32 71.9% [54.6–84.4%] vs next best gemma3:12b 22/32 68.8% [51.4–82.0%] · 32 paired items · exact McNemar p=1
leader shown from per-item rows; no signed per-model card yet for this run — the signed card for llama3.2:3b on this axis records a different measurement (accuracy 0.8), so it does not back the number shown · the card on record
Our own council specialist held the point lead in the full 19-model fleet and is excluded: a neutral measurement body does not rank its own models against the vendors it measures. The leader shown is the external-only re-rank of the same published rows (6 base models).
continuity · TIE
No model separated from the next best on this axis (exact McNemar, p≥0.05, n=33; p=0.6875).
Leader gemma3:12b 20/33 60.6% [43.7–75.3%] vs next best deepseek-r1:8b 18/33 54.5% [38.0–70.2%] · 33 paired items · exact McNemar p=0.6875
leader shown from per-item rows; no signed per-model card yet
Our own council specialist held the point lead in the full 19-model fleet and is excluded: a neutral measurement body does not rank its own models against the vendors it measures. The leader shown is the external-only re-rank of the same published rows (6 base models).
conformance · TIE
No model separated from the next best on this axis (exact McNemar, p≥0.05, n=35; p=0.625).
Leader mistral:7b 25/35 71.4% [54.9–83.7%] vs next best llama3.2:3b 23/35 65.7% [49.2–79.2%] · 35 paired items · exact McNemar p=0.625
leader shown from per-item rows; no signed per-model card yet
Our own council specialist held the point lead in the full 19-model fleet and is excluded: a neutral measurement body does not rank its own models against the vendors it measures. The leader shown is the external-only re-rank of the same published rows (6 base models).
openness · TIE
No model separated from the next best on this axis (exact McNemar, p≥0.05, n=32; p=0.424).
Leader gemma3:12b 27/32 84.4% [68.2–93.1%] vs next best mistral:7b 23/32 71.9% [54.6–84.4%] · 32 paired items · exact McNemar p=0.424
leader shown from per-item rows; no signed per-model card yet
Our own council specialist held the point lead in the full 19-model fleet and is excluded: a neutral measurement body does not rank its own models against the vendors it measures. The leader shown is the external-only re-rank of the same published rows (6 base models).
care · TIE
No model separated from the next best on this axis (exact McNemar, p≥0.05, n=200; p=0.427).
Leader qwen2.5:0.5b-instruct 81/200 40.5% [33.9–47.4%] vs next best qwen2.5:3b 74/200 37.0% [30.6–43.9%] · 199 paired items · exact McNemar p=0.427
The test counts 200 rows per model over 199 distinct paired items; the board's n (199) counts unique scored texts. The rows are published as frozen, duplicates included.
leader shown from per-item rows; no signed per-model card yet for this run — the signed card for qwen2.5:0.5b-instruct on this axis records a different measurement (accuracy 0.0), so it does not back the number shown · the card on record
Our own council specialist held the point lead in the full 19-model fleet and is excluded: a neutral measurement body does not rank its own models against the vendors it measures. The leader shown is the external-only re-rank of the same published rows (6 base models).
Rows published, no determination
- machinery-conformity — UNTESTED. The published rows give TIE (exact McNemar p=0.5811, n=33), but this axis has no signed card of any model in the public card index (/signed/card_index.json). The board publishes a separation determination only on axes that carry signed cards, so this one stays UNTESTED rather than resting on rows alone.
- cross-reality — UNTESTED. The published rows give TIE (exact McNemar p=0.0654, n=32), but this axis has no signed card of any model in the public card index (/signed/card_index.json). The board publishes a separation determination only on axes that carry signed cards, so this one stays UNTESTED rather than resting on rows alone.
- detector-interop — UNTESTED. The published rows give TIE (exact McNemar p=0.4531, n=33), but this axis has no signed card of any model in the public card index (/signed/card_index.json). The board publishes a separation determination only on axes that carry signed cards, so this one stays UNTESTED rather than resting on rows alone.
- art5-safeguard — UNTESTED. The published rows give TIE (exact McNemar p=1.0, n=36), but this axis has no signed card of any model in the public card index (/signed/card_index.json). The board publishes a separation determination only on axes that carry signed cards, so this one stays UNTESTED rather than resting on rows alone.
- swarm — UNTESTED. The published swarm rows are the retired 3-prompt PROTOCOL bank (40 rows per model over 3 distinct items), not the wave-2b bank the board's swarm row serves; rows from one bank cannot decide a separation determination about another.
- affect — UNTESTED. The published rows give TIE (exact McNemar p=1.0, n=41), but this axis has no signed card of any model in the public card index (/signed/card_index.json). The board publishes a separation determination only on axes that carry signed cards, so this one stays UNTESTED rather than resting on rows alone.
Think a row, a grade or a named model is wrong? Object or ask for a re-check.
Attestation · live from GET /api/gspc · click any row for traces
Measurement freshness · derived from GET /api/gspc · measured_on
behavioural axes 2026-08-12 · jail 2026-08-18 · financial-fact axes 2026-08-25
living_stamp.gold_run 18 Aug 2026 — payload stamp, not a live re-measure. Board counts stay derived from GET /api/gspc; no new MEASURED invented here.
These run dates are weeks old. Freshness is labelled; the board is not re-stamped from this UI.
Living Stamp — SIGNED
Do not treat this as a valid attestation. Check site_attestation instead.
Progress · 23 axis · 23 measured
N→N+1 drift · UNCHECKABLE
No published board time series for N→N+1 drift. Empty stays empty — do not invent drift numbers or a Merkle seal. Cite GET /root.json and GET /api/gspc for the living snapshot only. Living snapshot only — cite GET /root.json.
23 axes are on the board and every one carries a measurement — no declared slot is empty. Both counts are DERIVED from the axis array, never typed; if a future slot is added with no run behind it, this line separates the two again on its own. A measurement is not a separated leader: 0 of 14 model-comparison axes separated a leader · 8 TIE · 6 UNTESTED. A point-estimate lead is not a measured advantage, and UNTESTED is not a tie.
Arena Elo — signed
Per-axis Elo from the published arena snapshot (7 models, 28 arena axis — the arena's own set, not the board's count above). Snapshot: . New arena rounds do not change GSPC board axes without admission. The board DID signature is present; use the check below to verify these bytes.
| Model | Elo | Games | Win-rate | 95% CI |
|---|---|---|---|---|
| qwen3:8b | 1629 | 232 | 0.672 | 0.61–0.73 |
| llama3.1:8b | 1563.6 | 64 | 0.5 | 0.381–0.619 |
| phi3.5:3.8b | 1558.2 | 199 | 0.548 | 0.478–0.615 |
| phi4:14b | 1527.3 | 309 | 0.563 | 0.507–0.617 |
| mistral:7b | 1438.3 | 580 | 0.603 | 0.563–0.642 |
| nemotron-3-nano:30b | 1389.6 | 305 | 0.469 | 0.414–0.525 |
| gemma3:12b | 1345.4 | 500 | 0.258 | 0.222–0.298 |
In-lane measurements — not board rows
2 slots measured in-lane. Published as measured_in_lane on GET /api/gspc. NOT stamped onto the board count. public_count stays “23 axis · 23 measured” — quoted live, never typed. Click any card for traces + graphs.
Measurement, not certification. Leaders shown are point estimates (swarm quotes its 95% lower bound); only SEPARATED leads are statistically real — the live count is totals.separated_leads on GET /api/gspc. Jail is a measured floor when the stamp publishes one, never a hidden score. Full per-axis notes, fleet means and harm tails: GET /api/gspc. The living stamp carried there is marked UNVERIFIABLE — it does not reproduce under any published rule and is not a checkable attestation; the attestation over that payload that does verify is site_attestation.