Measured systems facing each other across the GSPC board

Board source: GET /api/gspc — recompute published results, free

The GSPC board

23 axes measured · 14 model fleets · 0 separated leaders · 9 public leader scores · 9 fact runs · TIE is TIE · not a certificate. (14 model-comparison + 9 fact runs — deterministic fact checks have no model fleet, leader or accuracy) · deterministic grading on frozen, published splits · a TIE means the leader's edge is statistically indistinguishable (McNemar p≥0.05) — ties are never counted as wins. Empty cells stay empty. Measured axes are not the same as public leader scores.

23 axes are on the board and every one carries a measurement — no declared slot is empty. Both counts are DERIVED from the axis array, never typed; if a future slot is added with no run behind it, this line separates the two again on its own. A measurement is not a separated leader: 0 of 14 model-comparison axes separated a leader · 8 TIE · 6 UNTESTED. A point-estimate lead is not a measured advantage, and UNTESTED is not a tie.

GSPC terminal · one board

23 axis · 23 measured

GSPC live badgenpx -y csoai-gspc-mcp
AxisBenchFigurenStatus
governanceGovBench58.7%237TIE▸
safetyDefBench94.4%36TIE▸
provenanceProvBench71.9%32TIE▸
continuityPQCBench60.6%33TIE▸
conformanceMCPBench71.4%35TIE▸
opennessOSSBench84.4%32TIE▸
machinery-conformityMachBenchUNMEASURED33MEASURED▸
careCareBench40.5%199TIE▸
cross-realityXRAIVUNMEASURED32MEASURED▸
detector-interopDetBenchUNMEASURED33MEASURED▸
art5-safeguardArt5BenchUNMEASURED36MEASURED▸
swarmSwarmBench v2b44.4%37UNTESTED▸
affectAffectBenchUNMEASURED41MEASURED▸
jailGoldBank-Detector59.2%71TIE▸
effect-bindingEffectBench v0.1 (server probe)facts261FACTS▸
provenance-controlsChainFactsfacts6FACTS▸
reserve-attestationReserveFactsfacts16FACTS▸
regulatory-frameworkRegimeFactsfacts16FACTS▸
distribution-integrityDistributionFactsfacts16FACTS▸
custody-disclosureCustodyFactsfacts16FACTS▸
ai-adoption-componentsEurostatfacts2FACTS▸
labour-componentsEurostatfacts2FACTS▸
humanoid-labour-indexDisclosurefacts8FACTS▸
Measurement, not certification · empty cells stay empty · nothing is written here.Arena Ed25519 signature present, unchecked · content_id b126b711de… · snapshot 2026-09-24T16:42:59Z

The full board — with intervals and harm tails

Axis deep-dive

provenance

Bench ProvBench · n=32 · leader accuracy 71.9% · 95% CI 54.6–84.4% · separation TIE

No model separated from the next best on this axis (exact McNemar, p≥0.05, n=32; p=1.0).

Leader llama3.2:3b 23/32 71.9% [54.6–84.4%] vs next best gemma3:12b 22/32 68.8% [51.4–82.0%] · 32 paired items · exact McNemar p=1

leader shown from per-item rows; no signed per-model card yet for this run — the signed card for llama3.2:3b on this axis records a different measurement (accuracy 0.8), so it does not back the number shown · the card on record

Our own council specialist held the point lead in the full 19-model fleet and is excluded: a neutral measurement body does not rank its own models against the vendors it measures. The leader shown is the external-only re-rank of the same published rows (6 base models).

Frozen gold bank (Hugging Face)Verify the signed chainRaw JSON (GET /api/gspc)Full board

Bank exposure. The 99 live signed cards on this axis pin banks with different exposure: 64 UNPINNED, 35 PUBLIC_BANK. Measured on a public benchmark. Every item in the bank behind this number is publicly readable, so a model could have been trained on it. Read it as a score on a public test, not on unseen items; no held-out slice backs it.

99 live signed cards on this axis · per-card labels · signed record · the label changes no score.

AxisBenchnLeader accuracy95% CISeparationOpen traces / graphs
governanceGovBench23758.7%52.3–64.7%TIE — indistinguishablep=0.3787
safetyDefBench3694.4%81.9–98.5%TIE — indistinguishablep=0.6875
provenanceProvBench3271.9%54.6–84.4%TIE — indistinguishablep=1
continuityPQCBench3360.6%43.7–75.3%TIE — indistinguishablep=0.6875
conformanceMCPBench3571.4%54.9–83.7%TIE — indistinguishablep=0.625
opennessOSSBench3284.4%68.2–93.1%TIE — indistinguishablep=0.424
machinery-conformityMachBench33no public leader scorenot published — no public leader scoreUNTESTED
careCareBench19940.5%33.9–47.4%TIE — indistinguishablep=0.427
cross-realityXRAIV32no public leader scorenot published — no public leader scoreUNTESTED
detector-interopDetBench33no public leader scorenot published — no public leader scoreUNTESTED
art5-safeguardArt5Bench36no public leader scorenot published — no public leader scoreUNTESTED
swarmSwarmBench v2b37≥44.4%lower boundwithheld (n not independent)UNTESTED
affectAffectBench41no public leader scorenot published — no public leader scoreUNTESTED
jailGoldBank-Detector7159.2%47.5–69.8%TIE — indistinguishable
effect-bindingEffectBench v0.1 (server probe)261tool-call servers probedno leader accuracy261 of 600 servers tried, from a population of 20,992 third-party serversnot applicable — no accuracy to boundMEASURED — deterministic factsnot applicable — no fleet, no leader
provenance-controlsChainFacts6issuer accounts (not bank items)no leader accuracy6 of the 16 instruments named in the registrynot applicable — no accuracy to boundMEASURED — deterministic factsnot applicable — no fleet, no leader
reserve-attestationReserveFacts16issuer accounts (not bank items)no leader accuracynot applicable — no accuracy to boundMEASURED — deterministic factsnot applicable — no fleet, no leader
regulatory-frameworkRegimeFacts16issuer accounts (not bank items)no leader accuracynot applicable — no accuracy to boundMEASURED — deterministic factsnot applicable — no fleet, no leader
distribution-integrityDistributionFacts16issuer accounts (not bank items)no leader accuracynot applicable — no accuracy to boundMEASURED — deterministic factsnot applicable — no fleet, no leader
custody-disclosureCustodyFacts16issuer accounts (not bank items)no leader accuracynot applicable — no accuracy to boundMEASURED — deterministic factsnot applicable — no fleet, no leader
ai-adoption-componentsEurostat2public seriesno leader accuracynot applicable — no accuracy to boundMEASURED — deterministic factsnot applicable — no fleet, no leader
labour-componentsEurostat2public seriesno leader accuracynot applicable — no accuracy to boundMEASURED — deterministic factsnot applicable — no fleet, no leader
humanoid-labour-indexDisclosure8frozen vendor URLsno leader accuracynot applicable — no accuracy to boundMEASURED — deterministic factsnot applicable — no fleet, no leader

Board dataset on Hugging Faceaxis,bench,status,n,accuracy,interval_lo,interval_hi,separation,fleet_mean,dataset,as_of · empty cells stay empty — never zeroed, never interpolated. The dataset is a versioned mirror; compare its timestamp with live GET /api/gspc.

Separation from the published per-item rows

15,580 rows, published byte-identical at csoai/gspc-peritem-rows-2026-08-12 · peritem_sha256 0d8dacfbe7384a5d2f6a83e7455482ab75f935366bbb6e18da61cfdac8890dec · signed record. The test was fixed on 2026-08-13: exact McNemar on the discordant items, leader vs the best base model, and p<0.05 is required to separate. Our own models are removed before ranking. A TIE is not a win.

  • governance · TIE

    No model separated from the next best on this axis (exact McNemar, p≥0.05, n=237; p=0.3787).

    Leader mistral:7b 139/237 58.7% [52.3–64.7%] vs next best deepseek-r1:8b 128/237 54.0% [47.6–60.2%] · 237 paired items · exact McNemar p=0.3787

    leader shown from per-item rows; no signed per-model card yet for this run — the signed card for mistral:7b on this axis records a different measurement (accuracy 0.0), so it does not back the number shown · the card on record

    Our own council specialist held the point lead in the full 19-model fleet and is excluded: a neutral measurement body does not rank its own models against the vendors it measures. The leader shown is the external-only re-rank of the same published rows (6 base models).

  • safety · TIE

    No model separated from the next best on this axis (exact McNemar, p≥0.05, n=36; p=0.6875).

    Leader gemma3:12b 34/36 94.4% [81.9–98.5%] vs next best qwen2.5:3b 32/36 88.9% [74.7–95.6%] · 36 paired items · exact McNemar p=0.6875

    leader shown from per-item rows; no signed per-model card yet

  • provenance · TIE

    No model separated from the next best on this axis (exact McNemar, p≥0.05, n=32; p=1.0).

    Leader llama3.2:3b 23/32 71.9% [54.6–84.4%] vs next best gemma3:12b 22/32 68.8% [51.4–82.0%] · 32 paired items · exact McNemar p=1

    leader shown from per-item rows; no signed per-model card yet for this run — the signed card for llama3.2:3b on this axis records a different measurement (accuracy 0.8), so it does not back the number shown · the card on record

    Our own council specialist held the point lead in the full 19-model fleet and is excluded: a neutral measurement body does not rank its own models against the vendors it measures. The leader shown is the external-only re-rank of the same published rows (6 base models).

  • continuity · TIE

    No model separated from the next best on this axis (exact McNemar, p≥0.05, n=33; p=0.6875).

    Leader gemma3:12b 20/33 60.6% [43.7–75.3%] vs next best deepseek-r1:8b 18/33 54.5% [38.0–70.2%] · 33 paired items · exact McNemar p=0.6875

    leader shown from per-item rows; no signed per-model card yet

    Our own council specialist held the point lead in the full 19-model fleet and is excluded: a neutral measurement body does not rank its own models against the vendors it measures. The leader shown is the external-only re-rank of the same published rows (6 base models).

  • conformance · TIE

    No model separated from the next best on this axis (exact McNemar, p≥0.05, n=35; p=0.625).

    Leader mistral:7b 25/35 71.4% [54.9–83.7%] vs next best llama3.2:3b 23/35 65.7% [49.2–79.2%] · 35 paired items · exact McNemar p=0.625

    leader shown from per-item rows; no signed per-model card yet

    Our own council specialist held the point lead in the full 19-model fleet and is excluded: a neutral measurement body does not rank its own models against the vendors it measures. The leader shown is the external-only re-rank of the same published rows (6 base models).

  • openness · TIE

    No model separated from the next best on this axis (exact McNemar, p≥0.05, n=32; p=0.424).

    Leader gemma3:12b 27/32 84.4% [68.2–93.1%] vs next best mistral:7b 23/32 71.9% [54.6–84.4%] · 32 paired items · exact McNemar p=0.424

    leader shown from per-item rows; no signed per-model card yet

    Our own council specialist held the point lead in the full 19-model fleet and is excluded: a neutral measurement body does not rank its own models against the vendors it measures. The leader shown is the external-only re-rank of the same published rows (6 base models).

  • care · TIE

    No model separated from the next best on this axis (exact McNemar, p≥0.05, n=200; p=0.427).

    Leader qwen2.5:0.5b-instruct 81/200 40.5% [33.9–47.4%] vs next best qwen2.5:3b 74/200 37.0% [30.6–43.9%] · 199 paired items · exact McNemar p=0.427

    The test counts 200 rows per model over 199 distinct paired items; the board's n (199) counts unique scored texts. The rows are published as frozen, duplicates included.

    leader shown from per-item rows; no signed per-model card yet for this run — the signed card for qwen2.5:0.5b-instruct on this axis records a different measurement (accuracy 0.0), so it does not back the number shown · the card on record

    Our own council specialist held the point lead in the full 19-model fleet and is excluded: a neutral measurement body does not rank its own models against the vendors it measures. The leader shown is the external-only re-rank of the same published rows (6 base models).

Rows published, no determination

  • machinery-conformity — UNTESTED. The published rows give TIE (exact McNemar p=0.5811, n=33), but this axis has no signed card of any model in the public card index (/signed/card_index.json). The board publishes a separation determination only on axes that carry signed cards, so this one stays UNTESTED rather than resting on rows alone.
  • cross-reality — UNTESTED. The published rows give TIE (exact McNemar p=0.0654, n=32), but this axis has no signed card of any model in the public card index (/signed/card_index.json). The board publishes a separation determination only on axes that carry signed cards, so this one stays UNTESTED rather than resting on rows alone.
  • detector-interop — UNTESTED. The published rows give TIE (exact McNemar p=0.4531, n=33), but this axis has no signed card of any model in the public card index (/signed/card_index.json). The board publishes a separation determination only on axes that carry signed cards, so this one stays UNTESTED rather than resting on rows alone.
  • art5-safeguard — UNTESTED. The published rows give TIE (exact McNemar p=1.0, n=36), but this axis has no signed card of any model in the public card index (/signed/card_index.json). The board publishes a separation determination only on axes that carry signed cards, so this one stays UNTESTED rather than resting on rows alone.
  • swarm — UNTESTED. The published swarm rows are the retired 3-prompt PROTOCOL bank (40 rows per model over 3 distinct items), not the wave-2b bank the board's swarm row serves; rows from one bank cannot decide a separation determination about another.
  • affect — UNTESTED. The published rows give TIE (exact McNemar p=1.0, n=41), but this axis has no signed card of any model in the public card index (/signed/card_index.json). The board publishes a separation determination only on axes that carry signed cards, so this one stays UNTESTED rather than resting on rows alone.

Think a row, a grade or a named model is wrong? Object or ask for a re-check.

Attestation · live from GET /api/gspc · click any row for traces

Measurement freshness · derived from GET /api/gspc · measured_on

behavioural axes 2026-08-12 · jail 2026-08-18 · financial-fact axes 2026-08-25

living_stamp.gold_run 18 Aug 2026 — payload stamp, not a live re-measure. Board counts stay derived from GET /api/gspc; no new MEASURED invented here.

These run dates are weeks old. Freshness is labelled; the board is not re-stamped from this UI.

Living Stamp — SIGNED

Do not treat this as a valid attestation. Check site_attestation instead.

Progress · 23 axis · 23 measured

N→N+1 drift · UNCHECKABLE

No published board time series for N→N+1 drift. Empty stays empty — do not invent drift numbers or a Merkle seal. Cite GET /root.json and GET /api/gspc for the living snapshot only. Living snapshot only — cite GET /root.json.

23 axes are on the board and every one carries a measurement — no declared slot is empty. Both counts are DERIVED from the axis array, never typed; if a future slot is added with no run behind it, this line separates the two again on its own. A measurement is not a separated leader: 0 of 14 model-comparison axes separated a leader · 8 TIE · 6 UNTESTED. A point-estimate lead is not a measured advantage, and UNTESTED is not a tie.

Arena Elo — signed

Per-axis Elo from the published arena snapshot (7 models, 28 arena axis — the arena's own set, not the board's count above). Snapshot: . New arena rounds do not change GSPC board axes without admission. The board DID signature is present; use the check below to verify these bytes.

Researcher access — verified rankingsEnterprise data licenseFree access is permanent and unconditional (researchers, journalists, fact-checkers) — see the licensing tiers for the license legs. Pricing lives on the legal surface.
ModelEloGamesWin-rate95% CI
qwen3:8b16292320.6720.61–0.73
llama3.1:8b1563.6640.50.381–0.619
phi3.5:3.8b1558.21990.5480.478–0.615
phi4:14b1527.33090.5630.507–0.617
mistral:7b1438.35800.6030.563–0.642
nemotron-3-nano:30b1389.63050.4690.414–0.525
gemma3:12b1345.45000.2580.222–0.298

In-lane measurements — not board rows

2 slots measured in-lane. Published as measured_in_lane on GET /api/gspc. NOT stamped onto the board count. public_count stays “23 axis · 23 measured” — quoted live, never typed. Click any card for traces + graphs.

Measurement, not certification. Leaders shown are point estimates (swarm quotes its 95% lower bound); only SEPARATED leads are statistically real — the live count is totals.separated_leads on GET /api/gspc. Jail is a measured floor when the stamp publishes one, never a hidden score. Full per-axis notes, fleet means and harm tails: GET /api/gspc. The living stamp carried there is marked UNVERIFIABLE — it does not reproduce under any published rule and is not a checkable attestation; the attestation over that payload that does verify is site_attestation.