Council of AI — position paper

What an IVO's evidence should look like

California SB 813 (Chapter 179, signed September 9, 2026) creates a state statutory framework for Independent Verification Organizations (IVOs) assessing AI systems. The Government Operations Agency must develop IVO designation criteria by January 1, 2028. This paper maps our evidence practice to the four criteria in §8898.1(c)(2).

This is a CSOAI editorial act. It is not a California Government Operations Agency endorsement, not legal advice, and not a claim that CSOAI is a designated IVO. IVO designation requires GovOps application.

The statutory criteria

SB 813 §8898.1(c)(2) requires the agency to consider, at a minimum, whether an IVO meets four criteria. The standards are unwritten — GovOps must develop them through stakeholder working groups. This paper proposes evidence standards for each criterion, grounded in what we actually publish and what a stranger can independently verify.

Source: SB 813, Chapter 179, California Legislature

A

Assessing risks and identifying metrics/methodologies

§8898.1(c)(2)(A)

An IVO must demonstrate expertise in assessing the risks posed by an AI system or model and identifying the metrics and methodologies that form the basis for that assessment.

Our evidence practice:

  • Frozen, published instruments: each GSPC axis uses a versioned, hash-pinned bank of test items. The instrument is published before any model is graded against it, and the bank never changes mid-measurement.
  • Deterministic grading: no model judges another model. Every score is computed by a deterministic grader (keyword match, exact match, regex, numeric comparison) against the frozen bank. The grader code is published.
  • Per-item evidence: every signed measurement card carries the full item-level evidence — the model's response to each test item, the expected answer, and the grader's deterministic verdict. A stranger can re-grade every item from the published bytes.
  • Three-state honesty: every cell is MEASURED (a real run exists), UNMEASURED (no run — the slot exists but is empty), or INDEXED (the subject is known but not yet probed). Zero-filling is explicitly forbidden.
Verify: curl -s https://councilof.ai/api/gspc | jq '.totals'
B

Personnel with sufficient technical expertise

§8898.1(c)(2)(B)

An IVO must employ or engage personnel with sufficient technical expertise to conduct its assessments.

Our evidence practice:

  • The measurement mill runs on dedicated GPU compute with a GSPC worker that grades models against frozen banks. Its state is not typed here: the worker's own health is read live at /api/worker and shown below.
  • The public root is maintained by a signed, multi-witness process: an Ed25519 signature, an OpenTimestamps proof and a Rekor transparency-log entry, with XRPL and EAS anchors recorded as NOT_YET until they exist. The current witness states are read live from /interop/root-witness-latest.json and shown below — a pending OpenTimestamps proof is not a Bitcoin timestamp.
  • The corrections ledger at /api/corrections records every error we have made in our own published figures, how it was caught, and the fix. Corrections are permanent — we never delete or edit a correction.
  • The measurement methodology is published at /methodology. The frozen banks are published on Hugging Face under the csoai organisation. The card verification code is published at /signed/verify-card.mjs.
  • /api/worker → status OFFLINE · worker state not published · read_at not published
  • root as_of 2026-09-28T07:32:07Z · OpenTimestamps STAMPED_PENDING_BITCOIN (Bitcoin blocks: none yet) · Rekor WITNESSED (logIndex 2981565650) · XRPL memo NOT_YET · EAS NOT_YET
Verify: curl -s https://councilof.ai/api/worker | jq '{status, worker.state}'
C

Managing conflicts of interest

§8898.1(c)(2)(C)

An IVO must identify and manage potential conflicts of interest. Payment from the assessed party is allowed at reasonable market rates, but payment must not be conditioned on the results of the assessment.

Our evidence practice:

  • No issuer-pays: the entities we measure never pay for their measurement. The board's own totals are read live at /api/gspc (its count line is derived from the axis array, never typed) — none of the measured parties paid for or influenced their measurement.
  • Revenue comes from machine-readable data delivery (x402 doors), not from favourable findings. The settlement ledger is public at /api/revenue; its one_number (distinct non-self payers) is read live and shown below, never typed.
  • Self-settlements are explicitly excluded from revenue counts. A wallet we control paying us is recorded for audit but is neither revenue nor a buyer.
  • Corrections are free forever. A measured party can request a correction at no cost, and the correction is published on the same ledger regardless of who requested it.
  • The independence-conditions page (/evaluator-access) publishes our five conditions: no lab money, no gag clauses, methods on the card, corrections ledger governs, access/redaction terms published.
  • /api/revenue → one_number MEASURED: 1 distinct non-self payer(s) · 1 settlement(s) · 22 self-settlement(s) excluded
Verify: curl -s https://councilof.ai/api/revenue | jq '.one_number'
D

Independence from the party being assessed

§8898.1(c)(2)(D)

An IVO must maintain independence from the party being assessed — no operational or management dependence, free from the assessed party's control in reaching conclusions.

Our evidence practice:

  • We do not require, request, or receive API access from the entities we measure. Every measurement uses publicly available endpoints, published models, and public data. If a model is unreachable, the cell reads UNMEASURED — we never fill it from a private channel.
  • The measurement card's signature is over the exact bytes that were graded. The card is content-addressed (sha256 of the canonical body), so altering any byte invalidates the signature. The assessed party cannot modify a card after signing.
  • The public root commits to the full leaf list. Every card in the root is independently verifiable — fetch the card, recompute the hash, check the signature against the published DID key. The assessed party has no role in this process.
  • When we get something wrong, the corrections ledger records it publicly. The assessed party does not approve or suppress corrections. The correction is published whether the error favours or disfavours the measured entity.
Verify: curl -s https://councilof.ai/root.json | jq '{card_count, as_of, merkle_root}'

The five conditions

Before accepting access to any AI system for evaluation, we publish these conditions at /evaluator-access:

  1. No money from the developer being evaluated.
  2. No gag clauses — the right to publish findings without editorial control.
  3. Methods on the card — the measurement methodology is published with every result.
  4. Corrections ledger governs — errors are published publicly and permanently.
  5. Access and redaction terms published — when we receive access, the terms are visible.

These conditions are not new — they are what SB 813 §8898.1(c)(2)(C) and (D) already require. We publish them so the rulemaking has a concrete reference.

What this is not

  • Not a claim that CSOAI is a designated IVO. Designation requires GovOps application.
  • Not legal advice. This is a technical evidence-standards proposal.
  • Not a certification. We measure, we do not certify. A grade is never sold.
  • Not an endorsement by the State of California. SB 813 §8898.4(a)(2) explicitly states that publication of criteria does not constitute state recommendation.

Issued by CSOAI Ltd (England & Wales, Companies House 16939677), 3rd Floor, 86–90 Paul Street, London EC2A 4NE. This position paper is published under CC BY 4.0. Corrections: /api/corrections.

Statute: SB 813, Chapter 179 (California, signed Sep 9, 2026). As_of: Sep 15, 2026.