External benchmark · 170/170 held-out scenarios · losses published
Measured against the EU AI Act Benchmark — including where we fail.
We ran our engines against the EU-funded AI Act Evaluation Benchmark — all 170 held-out scenarios — and we publish everything: the score, the interval, the per-scenario data, and the caveats. Measured, not claimed.
The number
Run 2026-08-02, $0. Two independent substrates — Apple M4 (local ollama) and NVIDIA T4 (Kaggle) — same frozen split, same harness, same weights. The intervals overlap almost exactly: this is a replicated measurement, not an anecdote.
| Engine (variant) | M4 set-F1 [95% CI] | T4 set-F1 [95% CI] | Exact match | Scenarios |
|---|---|---|---|---|
clan-refusal-gate | 0.153 [0.128, 0.182] | 0.151 [0.127, 0.180] | 0/170 | 170/170 measured, both substrates |
clan-law-refusing | 0.156 [0.131, 0.184] | 0.147 [0.123, 0.176] | 0/170 | 170/170 measured, both substrates |
artefacts: coai-dashboard/benchmark-results/aiact_benchmark/aiact_20260802_071146.json (M4) · coai-dashboard/benchmark-results/aiact_benchmark/aiact_t4_20260802_115217.json (Kaggle T4) — per-scenario scores, CIs, and run metadata included.
What this measures
Honest scope — what the stick is, and whose stick it is.
The task
Given a described AI system, list which EU AI Act articles apply — scored as set-F1 against the benchmark's gold article lists.
The benchmark
AI Act Evaluation Benchmark (davidath), arXiv 2603.09435, data CC-BY-4.0 — an EU-funded measuring instrument. We use it as a stick to measure ourselves; we did not build it and do not maintain it.
The split
Frozen split v1 — held-out selected by hash of the scenario text (sha256 % 2), fixed before any run, reproducible by anyone. 170 of 339 scenarios are held-out; we scored all 170, never a cherry-picked subset.
- The benchmark's gold labels are LLM-generated (upstream disclosure) — we measure agreement with the benchmark, not legal truth.
- Article retrieval is one task. It does not measure compliance judgement, obligation drafting, or deployment safety.
- These are small council-tuned variants (≤4B class). The size ladder (0.5B → 8B, second substrate: Kaggle T4) publishes next; we expect the ordering to change and we will publish that too.
The statistics — why you can trust the interval
BCa bootstrap
2,000 resamples, fixed seed — near-nominal coverage for skewed score distributions at this sample size; plain percentile bootstrap undercovers.
Wilson score interval
For the exact-match proportion — correct behaviour near 0 and 1 where bootstrap fails. This run: 0/170 exact, 95% Wilson [0.000, 0.022], both engines.
Per-scenario scores published
Every per-scenario set-F1 is in the artefact — anyone can recompute our CIs or run paired tests against us.
Three-outcome honesty
A scenario is MEASURED, UNPARSEABLE (scored 0 and counted), or UNMEASURED (infrastructure failure — counted, never folded into the score). This run: 170 measured, 0 unparseable, 0 unmeasured, both engines.
The harness
Open, deterministic, stdlib-only: scripts/eat_aiact_benchmark.py in the coai-dashboard repo. Same prompts, same hashing, same statistics on every substrate — Apple M4 (this result) and Kaggle T4 (landing) — because two independent substrates measuring the same thing is the difference between a result and an anecdote.
No harmonised AI Act standards are published in the OJEU as of August 2026; we anchor on ISO/IEC 23894, TR 24027, EN ISO/IEC 42001 and CEN/CLC/TR 17894, and we say so.
What this page does not claim
- Not a compliance certification. Article retrieval is one task, scored against one benchmark.
- Not legal truth. The gold labels are LLM-generated; we measure agreement with the benchmark.
- Not a leaderboard position. We publish our own engines only; no competitor was run on this page.
- Not final. The second substrate (Kaggle T4) has now replicated these intervals (see table). The size ladder (0.5B → 8B) publishes next — ordering may change, and that will be published too.