Research & Transparency

We publish our wrong turns, not just our wins.

A governance company that only publishes flattering results is not credible. Below is a running, honest account of research findings behind our internal research lineage — including claims we made, then corrected or retracted after closer scrutiny. Confirmed findings link to the technical detail; retracted ones explain exactly what was wrong and why.

Preprint · September 2026Same model, same prompts, different answersItem-level cross-hardware reproducibility of LLM evaluation results: the same models, prompts and settings re-run on a second GPU runtime and compared item by item, with the full per-item dataset and an offline verifier.
Lineage diversity beats topology shape
CONFIRMED

Across three independent measurement passes, diverse model lineages (e.g. Qwen + Llama + DeepSeek + Gemma + Mistral) consistently outperformed identical-lineage configurations of the same size on a governance-quality battery. The gap between diverse and identical configurations was roughly 6× larger than the gap between different topology shapes (ring vs. pyramid vs. triangle). Practical upshot: which models you combine matters far more than how you arrange them.

A one-size 'containment = 1.00 under attack' claim
RETRACTED & CORRECTED

An early adversarial test reported perfect containment (1.00) even under a simulated majority attack. On review, the test was tautological — it hard-gated on a condition that was true of every attack by construction, so it proved nothing about actual robustness. We retracted the claim and reran a corrected test that separates "obvious" breaches (which are gated to zero regardless of votes, by design) from "laundered" harm that looks confident but is actually harmful (where the real result is 58–79% containment under 2–3 compromised voting nodes — real, useful, and explicitly not perfect).

A fabricated citation, caught and removed
RETRACTED

An internal research note once cited a named news publication in support of a claim. On review the citation did not exist — it was an invented reference. It has been struck from every downstream document. The underlying technical claim it was attached to stands on its own on running code and measured results, not on any citation, and is unaffected.

An acceptance claim, downgraded
CORRECTED

A research note referenced an academic conference acceptance. This was overstated — the correct framing is that submission to that venue is a target, not a confirmed acceptance. The claim has been downgraded to aspirational language throughout.

Independent cross-verification of a research pass
PARTIALLY VERIFIED

A separate research pass reported that diverse 5-model configurations won on both a clean-data metric and a containment metric, using a distinct code path from our own measurement. We could independently verify only one of the two measurement passes behind this finding on the tree at the time it was reported; the second is treated as a consistent but unverified secondary report until its source code is confirmed on disk. We flag this rather than quietly merging both into one number.

A capability benchmark we do not yet have
OPEN — NOT DONE

We have not run a head-to-head capability benchmark (e.g. GSM8K, MMLU) of our fine-tuned model against frontier models. The governance-topology results above measure decision-quality, safety, and cost under a stated error model — they are a real, useful, and different thing from a capability benchmark, and we do not present one as a substitute for the other. This is an open item, not a hidden one.

Why we do this

If we cannot survive our own adversarial review, we have no business selling assurance to anyone else. Every finding above was caught and corrected before — or immediately after — it appeared in any customer-facing material. That correction loop, not a flawless research record, is the actual credibility signal.

Updated as findings are confirmed or corrected. Last reviewed 2026-07-12.