Preprint · September 2026

Same model, same prompts, different answers

Item-level cross-hardware reproducibility of LLM evaluation results

Nicholas Templeman, CSOAI Ltd (Council of AI), United Kingdom.

Measurement, not certification. This work measures whether a published evaluation result survives a change of machine. It scores no model, company, GPU vendor or cloud provider, and it does not say which runtime is right. None of the models is a CSOAI model. Everything below is an aggregate over all of them.

Abstract

Evaluation results for large language models are usually published as totals: a model answered k of n items correctly. We ask whether those results are a property of the model, the prompts and the grader, or also of the machine that ran them. We re-ran 154 published evaluation results (“cards”), covering 11 open-weight models on 14 evaluation axes and 9,757 items, on a second runtime: a Kaggle Tesla T4 instead of the RunPod RTX 3090 that produced them, with the same model manifest digest, the same prompts and item banks (pinned by SHA-256), the same instrument, and fixed greedy decoding.

Raw outputs were byte-identical on 84.4% of items; grades agreed on 97.3%. The 260 grade flips went in both directions (122 up, 138 down) and so mostly cancel in totals: 62 cards reproduced their total exactly, but on 11 of those an item's grade had changed. Only 54 cards were grade-identical item by item. A further 83 items kept their grade but swapped between “parse error” and “wrong answer”, which changes the denominator of the reported accuracy. Re-running the T4 job on the same T4 was not byte-deterministic for 20 of 150 cards. Grade flips arise overwhelmingly when the two runtimes diverge on the first generated token.

We argue that evaluation results should be admitted for publication only on item-level agreement across independent runtimes, that the runtime must be reported with every result, and that totals alone cannot establish reproducibility. The per-item data for all three runs, the runtime declarations and the signed Merkle roots that bind them are public.

Key figures

RTX 3090 against Tesla T4, same digest, prompts and decode settings. Read from paper/numbers.json at dataset commit 57f1166f; SHA-256 computed in your browser: 7a63270ae58dd6fb2653e48d007db068a73904ee81c28ba25e36cf1bbbf90dc9.

Compared
154 cards · 11 models · 14 axes · 9,757 items
Raw output byte-identical
84.4% of items (8,231 of 9,757)
Same grade on both runtimes
97.3% of items (9,497 of 9,757)
Grade flips
260 (2.66%, 95% CI 2.36–3.00%): 122 up, 138 down; sign test p = 0.35
Totals reproduced exactly
62 of 154 cards, of which 11 hid a changed item
Same grade on every item
54 of 154 cards
Parse error ↔ wrong answer swaps
83 items on 31 cards (grade unchanged, denominator changed)
Same T4, run twice
20 of 150 cards not byte-identical; 13 changed their totals

Under an admission rule that requires the same totals and the same grade on every item, 51 of 154 cards pass, against 62 under a totals-only rule. For items whose grade flipped, the median length of the common prefix of the two outputs is 0 characters: the runtimes disagree on the first token, which in a label task is the answer.

What this does not show

  • Not a ranking. Flips run both ways in near-equal numbers; neither runtime is “better”, and no model, company or vendor is scored.
  • Not a cause. GPU architecture, GPU count, driver, flash-attention setting and possibly the server build all differ between the runtimes, and several primary-side values were not recorded. Effects are attributed to the runtime as a whole.
  • Not other engines or larger models. One serving engine (Ollama with llama.cpp), single-request serving, small to mid-size quantized open-weight models.
  • Not a third runtime. A Tesla P100 reproduction was queued while the analysis was prepared. It has not been admitted or signed, and no result from it is reported.
  • Not independent. All runs are ours. The data and signed roots let anyone re-check the comparison; a reproduction by a different operator would be stronger evidence.

How to verify

Download the dataset, then run its offline verifier:

pip install -U cryptography huggingface_hub
hf download csoai/cross-hardware-reproducibility --repo-type dataset --local-dir xhw
cd xhw && python3 verify.py

It checks the capsule hashes, the RFC 6962 Merkle roots, the Ed25519 signatures against the board key in csoai.org/.well-known/did.json, that the per-item data matches what the signed declarations bind, that recounting it reproduces every card's declared counts, and every file against manifest.jsonl. It prints ALL CHECKS PASS only if all of them do.

To regenerate every number and table in the paper from the dataset alone:

python3 code/figures/from_dataset.py --dataset . --work /tmp/xhw
python3 code/figures/make_figures.py --decl /tmp/xhw/decl-nemo \
  --decl /tmp/xhw/decl-batch2 --items /tmp/xhw/items
cmp code/build/numbers.json paper/numbers.json

Signed roots

The runtime declarations are bound into two signed batches of measurement capsules, signed under did:web:csoai.org#board-attestation-1. Read from the dataset's capsule records:

Batch 1 (14 capsules)
5ae00c1f4c2ed290fb91206f29372d0df3e61478c568d4797855502e751ebff6
Batch 2 (140 capsules)
78226433e9e433b311e067ee56fcbaf415f6dc57301254f3bb4922a0b8053cf6

The same kind of record is published daily in the measurement capsules index.

Licence and citation

Paper and data: CC BY 4.0. Model outputs are included as research evidence; each model remains under its own licence. The analysis code and draft text were produced with AI assistance; every number is recomputed from the published per-item data by the included scripts, and the author is responsible for the content.

@misc{templeman2026crosshardware,
  title  = {Same model, same prompts, different answers: item-level
            cross-hardware reproducibility of LLM evaluation results},
  author = {Templeman, Nicholas},
  year   = {2026},
  note   = {CSOAI Ltd (Council of AI). Preprint and data:
            https://huggingface.co/datasets/csoai/cross-hardware-reproducibility}
}

Corrections and objections

If you think something in the paper or the data is wrong, or you want a specific output withheld, see how to object or email nicholas@csoai.org. A correction is published as a new, linked version and dated in the corrections ledger; signed files are never edited in place.