Preprint · September 2026
Same model, same prompts, different answers
Item-level cross-hardware reproducibility of LLM evaluation results
Nicholas Templeman, CSOAI Ltd (Council of AI), United Kingdom.
Measurement, not certification. This work measures whether a published evaluation result survives a change of machine. It scores no model, company, GPU vendor or cloud provider, and it does not say which runtime is right. None of the models is a CSOAI model. Everything below is an aggregate over all of them.
Abstract
Evaluation results for large language models are usually published as totals: a model answered k of n items correctly. We ask whether those results are a property of the model, the prompts and the grader, or also of the machine that ran them. We re-ran 154 published evaluation results (“cards”), covering 11 open-weight models on 14 evaluation axes and 9,757 items, on a second runtime: a Kaggle Tesla T4 instead of the RunPod RTX 3090 that produced them, with the same model manifest digest, the same prompts and item banks (pinned by SHA-256), the same instrument, and fixed greedy decoding.
Raw outputs were byte-identical on 84.4% of items; grades agreed on 97.3%. The 260 grade flips went in both directions (122 up, 138 down) and so mostly cancel in totals: 62 cards reproduced their total exactly, but on 11 of those an item's grade had changed. Only 54 cards were grade-identical item by item. A further 83 items kept their grade but swapped between “parse error” and “wrong answer”, which changes the denominator of the reported accuracy. Re-running the T4 job on the same T4 was not byte-deterministic for 20 of 150 cards. Grade flips arise overwhelmingly when the two runtimes diverge on the first generated token.
We argue that evaluation results should be admitted for publication only on item-level agreement across independent runtimes, that the runtime must be reported with every result, and that totals alone cannot establish reproducibility. The per-item data for all three runs, the runtime declarations and the signed Merkle roots that bind them are public.
Key figures
RTX 3090 against Tesla T4, same digest, prompts and decode settings. Read from paper/numbers.json at dataset commit 57f1166f; SHA-256 computed in your browser: 7a63270ae58dd6fb2653e48d007db068a73904ee81c28ba25e36cf1bbbf90dc9.
- Compared
- 154 cards · 11 models · 14 axes · 9,757 items
- Raw output byte-identical
- 84.4% of items (8,231 of 9,757)
- Same grade on both runtimes
- 97.3% of items (9,497 of 9,757)
- Grade flips
- 260 (2.66%, 95% CI 2.36–3.00%): 122 up, 138 down; sign test p = 0.35
- Totals reproduced exactly
- 62 of 154 cards, of which 11 hid a changed item
- Same grade on every item
- 54 of 154 cards
- Parse error ↔ wrong answer swaps
- 83 items on 31 cards (grade unchanged, denominator changed)
- Same T4, run twice
- 20 of 150 cards not byte-identical; 13 changed their totals
Under an admission rule that requires the same totals and the same grade on every item, 51 of 154 cards pass, against 62 under a totals-only rule. For items whose grade flipped, the median length of the common prefix of the two outputs is 0 characters: the runtimes disagree on the first token, which in a label task is the answer.
What this does not show
- Not a ranking. Flips run both ways in near-equal numbers; neither runtime is “better”, and no model, company or vendor is scored.
- Not a cause. GPU architecture, GPU count, driver, flash-attention setting and possibly the server build all differ between the runtimes, and several primary-side values were not recorded. Effects are attributed to the runtime as a whole.
- Not other engines or larger models. One serving engine (Ollama with llama.cpp), single-request serving, small to mid-size quantized open-weight models.
- Not a third runtime. A Tesla P100 reproduction was queued while the analysis was prepared. It has not been admitted or signed, and no result from it is reported.
- Not independent. All runs are ours. The data and signed roots let anyone re-check the comparison; a reproduction by a different operator would be stronger evidence.
How to verify
Download the dataset, then run its offline verifier:
pip install -U cryptography huggingface_hub hf download csoai/cross-hardware-reproducibility --repo-type dataset --local-dir xhw cd xhw && python3 verify.py
It checks the capsule hashes, the RFC 6962 Merkle roots, the Ed25519 signatures against the board key in csoai.org/.well-known/did.json, that the per-item data matches what the signed declarations bind, that recounting it reproduces every card's declared counts, and every file against manifest.jsonl. It prints ALL CHECKS PASS only if all of them do.
To regenerate every number and table in the paper from the dataset alone:
python3 code/figures/from_dataset.py --dataset . --work /tmp/xhw python3 code/figures/make_figures.py --decl /tmp/xhw/decl-nemo \ --decl /tmp/xhw/decl-batch2 --items /tmp/xhw/items cmp code/build/numbers.json paper/numbers.json
Signed roots
The runtime declarations are bound into two signed batches of measurement capsules, signed under did:web:csoai.org#board-attestation-1. Read from the dataset's capsule records:
- Batch 1 (14 capsules)
- 5ae00c1f4c2ed290fb91206f29372d0df3e61478c568d4797855502e751ebff6
- Batch 2 (140 capsules)
- 78226433e9e433b311e067ee56fcbaf415f6dc57301254f3bb4922a0b8053cf6
The same kind of record is published daily in the measurement capsules index.
Licence and citation
Paper and data: CC BY 4.0. Model outputs are included as research evidence; each model remains under its own licence. The analysis code and draft text were produced with AI assistance; every number is recomputed from the published per-item data by the included scripts, and the author is responsible for the content.
@misc{templeman2026crosshardware,
title = {Same model, same prompts, different answers: item-level
cross-hardware reproducibility of LLM evaluation results},
author = {Templeman, Nicholas},
year = {2026},
note = {CSOAI Ltd (Council of AI). Preprint and data:
https://huggingface.co/datasets/csoai/cross-hardware-reproducibility}
}Corrections and objections
If you think something in the paper or the data is wrong, or you want a specific output withheld, see how to object or email nicholas@csoai.org. A correction is published as a new, linked version and dated in the corrections ledger; signed files are never edited in place.