The claim is a hash, so you can break it
Everything on this site reduces to one falsifiable statement: under the numeric contract, this model produces the digest 202f51a312e5c0e685354355 on any conforming backend. You do not have to take that on trust.
No GPU, no network, no account. If your hardware breaks the contract, that is the more interesting result and the programme wants the report.
Run it
$ git clone https://github.com/Nodus-Lab/nodus-verifiable-compute
$ cd nodus-verifiable-compute
$ pip install numpy
$ python3 conformance/check.py \
--reference conformance/reference.json
Reports where your backend stops being exact, and whether an integer matmul and a full integer layer reproduce the reference bit for bit.
$ ./models/fetch.sh
$ python3 conformance/check_qwen.py --backend numpy \
--reference 202f51a312e5c0e685354355
BIT-IDENTICAL to reference
…or it does not. On an NVIDIA box, --backend cuda runs the same check through torch with TF32 left on.
Four tests, in order of severity
Operand bound
Finds the bit width at which your backend stops producing exact integer products. On Accelerate and OpenBLAS nothing fails below 2²²; on a Metal or TF32 matmul it stops at 2¹¹.
Integer matmul
Reproduces the reference integer matrix product bit for bit, under the tiling bound the contract declares.
Full integer layer
Softmax, GELU, layernorm and requantisation as lookup tables — the whole transformer layer, checked against the reference digest.
The real model
Qwen2.5-0.5B over the first 96 tokens of wikitext-2, hashed to a single digest you can compare against every machine in the table.
Setup
$ python3 -m venv .venv && ./.venv/bin/pip install numpy psutil pypdf
That is enough for the conformance kit and most experiments. The zkVM baselines additionally need SP1 6.5.0 — curl -L https://sp1up.succinct.xyz | bash && sp1up. Model weights are fetched rather than vendored, and each experiment’s methodology.md carries its own reproduction commands.
conformance/README.md
One script, two dependencies (python3, numpy), no GPU and no network. It checks
whether a machine conforms to the numeric contract in
../docs/numeric-contract.md.
python3 check.py --emit-reference > reference.json # on a reference machine
python3 check.py --reference reference.json # on the machine under test
reference.json in this folder was generated on an Apple M5 (macOS 26.2, NumPy
2.5.2, Accelerate). Everything is derived from fixed seeds, so no data moves
between machines.
What the four tests do
| test | what a failure means | |
|---|---|---|
| T1 | sweeps operand magnitude with the accumulator held small | the backend rounds matmul inputs — NC-3b's bound must be tightened to whatever this reports |
| T2 | exact int8 matmul, hashed | the contract does not hold here |
| T3 | float32 matmul | nothing — float results are supposed to differ; the number is the divergence the contract removes |
| T4 | a whole integer layer: matmul, requantise, LUT GELU, matmul, integer layernorm | the contract does not hold for a realistic pipeline |
T1 is the one worth watching. It deliberately restricts the second operand to {-1,0,1} and the reduction length to 4, so the accumulator keeps ~22 bits of headroom and any failure is an operand-precision limit rather than an accumulator one. That distinction is what E018 got wrong on Apple's Metal GPU, where the real bound turned out to be 2^11 rather than the 2^24 accumulator bound the contract originally assumed.
Real-model check
check_qwen.py runs Qwen2.5-0.5B (24 layers, vocab 151,936) under NC-0.5 and
hashes the logits. numpy only — it parses safetensors itself and the token ids
are hardcoded, so nothing about the input or the model loading can differ
between machines.
python3 check_qwen.py --backend numpy --reference 202f51a312e5c0e685354355
python3 check_qwen.py --backend cuda --reference 202f51a312e5c0e685354355 # needs torch
| machine | real-model logits hash | |
|---|---|---|
| Apple M5, CPU (Accelerate, ARM) | 202f51a3… |
reference |
| Apple M5, Metal GPU | 202f51a3… |
identical |
| Intel Xeon (OpenBLAS, Debian 13) | 202f51a3… |
identical |
| Intel Xeon (OpenBLAS, Ubuntu 22.04) | 202f51a3… |
identical |
| NVIDIA L4, CUDA 12.4, TF32 on | 202f51a3… |
identical |
Results so far
| machine | T1 bound | contract holds? |
|---|---|---|
| Apple M5, CPU (Accelerate, ARM) | 2^22 (test ceiling) | reference |
| Apple M5, Metal GPU | 2^11 | yes, once operands are kept ≤ 2^11 |
| Intel Xeon 8581C (OpenBLAS, x86-64, Linux) | 2^22 (test ceiling) | yes — bit-identical |
| NVIDIA L4, TF32 off | 2^19 (test ceiling) | yes — bit-identical |
| NVIDIA L4, TF32 on | 2^11 | yes — bit-identical |
The last row was a prediction before it was a measurement. E021 found neither CPU had an operand limit, which made the bound look like a property of GPU matmul paths; TF32 carries a 10-bit explicit mantissa, 11 with the implicit bit, so E022 predicted 2^11 for NVIDIA — the same figure as Metal, from an unrelated vendor's design — and measured exactly that. Two of the CPU bounds are test ceilings rather than limits: those backends were still exact where the sweep stopped.