NodusLab
Conformance kit

The claim is a hash, so you can break it

Everything on this site reduces to one falsifiable statement: under the numeric contract, this model produces the digest 202f51a312e5c0e685354355 on any conforming backend. You do not have to take that on trust.

No GPU, no network, no account. If your hardware breaks the contract, that is the more interesting result and the programme wants the report.

Two commands

Run it

0.3 sthe portable kit — numpy only
$ git clone https://github.com/Nodus-Lab/nodus-verifiable-compute
$ cd nodus-verifiable-compute
$ pip install numpy

$ python3 conformance/check.py \
        --reference conformance/reference.json

Reports where your backend stops being exact, and whether an integer matmul and a full integer layer reproduce the reference bit for bit.

8 sthe real model — 942 MB, digest-pinned
$ ./models/fetch.sh

$ python3 conformance/check_qwen.py --backend numpy \
        --reference 202f51a312e5c0e685354355

  BIT-IDENTICAL to reference

…or it does not. On an NVIDIA box, --backend cuda runs the same check through torch with TF32 left on.

What it checks

Four tests, in order of severity

T1

Operand bound

Finds the bit width at which your backend stops producing exact integer products. On Accelerate and OpenBLAS nothing fails below 2²²; on a Metal or TF32 matmul it stops at 2¹¹.

T2

Integer matmul

Reproduces the reference integer matrix product bit for bit, under the tiling bound the contract declares.

T3

Full integer layer

Softmax, GELU, layernorm and requantisation as lookup tables — the whole transformer layer, checked against the reference digest.

T4

The real model

Qwen2.5-0.5B over the first 96 tokens of wikitext-2, hashed to a single digest you can compare against every machine in the table.

Environment

Setup

$ python3 -m venv .venv && ./.venv/bin/pip install numpy psutil pypdf

That is enough for the conformance kit and most experiments. The zkVM baselines additionally need SP1 6.5.0 — curl -L https://sp1up.succinct.xyz | bash && sp1up. Model weights are fetched rather than vendored, and each experiment’s methodology.md carries its own reproduction commands.

Full text

conformance/README.md

One script, two dependencies (python3, numpy), no GPU and no network. It checks whether a machine conforms to the numeric contract in ../docs/numeric-contract.md.

python3 check.py --emit-reference > reference.json    # on a reference machine
python3 check.py --reference reference.json           # on the machine under test

reference.json in this folder was generated on an Apple M5 (macOS 26.2, NumPy 2.5.2, Accelerate). Everything is derived from fixed seeds, so no data moves between machines.

What the four tests do

test what a failure means
T1 sweeps operand magnitude with the accumulator held small the backend rounds matmul inputs — NC-3b's bound must be tightened to whatever this reports
T2 exact int8 matmul, hashed the contract does not hold here
T3 float32 matmul nothing — float results are supposed to differ; the number is the divergence the contract removes
T4 a whole integer layer: matmul, requantise, LUT GELU, matmul, integer layernorm the contract does not hold for a realistic pipeline

T1 is the one worth watching. It deliberately restricts the second operand to {-1,0,1} and the reduction length to 4, so the accumulator keeps ~22 bits of headroom and any failure is an operand-precision limit rather than an accumulator one. That distinction is what E018 got wrong on Apple's Metal GPU, where the real bound turned out to be 2^11 rather than the 2^24 accumulator bound the contract originally assumed.

Real-model check

check_qwen.py runs Qwen2.5-0.5B (24 layers, vocab 151,936) under NC-0.5 and hashes the logits. numpy only — it parses safetensors itself and the token ids are hardcoded, so nothing about the input or the model loading can differ between machines.

python3 check_qwen.py --backend numpy --reference 202f51a312e5c0e685354355
python3 check_qwen.py --backend cuda  --reference 202f51a312e5c0e685354355   # needs torch
machine real-model logits hash
Apple M5, CPU (Accelerate, ARM) 202f51a3… reference
Apple M5, Metal GPU 202f51a3… identical
Intel Xeon (OpenBLAS, Debian 13) 202f51a3… identical
Intel Xeon (OpenBLAS, Ubuntu 22.04) 202f51a3… identical
NVIDIA L4, CUDA 12.4, TF32 on 202f51a3… identical

Results so far

machine T1 bound contract holds?
Apple M5, CPU (Accelerate, ARM) 2^22 (test ceiling) reference
Apple M5, Metal GPU 2^11 yes, once operands are kept ≤ 2^11
Intel Xeon 8581C (OpenBLAS, x86-64, Linux) 2^22 (test ceiling) yes — bit-identical
NVIDIA L4, TF32 off 2^19 (test ceiling) yes — bit-identical
NVIDIA L4, TF32 on 2^11 yes — bit-identical

The last row was a prediction before it was a measurement. E021 found neither CPU had an operand limit, which made the bound look like a property of GPU matmul paths; TF32 carries a 10-bit explicit mantissa, 11 with the implicit bit, so E022 predicted 2^11 for NVIDIA — the same figure as Metal, from an unrelated vendor's design — and measured exactly that. Two of the CPU bounds are test ceilings rather than limits: those backends were still exact where the sweep stopped.