NodusLab
Measured results

What the evidence actually costs

Every figure below is measured against the same baseline: running the job again. R_verify = 1 by definition, so anything above that line has to justify itself with something re-execution cannot provide.

E021 · E022 · E023

Five machines, four vendors, one set of bytes

The measurement the rest of the programme rests on: held to the numeric contract, every honest backend tested returns the same logits, bit for bit.

Qwen2.5-0.5B · 24 layers · vocab 151,936 · NC-0.5 · first 96 tokens of wikitext-2E023 · E021 · E022
MachineStackLogits digest
Apple M5 CPUARM · Accelerate · macOS202f51a3…
Apple M5 GPUMetal · MLX202f51a3…
Intel Xeonx86 · OpenBLAS · Debian 13202f51a3…
Intel Xeonx86 · OpenBLAS · Ubuntu 22.04202f51a3…
NVIDIA L4CUDA 12.4 · TF32 enabled202f51a3…
Byte-identical on all five

Two CPU vendors, two GPU vendors, three operating systems, three Python versions — every next-token prediction matching. TF32 was left on deliberately: int8 operands sit four bits inside the contract’s operand bound, so the reduced-precision tensor-core path cannot lose anything. The contract never asks a provider to cripple their hardware, which is the difference between a specification that can be complied with and one that is merely stated.

R_verify

Nine orders of magnitude between two mechanisms

MechanismWorkloadRSource
General-purpose zkVM (SP1)y = 3x + 57 × 10⁴×E001
Freivalds' algorithm8192² matmul0.023E003
Sumcheck, single matmuln = 4096 over F_p0.634E026
Sumcheck, two-layer chainn = 40960.465E026
Re-executionanything1.000baseline
The difference is not implementation quality. A zkVM pays for every load and branch in an execution trace; Freivalds exploits the algebraic structure of the statement. The zkVM figure is R_prove — its verification is flat and cheap, but the proof has to exist first.
E009-pre · E014

Why floating point leaves no threshold to set

Freivalds’ guarantee assumes exact arithmetic. In float it degrades to a tolerance, and a tolerance is the attacker’s budget. Crossing silicon vendors costs 5 × 10⁻⁷; crossing from CPU to GPU on the same machine costs 1,600× more.

Mean relative divergence · real-valued operands

Log scale. Everything left of the hatched band is honest; everything right of it is a cheat.

3.7× — the entire enforcement window
10⁻⁷10⁻⁶10⁻⁵10⁻⁴10⁻³10⁻²4.83 × 10⁻⁷two honest CPUs
different vendors and BLAS
7.70 × 10⁻⁴honest CPU vs honest GPU
same machine
2.84 × 10⁻³a bfloat16 cheat
A threshold must sit inside the band. There is no value in it that admits every honest GPU and rejects every cheat — and admitting GPUs is the entire point of a compute network. E009-pre
The same comparison · integer operands

Every honest backend returns the bit-exact product. The window has no upper edge.

usable window — unbounded
10⁻⁷10⁻⁶10⁻⁵10⁻⁴10⁻³10⁻²0.0 — exactevery honest backend
CPU and GPU alike
1.41 × 10⁻³the same cheat
We expected a supply/security trade-off and measured the opposite. The strictest contract admits every honest backend tested. E014 · E015 · E017 · E018 · E023
The GPU’s imprecision was never about the GPU. It was about representing real numbers.
E016 · E024

The bill we had not counted

Verifying a layer requires the verifier to hold that layer’s intermediates. Against a 0.02 MB answer that is a factor of three thousand, and it makes every configuration lose. Compute was never the binding cost.

49.1 MBintermediates the verifier must receive, against a 0.02 MB answer3,010× the payload · E016
$164/hrthe compute price at which verifying every layer breaks evenDERIVED · a real GPU-hour is about $1
4.2%audit traffic once you challenge one random layer instead of all of thembandwidth flat, compute saved scales with depth

The fix is to audit one random layer rather than all of them: bandwidth stays flat while the compute saved scales with depth. Built and run against the real model — Merkle commit, uniform post-commitment challenge, open, recompute. This is not a proof and we do not call it one: a single-layer cheat escapes with probability 1 − 1/L.

E024 — the layer-audit protocol, run against Qwen2.5-0.5B
honest provider, all 24 layers challengedaccepted
single-layer cheat, 200 trialsdetected 4.5% (1/L = 4.2%)
cheat everywhere, answer the challenge honestlyrejected by the commitment
audit traffic vs verifying every layer4.2%
For a provider skipping m of L layers, both the saving and the detection probability are m/L, so they cancel. A stake equal to the job’s compute value makes cheating negative-expected-value, at any depth and any level of greed.
E027 — a conclusive fraud proof by bisection

Two parties who disagree binary-search to the first layer whose digest differs; an arbiter recomputes that layer alone.

E024 random auditE027 bisection
bandwidth on the wire4.2%4.24%
arbiter's compute4.2%4.2% — one layer of 24
detection4.2%certain
the verdict it produces"cheating is −EV""this provider cheated, at layer 3"
what it requiresa stake worth more than the joba challenger who ran the job
5 rounds = ⌈log₂ 24⌉, 1,920 B of digests and Merkle paths, and the guilty layer correctly named in 24 of 24 cases. A challenger who disputes an honest job is ruled against by the same machinery. Bisection does not make verification cheap — it makes adjudication cheap, and the ruling certain. The whole extra cost is that someone has to re-execute.
E019 · E020 · E026

Why the interactive proof does not rescue it

Sumcheck replaces a 16.8 MB intermediate with a 792-byte transcript — 21,183× — and the prover pays only 10.4% at n = 4096, of which the protocol itself is 0.5%. The rest was our own implementation, and the measurement found two of its own bugs on the way down from 36.1%.

On a transformer layer it saves nothing, because every matmul output feeds a lookup-table non-linearity the verifier must recompute anyway. The tension is structural: the contract makes non-linearities lookup tables because a gather has no arithmetic and so cannot diverge, and a sumcheck passes only through low-degree arithmetic.

The property that buys determinism blocks the proof.

Making a transformer sumcheck-friendly means changing the model, and that was measured too. Replacing GELU with x² is free. Replacing softmax is the whole cost, and it grows with depth — +6.8% at 2 layers, +19.8% at 8 — so it is a divergence, not a fixed tax a bigger model amortises. The polynomial route is closed.

Taxonomy

Six claims, and which ones evidence can actually support

Claim O

Origin

The artefact was produced by party P, was not modified, and is not a replay.

A prerequisite, not an achievement. Nodus v1 has exactly this and nothing else.

Claim A

Correctness

output = f(model, input) for the specified f, model, input.

Only meaningful when f is well defined. On heterogeneous accelerators in float, f is a family of functions differing in rounding — which is what the numeric contract exists to fix.

Claim B

Execution

The provider actually ran the computation, on this occasion, to obtain this result.

Cryptographic proofs of A do not imply B. Freshness and binding constructions are what speak to it.

Claim C

Work

The provider performed the claimed quantity and type of computational work.

Conditionally achievable in software, and always resting on an unproven hardness conjecture.

Claim D

Physical resource usage

A specific machine consumed the claimed GPU-seconds, cycles, or energy.

Not achievable trustlessly. A proof is a mathematical object and carries no information about the hardware that produced it.

Claim E

Useful computation

The provider did the genuinely required work rather than any of the shortcuts.

A conjunction, not a primitive: E ≈ A ∧ B ∧ model binding ∧ freshness.

Nodus v1’s HMAC-authenticated receipts establish Claim O and nothing else. Authentication is not proof of computation.