Origin
The artefact was produced by party P, was not modified, and is not a replay.
A prerequisite, not an achievement. Nodus v1 has exactly this and nothing else.
Every figure below is measured against the same baseline: running the job again. R_verify = 1 by definition, so anything above that line has to justify itself with something re-execution cannot provide.
The measurement the rest of the programme rests on: held to the numeric contract, every honest backend tested returns the same logits, bit for bit.
| Machine | Stack | Logits digest |
|---|---|---|
| Apple M5 CPU | ARM · Accelerate · macOS | 202f51a3… |
| Apple M5 GPU | Metal · MLX | 202f51a3… |
| Intel Xeon | x86 · OpenBLAS · Debian 13 | 202f51a3… |
| Intel Xeon | x86 · OpenBLAS · Ubuntu 22.04 | 202f51a3… |
| NVIDIA L4 | CUDA 12.4 · TF32 enabled | 202f51a3… |
Two CPU vendors, two GPU vendors, three operating systems, three Python versions — every next-token prediction matching. TF32 was left on deliberately: int8 operands sit four bits inside the contract’s operand bound, so the reduced-precision tensor-core path cannot lose anything. The contract never asks a provider to cripple their hardware, which is the difference between a specification that can be complied with and one that is merely stated.
| Mechanism | Workload | R | Source |
|---|---|---|---|
| General-purpose zkVM (SP1) | y = 3x + 5 | 7 × 10⁴× | E001 |
| Freivalds' algorithm | 8192² matmul | 0.023 | E003 |
| Sumcheck, single matmul | n = 4096 over F_p | 0.634 | E026 |
| Sumcheck, two-layer chain | n = 4096 | 0.465 | E026 |
| Re-execution | anything | 1.000 | baseline |
Freivalds’ guarantee assumes exact arithmetic. In float it degrades to a tolerance, and a tolerance is the attacker’s budget. Crossing silicon vendors costs 5 × 10⁻⁷; crossing from CPU to GPU on the same machine costs 1,600× more.
Log scale. Everything left of the hatched band is honest; everything right of it is a cheat.
Every honest backend returns the bit-exact product. The window has no upper edge.
The GPU’s imprecision was never about the GPU. It was about representing real numbers.
Verifying a layer requires the verifier to hold that layer’s intermediates. Against a 0.02 MB answer that is a factor of three thousand, and it makes every configuration lose. Compute was never the binding cost.
The fix is to audit one random layer rather than all of them: bandwidth stays flat while the compute saved scales with depth. Built and run against the real model — Merkle commit, uniform post-commitment challenge, open, recompute. This is not a proof and we do not call it one: a single-layer cheat escapes with probability 1 − 1/L.
| honest provider, all 24 layers challenged | accepted |
| single-layer cheat, 200 trials | detected 4.5% (1/L = 4.2%) |
| cheat everywhere, answer the challenge honestly | rejected by the commitment |
| audit traffic vs verifying every layer | 4.2% |
Two parties who disagree binary-search to the first layer whose digest differs; an arbiter recomputes that layer alone.
| E024 random audit | E027 bisection | |
|---|---|---|
| bandwidth on the wire | 4.2% | 4.24% |
| arbiter's compute | 4.2% | 4.2% — one layer of 24 |
| detection | 4.2% | certain |
| the verdict it produces | "cheating is −EV" | "this provider cheated, at layer 3" |
| what it requires | a stake worth more than the job | a challenger who ran the job |
Sumcheck replaces a 16.8 MB intermediate with a 792-byte transcript — 21,183× — and the prover pays only 10.4% at n = 4096, of which the protocol itself is 0.5%. The rest was our own implementation, and the measurement found two of its own bugs on the way down from 36.1%.
On a transformer layer it saves nothing, because every matmul output feeds a lookup-table non-linearity the verifier must recompute anyway. The tension is structural: the contract makes non-linearities lookup tables because a gather has no arithmetic and so cannot diverge, and a sumcheck passes only through low-degree arithmetic.
The property that buys determinism blocks the proof.
Making a transformer sumcheck-friendly means changing the model, and that was measured too. Replacing GELU with x² is free. Replacing softmax is the whole cost, and it grows with depth — +6.8% at 2 layers, +19.8% at 8 — so it is a divergence, not a fixed tax a bigger model amortises. The polynomial route is closed.
The artefact was produced by party P, was not modified, and is not a replay.
A prerequisite, not an achievement. Nodus v1 has exactly this and nothing else.
output = f(model, input) for the specified f, model, input.
Only meaningful when f is well defined. On heterogeneous accelerators in float, f is a family of functions differing in rounding — which is what the numeric contract exists to fix.
The provider actually ran the computation, on this occasion, to obtain this result.
Cryptographic proofs of A do not imply B. Freshness and binding constructions are what speak to it.
The provider performed the claimed quantity and type of computational work.
Conditionally achievable in software, and always resting on an unproven hardness conjecture.
A specific machine consumed the claimed GPU-seconds, cycles, or energy.
Not achievable trustlessly. A proof is a mathematical object and carries no information about the hardware that produced it.
The provider did the genuinely required work rather than any of the shortcuts.
A conjunction, not a primitive: E ≈ A ∧ B ∧ model binding ∧ freshness.
Nodus v1’s HMAC-authenticated receipts establish Claim O and nothing else. Authentication is not proof of computation.