The Nodus Paradigm
Introduction
Nodus Verifiable Compute asks one question and answers it with numbers. Nothing here is deployed; no production node can currently run a contract-conformant job. What exists is twenty-seven experiments, a specification, and a portable kit that lets anyone break the central claim on their own machine.
The purpose of this page is to state what the programme is asking, what it refuses to assume and what would make it a failure. Each of those is written down in advance, because a research programme that decides afterwards what it was testing has not tested anything.
The question
Can we build a verification layer for distributed AI compute where the cost of obtaining trustworthy evidence is significantly lower than re-executing the computation, while reducing the amount of trust placed in the compute provider?
Two quantities appear in that sentence and the whole programme is about their ratio. R is the cost of obtaining trustworthy evidence divided by the cost of the native computation, and it is the headline metric of every experiment here. It is not the only one — evidence also has a strength, and strength without a cost figure is marketing — but no architecture is interesting unless we can state both its R and exactly what it buys.
There is always a trivially available mechanism: run the job again yourself. Its R is about one, or two for a fully replicated system. It requires that the verifier own hardware capable of the job. It establishes correctness, and only for deterministic workloads. It establishes nothing about who ran the work, when, on what, or what it cost them. Any proposal that costs more than this and offers nothing re-execution cannot — privacy, a weak verifier, tolerance of non-determinism, succinctness for a third party — is economically dominated, and is recorded as such.
What this programme refuses to assume
These are commitments, not opinions. Each one exists because it is a well-trodden way to waste a research year.
ZK is not assumed to be the answer. Measured proving costs for neural inference are 10⁶–10⁸× native. That gap may close; it may not. Either way it is a measurement, not a premise.
TEEs are not assumed to be the answer. A TEE moves trust from the provider to a hardware vendor, plus its firmware supply chain, plus its attestation service. That is less trust, not no trust, and the reduction has to be argued rather than asserted.
Hardware is not assumed to be the answer. “Build a proof-friendly accelerator” is a decade-scale bet. It may be correct. It cannot be the premise of a programme whose first job is to find out what is already cheap.
Determinism is not assumed. Floating-point non-determinism across heterogeneous accelerators breaks naive replication and naive output-hash equality. Nodus is explicitly heterogeneous, so this is a first-order problem rather than an edge case.
Correctness is not assumed to be what we are buying. The network pays for compute units, not for correct answers. Almost all of the verifiable-computation literature addresses the second one — and a valid proof of a cheaply-obtained correct answer is still a valid proof.
The result
The binding constraint turned out not to be cryptographic cost. It is that honest hardware does not agree with itself. Two honest providers running the same model on different devices diverge more than some cheats do, which leaves no threshold that separates them.
Constraining the job to exact integer arithmetic removes the problem instead of trading against it. Qwen2.5-0.5B, held to the numeric contract, produced logits with the digest 202f51a312e5c0e685354355 on five machines from four vendors — an Apple M5 on both CPU and GPU, Intel Xeon under two operating systems, and an NVIDIA L4 with TF32 left deliberately enabled. Byte-identical on all five. Once the disagreement is gone, verification is cheap by classical means and needs no cryptography at all.
It costs +2.43% perplexity — 95% CI 1.63–3.23, 120 chunks, paired — over the best achievable int8 quantisation, and nothing in provider hardware.
The working hypotheses
Each is mapped to an experiment and each states what would falsify it. Two were added mid-programme, after a result made the original set insufficient.
H1. A general zkVM has a large fixed cost that dominates small jobs, so R is a hyperbola rather than a constant. Confirmed. Fixed cost 12.29 s, flat to ~10⁶ cycles. E001 · E002
H2. For structured linear algebra, probabilistic verification reaches R ≪ 1 with a computable, adjustable soundness error. Confirmed for exact arithmetic — R_verify = 0.023 at 2⁻⁴⁰. Fails for floating point in a heterogeneous network: honest CPU/GPU divergence sits within 3.7× of a real precision cheat. E003 · E009
H3. Structure-exploiting verification beats general zkVMs by orders of magnitude on the same statement. Confirmed, ~10⁹× on R_verify for Claim A. Not yet tested against a cryptographic specialised argument. E001 vs E003
H4. Most practical security comes from cheap mechanisms — commitments, random audit, stake — with proof reserved for a high-value tail. Partially supported, and from an unexpected direction. Cost was never the obstacle; resolution was, and E014 removed it. Then E016 showed data movement, not cryptography, is what makes verifying every layer unaffordable. E009 · E012 · E016
H5. No mechanism buildable in 2026 establishes Claim D — physical resource usage — against a determined adversary without trusted hardware. Survives a literature search. Software forces work only under unproven hardness conjectures, and forcing work is still not evidence about a machine. E010 · E011 · E013
H6. The economically correct architecture is tiered: different evidence classes for different job values, not one mechanism for all jobs. Untested. E012
H7. The binding constraint is not the cost of verification but the numerical agreement among honest providers — so the decisive lever is specifying arithmetic, not choosing a proof system. Confirmed. Exact integer arithmetic costs no supply — all five honest backends return the bit-exact product — while separating cheats without bound. The heterogeneity problem was dissolved by changing the specification. E014
H8. Nodus need not solve Claim C or D at all: pricing jobs from the specification rather than from effort removes them from the settlement path. Supported. The only surviving exploit against exact verification was caching, which a hash lookup closes. Not yet quantified against real traffic. E013
What this does not claim
No new primitive. Freivalds is 1977, sumcheck is 1992, Merkle is 1979. The contribution is measurement and a specification.
SafetyNets still wins. Their 5% prover overhead beats our 10.4%, and their verification is 8–82× faster than re-execution where ours is 1.6×. What we add is that their approach does not transfer to a transformer, and why — their restriction to quadratic activations reads as load-bearing rather than dated, because their networks had no attention to lose.
Nothing is deployed. No production node can currently run a contract-conformant job. This is a measurement study, not a product.
No security proofs. Every claim is a measurement under stated assumptions, and the probabilistic ones are quoted with their soundness error.
One model, small samples. Qwen2.5-0.5B. Perplexity figures use 4–6 chunks of 128 tokens without error bars, except where E025 supplies them. The determinism results are exact-match and do not need them; the accuracy figures do.
House rules
The discipline the programme holds itself to, in the order it matters.
Authentication is never called proof of computation.
Probabilistic verification is never called a proof; quote its soundness error.
Correctness (Claim A) is never conflated with resource consumption (Claim D).
Every benchmark carries the re-execution baseline, measured on the same machine in the same run.
Numbers are labelled MEASURED / DERIVED / DECLARED / SPECULATIVE / UNKNOWN.
Failed experiments are written up. They are the cheap ones.
Search the literature before concluding something does not exist. We broke this rule once and rediscovered a 2017 result by construction; it is recorded in E016 rather than quietly fixed.
What would make this a failure
Stated in advance, so the goalposts cannot be moved afterwards. The programme has failed if it ends in any of the following.
Producing a roadmap and a literature review and no measurements.
Reporting proof-system benchmarks without the re-execution baseline beside them.
Describing a probabilistic scheme as a “proof”.
Claiming any single mechanism covers Claims A–E.
Building a ZKML implementation because it is interesting rather than because a measurement said it was the cheapest sufficient mechanism.
Read next
The whitepaperAll 27 experimentsThe documentsFalsify it yourself