NodusLab
E001 · CP-001

Simple arithmetic in a general zkVM

What does it cost to prove y = 3x + 5, and how does that cost change as the amount of proved work grows?

Checkpoint: CP-001 Stage: 1 Status: complete System under test: SP1 6.5.0 (CPU prover), riscv64im-succinct-zkvm-elf Machine: Apple M5, 10 cores, 25.8 GB, macOS 26.2

Question

What does it cost to prove y = 3x + 5, and how does that cost change as the amount of proved work grows?

Answer

core compressed
native computation 0.69 ns 0.69 ns
proving y = 3x + 5 11.19 s 58.12 s
verification 56.3 ms 38.3 ms
proof size 2.78 MB 1.27 MB
fixed cost per proof 12.29 s ≈ 57.7 s
asymptotic R_prove 7–9 × 10^4 ≈ 1 × 10^5
verification vs job size grows flat

Three findings, in order of importance:

  1. The cost is fixed, not marginal. Proving one operation and proving ten thousand both cost ~11 s. The fixed charge equals the marginal charge only at ~970,000 zkVM cycles.
  2. The asymptotic overhead is ~7 × 10^4 ×, and this is a lower bound — u64 arithmetic is the friendliest workload a RISC-V zkVM will ever see.
  3. Succinct verification is achievable and cheap; the system is not. Compressed proofs verify in a flat ~37 ms with a byte-identical 1.27 MB proof at any job size — but reaching the point where that beats re-execution costs ≈ 28 minutes of provider CPU to verify 35 ms of computation.

Why this is the right first experiment

It is the cheapest possible way to measure the thing that turns out to govern all the later economics: the fixed cost of a proof. If a proof has a large constant charge, then the entire ZK branch of the research is a question about amortisation — how much useful work can be folded under one proof — rather than a question about cryptographic efficiency. That reframing is worth more than any single benchmark number, and it costs one afternoon to establish.

It is also a deliberately optimistic setting. u64 wrapping arithmetic in a tight loop is the friendliest workload a RISC-V zkVM will ever see: no floating point, no large memory, no non-linearities, no data-dependent control flow. Every cost measured here is therefore a lower bound on what the same system would charge for neural network inference.

Files

file contents
hypothesis.md H1, H1a, H1b and the predictions recorded before running
methodology.md program, measurement procedure, fairness caveats, repro
implementation/sp1-arith/ the guest program and the Rust measurement driver
implementation/run.py sweep driver; emits harness records
results/ plots and derived tables
analysis.md the result
next_steps.md what the evidence says to do next

Raw records: ../../benchmarks/results/001-simple-arithmetic.jsonl

Claims established

Per docs/verification-taxonomy.md: Claim A only, at level L5. Not B, not C, not D, not E. The security block on every record spells out the trust assumptions and the residual attacks.

Research question (CP-001)

What does it cost to prove the smallest meaningful arithmetic statement in a general-purpose zkVM, and how does that cost change as the amount of work grows?

The roadmap's Stage 1 statement is y = 3x + 5. On its own that measures one point. The experiment therefore sweeps the same statement shape upward — y ← 3y + 5 applied k times — so that a single program isolates the two economically distinct quantities:

    prove_time(k)  ≈   a    +    b · cycles(k)
                       ▲            ▲
                  fixed cost    marginal cost
                  paid per job  paid per unit of work

Hypotheses

H1 (primary). A general zkVM has a large fixed cost that dominates small jobs, so R is a hyperbola in job size rather than a constant. Proving y=3x+5 costs essentially the same as proving ten thousand of them.

  • Falsified if proving time is approximately linear through the origin.

H1a. The fixed cost is large enough in absolute terms that there exists a minimum economically provable job size — below it, proof cost cannot be recovered from the job's value at any plausible price.

H1b. Verification cost for an SP1 core proof is not constant, because a core proof is a list of shard proofs and the verifier checks each one. If so, core proofs are structurally unsuited to a coordinator that must verify every job, and the relevant mode for Nodus is a compressed/recursive proof.

  • Falsified if verify time and proof size stay flat as cycles grow.

Predictions (recorded before running)

quantity prediction basis
prove time at k=1 1–60 s zkVM setup + minimum one shard
prove time flat until ~10^5–10^6 cycles typical SP1 shard size
R_prove at k=1 > 10^9 native is ~1 ns; proving is seconds
proof size (core) 10^5–10^7 bytes core proofs are not succinct
verify time 10–100 ms, growing with shard count linear in shards

What this experiment does not test

  • Whether SP1 is a good zkVM. One implementation, one machine, one workload.
  • Anything about neural networks. u64 wrapping arithmetic is the friendliest possible workload for a RISC-V zkVM: no floating point, no memory pressure, no non-linearities. Every number here is therefore an optimistic bound on what AI inference would cost in the same system. That framing is the point of running it first.
  • Claims B, C, D or E. See the security block attached to every record.

Program under proof

implementation/sp1-arith/program/src/main.rs, compiled to riscv64im-succinct-zkvm-elf:

let x     = sp1_zkvm::io::read::<u64>();
let iters = sp1_zkvm::io::read::<u32>();
let mut y = x;
for _ in 0..iters { y = y.wrapping_mul(3).wrapping_add(5); }
sp1_zkvm::io::commit(&x);
sp1_zkvm::io::commit(&iters);
sp1_zkvm::io::commit(&y);

iters = 1 is exactly y = 3x + 5. Wrapping arithmetic so that no input can panic; all values are u64, so the zkVM never touches floating point.

Exactly what the resulting proof asserts

There exists an execution of this ELF whose committed public values are (x, iters, y) and which halted successfully.

Notably it does not assert:

  • that the execution happened recently, or at all after the job was assigned (no freshness binding → Claim B unsupported);
  • that any particular machine did it (Claims C, D unsupported);
  • that the ELF is the program the coordinator meant — the verifying key binds the compiled artefact, so the compiler and build are inside the trust set.

Measurement procedure

Driver: implementation/run.py, using benchmarks/harness.py.

  1. Native baseline, same process, same run. The identical loop in host Rust, wrapped in std::hint::black_box on input and output so the optimiser cannot fold it. Repetition count = clamp(10^8 / iters, 1, 2·10^7) with a 10% warm-up pass; the reported figure is the mean.
  2. zkVM execute (no proving) to obtain the cycle count.
  3. Setup — proving/verifying key generation, timed separately because it is amortisable across all jobs running the same program.
  4. Prove — core mode by default; compressed as a separate sweep.
  5. Save proof to disk, measure the file size in bytes.
  6. Verify in-process against the verifying key.
  7. Peak RSS for the whole process via /usr/bin/time -l.

Output correctness is asserted inside the driver: the zkVM's committed y is compared against the natively computed y on every run, so a silent divergence fails the experiment rather than producing a pretty number.

Fairness notes (things that could make this misleading)

  • Peak RSS is process-wide, covering the native baseline, execution, proving and verification. It is an upper bound on prover memory, not a measurement of it. Labelled as such in the records.
  • The native baseline is absurdly fast (sub-nanosecond per iteration), which makes R enormous. That is not a rhetorical trick — it is the actual ratio, and it is the honest denominator. Where a reader wants a friendlier framing, the cycle-normalised figure (prove_seconds / cycles) is also reported and is the number that extrapolates to other workloads.
  • Setup is excluded from R. Including it would overstate per-job cost for a network that runs the same program repeatedly.
  • Single machine, single run per configuration. No error bars. Repeat counts should be added before any of these numbers is published. Treated as order-of-magnitude evidence, which is sufficient for the decision at hand.

Environment

  • Apple M5, 10 cores (10 physical / 10 logical), 25.8 GB, macOS 26.2 arm64.
  • SP1 cargo-prove 92b8eab (2026-08-26), SDK crates 6.5.0, CPU prover.
  • Guest target riscv64im-succinct-zkvm-elf, guest toolchain rustc 1.94.0-dev.
  • No GPU proving, no network prover, no Docker.

Reproducing

cd implementation/sp1-arith
PATH="$HOME/.sp1/bin:$PATH" cargo build --release --manifest-path script/Cargo.toml
cd ../..
../../.venv/bin/python implementation/run.py --sweep 1,10,100,1000,10000,100000,1000000

Status: complete (core mode); compressed mode reported in §6. Date: 2026-09-01. Machine: Apple M5, 10 cores, 25.8 GB, macOS 26.2. System: SP1 cargo-prove 92b8eab (2026-08-26), SDK 6.5.0, CPU prover. Raw records: ../../benchmarks/results/001-simple-arithmetic.jsonl Tables: results/summary.md


1. Answer to CP-001

Proving y = 3x + 5 in a general-purpose zkVM:

native computation 0.69 ns
zkVM cycles 5,186
proving 11.19 s
verification 56.3 ms
proof size 2.78 MB
peak RSS 6.8 GB
R_prove 1.6 × 10^10
R_verify 8.2 × 10^7

Proving ten thousand of the same operations costs 11.5 s — statistically indistinguishable. The cost is not for the work; it is for the existence of a proof.

Full sweep, four decades of work:

iters cycles native prove (s) verify (ms) proof (MB) peak RSS (GB) R_prove R_verify
1 5,186 0.7 ns 11.19 56.3 2.78 6.8 1.6e10 8.2e7
10 5,231 3.3 ns 11.00 55.8 2.78 8.4 3.3e9 1.7e7
100 5,681 45.1 ns 11.50 56.6 2.78 8.9 2.6e8 1.3e6
1,000 10,181 658 ns 15.92 72.7 2.78 8.0 2.4e7 1.1e5
10,000 55,181 8.54 µs 16.30 73.2 2.78 8.4 1.9e6 8.6e3
100,000 505,181 85.8 µs 21.82 77.0 2.78 8.0 2.5e5 897
1,000,000 5,005,181 860 µs 62.69 78.4 2.83 8.3 7.3e4 91.2
4,000,000 20,005,181 2.90 ms 252.78 142.0 5.88 13.9 8.7e4 49.0

A ninth configuration (16M iters, ~80M cycles) was started and deliberately stopped: peak RSS was tracking toward the machine's 25.8 GB and the run had begun to swap, which would have produced a contaminated timing. Four decades of clean data were judged sufficient for the conclusions below. Recorded here rather than silently omitted.


2. H1 — confirmed

prove_seconds = a + b · cycles, least squares over all eight points:

  • fixed cost a = 12.29 s per proof
  • marginal b = 11.9 µs/cycle (84 k cycles/s)
  • asymptotic marginal rate from the two largest points: 79 k cycles/s
  • the fixed charge equals the marginal charge at ≈ 970,000 cycles

Below ~10^6 cycles, most of what a provider pays is the fixed cost of producing a proof at all. R is a hyperbola in job size, not a constant, exactly as predicted.

H1a — confirmed, and the number is uncomfortable

There is a minimum economically provable job size, and it is far above a typical Nodus job. The comparison that makes this concrete uses v1's own unit (cost_model.py): 1 CU-v0 ≈ 0.456 machine-seconds on this machine.

The fixed cost alone of one SP1 core proof — 12.29 s — is equivalent to 27 CU of raw compute.

Nodus v1's production run settled 38.165 CU across 14 jobs, i.e. a median job of roughly 2.7 CU. One proof's fixed overhead therefore exceeds the entire compute cost of a typical v1 job by about an order of magnitude, before proving any of the job's actual work.

The asymptotic overhead — the number to carry forward

R_prove falls as jobs grow, then flattens:

  1e10 ┤●
       │ ●
  1e8  ┤   ●
       │      ●
  1e6  ┤         ●
       │            ●
  1e4  ┤               ●──●     ← flattens at ~7-9 × 10^4
       └──────────────────────
        10^3        10^7  cycles

MEASURED: the asymptotic proving overhead of SP1 on this workload is 7.3–8.7 × 10^4 ×. No amount of job size removes it — it is the steady-state price of turning execution into a proof.

Applying that ratio to Nodus's own unit: proving 1 CU-v0 of computation costs ≈ 9 hours of CPU (0.456 s × 7.3 × 10^4), at a declared local rate of about $0.04 per CU proved, against $5.5 × 10^-7 per CU computed.

This is a lower bound, and by a wide margin. The workload measured here — u64 wrapping arithmetic in a tight register-resident loop — is the friendliest thing a RISC-V zkVM will ever see. Real inference is floating-point, SIMD-heavy, and memory-bound: a zkVM has no vector units, must emulate floats in software, and charges for every memory access. The true ratio for inference is certainly worse, plausibly by another 1–3 orders of magnitude. Any claim that general zkVM proving of LLM inference is near-practical has to answer this measurement first.


3. H1b — falsified as stated, and the correction matters more

Predicted: core-proof verification is linear in shard count, so verify time and proof size grow roughly linearly with cycles.

Measured: they grow, but far sublinearly. Over a 3,858× increase in cycles (5,186 → 20,005,181), verification grew only 2.5× (56.3 → 142.0 ms) and proof size 2.1× (2.78 → 5.88 MB). The prediction was too strong and is retracted; the annotation in the early records has been corrected in the runner and the retraction noted there.

But the economically important consequence survives, for a subtler reason. Over the linear regime (5M → 20M cycles) verification grows at ≈ 4.24 ns per cycle while the native computation costs ≈ 0.145 ns per cycle. Verification cost grows ~29× faster than the work it verifies. Therefore:

DERIVED (from two points — weak, needs confirmation): R_verify for SP1 core proofs asymptotes to ≈ 29× and never crosses 1.

If that holds, an SP1 core proof is never cheaper for a coordinator to check than simply re-running the job — not for any job size. The declining R_verify column (8.2e7 → 49) is not converging on 1; it is converging on ~29.

Conclusion: core proofs are structurally the wrong product for Nodus. A coordinator that must verify every job needs verification cost independent of job size, which means recursion. §6 confirms that compressed proofs deliver exactly that — flat verification and byte-identical proof size across a 965× range of job sizes.


4. Prover memory is a real constraint, not a footnote

Peak RSS reached 13.9 GB at 20M cycles, and the 80M-cycle run had to be abandoned on a 25.8 GB machine. Memory grows sublinearly (8.3 GB at 5M → 13.9 GB at 20M) but it grows, and it binds sooner than time does.

For a compute network this reframes provider eligibility: a node able to run a model is not necessarily able to prove it ran the model. Proving capacity is a separate, RAM-dominated hardware requirement, and pricing it as though it were the same resource would be wrong.


5. What the proof actually establishes

Per docs/verification-taxonomy.md: Claim A, at level L5. Nothing else.

The statement is: there exists an execution of this ELF whose committed public values are (x, iters, y) and which halted successfully.

Trust set: the verifier holds the correct verifying key; the ELF is the intended program (the compiler and build are inside the trust boundary); Fiat-Shamir in the random-oracle model; no soundness bug in SP1 (unaudited by us).

What a malicious provider can still get away with

  1. Cache and replay. Nothing binds the proof to a fresh execution. An (input, output, proof) triple verifies forever. Claim B is unsupported, and this is the cheapest gap to close — a coordinator-chosen nonce inside the committed input costs nothing.
  2. Claim no work while doing none. The proof carries zero information about resources consumed. A prover with a faster method, a precomputed table, or a friend who already proved this statement produces a byte-identical, valid proof. A correctness proof is not a proof of work done — the central claim/payment mismatch of research-thesis.md §4, now demonstrated rather than argued.
  3. Outsource the proving entirely to a third party, or to a cheaper machine than the one that claimed the capacity.

6. Compressed proofs — the structural fix, measured and confirmed

SP1's compressed mode recursively folds the shard proofs into one. It should decouple verification cost from job size. It does, completely.

mode cycles prove (s) verify (ms) proof (KB) R_verify
compressed 5,186 58.12 38.3 1,272.6 5.1e7
compressed 55,181 57.75 38.1 1,272.6 4.2e3
compressed 5,005,181 104.32 35.2 1,272.6 33.9
(core, for comparison) 5,005,181 62.69 78.4 2,830 91.2

Across a 965× increase in cycles: verification time is flat (38.3 → 35.2 ms, no trend beyond noise) and proof size is exactly constant at 1,272.6 KB to the byte. This is what succinctness looks like, and it is the property Nodus needs: a coordinator's per-job verification cost that does not depend on how big the job was.

The price of it:

  • fixed proving cost rises from 12.29 s (core) to ≈ 57.7 s (compressed) — 4.7× — because the recursion itself has to be proved;
  • marginal proving rate is ≈ 106 k cycles/s (from the two largest points), comparable to core's 79 k cycles/s;
  • prover RSS is similar (9.5–11.4 GB).

So compression shifts cost from the coordinator to the provider, which is the right direction for a network with one coordinator and many providers.

The crossover, now computable

With verification pinned at ≈ 35 ms and native cost at ≈ 2.07 × 10^-10 s/cycle (measured in the same run), R_verify reaches 1 at:

≈ 1.7 × 10^8 zkVM cycles. Above that, it is genuinely cheaper for the coordinator to check the proof than to re-run the job.

But look at what it costs to get there. Producing that proof takes 57.7 s + 9.41 µs/cycle × 1.7×10^8 ≈ 28 minutes of provider CPU, to verify a computation whose native cost is 35 milliseconds. The system-wide overhead at the crossover point is ≈ 4.7 × 10^4 ×.

That is the honest summary of general-purpose ZK for this workload:

The verifier can be made cheap. The system cannot. Succinctness is real and it is achievable today, but it is bought at ~5 × 10^4 × the cost of the computation, paid by the provider. Whether that is worth it depends entirely on whether the job needs something re-execution cannot give — third-party verifiability or weight privacy. For a coordinator that already holds the model and could simply re-run it, it is not.

Note that 1.27 MB is still a recursive STARK, not a constant-size wrapped SNARK. The few-hundred-byte, on-chain-verifiable regime needs the Groth16/PlonK wrap, which requires a gnark toolchain not installed here. Queued.

7. Economic implications

  1. General zkVMs cannot price small jobs. A 12.3 s fixed cost versus a 2.7 CU (≈1.2 s) median job is not a tuning problem, it is a category error. If ZK is used at all it must be amortised across many jobs or many layers — which promotes batching and recursion (E007/E008) from "optimisations" to preconditions.
  2. The asymptotic 7×10^4 overhead is the number that decides the ZK branch. It must be beaten by ~4 orders of magnitude by specialised arguments before general-purpose proving of inference is worth discussing.
  3. Proving is a distinct hardware market. RAM-bound, not FLOP-bound. If Nodus ever pays for proofs, it is buying a different resource than the one CU-v0 measures.
  4. Nothing here supports Claims B, C or D, which is what a compute economy actually bills for. Even a perfect, free proof would not stop a provider billing for work it did not do.

8. Limitations

  • One zkVM, one machine, one run per configuration, no error bars. RISC Zero as a second baseline is queued specifically because a single implementation cannot support a claim about "general zkVMs".
  • Peak RSS is process-wide and therefore an upper bound on prover memory.
  • The R_verify asymptote in §3 rests on two points. Weak. Flagged as DERIVED.
  • u64 arithmetic is not a proxy for inference. Every number here is optimistic.

9. Next

See next_steps.md. In short: the evidence sent us to Experiment 003 immediately, and that experiment has already returned an answer nine orders of magnitude better on the same metric.

What the evidence said to do, and what we did about it

E001 measured a general zkVM at R_prove ≈ 7×10^4 asymptotically and R_verify converging to ~29 rather than to 1. That is not a promising basis for verifying inference, and the result made one thing urgent: find out whether anything achieves R_verify < 1 at all. Rather than proceed down the roadmap in order, we went straight to CP-003.

That was the right call. Experiment 003 measured Freivalds' algorithm at R_verify = 0.023 on the same machine — 43.7× cheaper than recomputing an 8192×8192 matmul, with soundness error ≤ 2^-40 and no cryptographic assumptions. Nine orders of magnitude better than the zkVM on the same metric and the same claim.

The programme's centre of gravity moves accordingly: structure-exploiting verification first, general-purpose proving as a baseline we no longer invest in.


Immediate follow-ups to E001 itself

1. Settle the verification-flatness question — DONE

Compressed verification is flat (38.3 / 38.1 / 35.2 ms) and proof size is byte-identical (1,272.6 KB) across a 965× range of cycles. Core proofs are confirmed as the wrong product. R_verify crosses 1 at ≈ 1.7 × 10^8 cycles — but the proof that achieves it costs ≈ 28 minutes of provider CPU to verify 35 ms of computation. See analysis.md §6.

The follow-up this creates: if the coordinator's cost can be made constant, the whole question becomes how many jobs one proof can cover. That promotes E007 (batching) and E008 (recursive aggregation) — see the queue below.

2. Groth16/PlonK wrap

1.27 MB is not a succinct proof. Getting to the few-hundred-byte, constant-time- verification regime needs the gnark wrap. Required before any claim about on-chain or third-party verification.

3. RISC Zero as a second baseline

One zkVM cannot support a statement about "general zkVMs". Until this exists, every E001 conclusion is properly written as "SP1 6.5.0 on CPU", not "zkVMs".

4. Repeat runs and error bars

Single run per configuration. Before any of these numbers leaves the repository, n ≥ 5 with reported spread.


The queue, re-ordered by E001+E003 evidence

priority work why
1 E009-pre — heterogeneous honest divergence Gates the tolerance in E003 and the replication threshold in v1. Cheap. Nothing in the L2/L3 branch can be deployed without it.
2 E003 §2 — Freivalds over a prime field Restores the exact 2^-40 guarantee that floating point destroys. May imply the whole network should prefer quantised integer inference — an architectural consequence derived from a verification requirement.
3 E003 §3 — chain of matmuls Turns "verify a matmul" into "verify an inference". Tests whether the advantage survives having to hold every intermediate activation.
4 Freshness binding everywhere Converts Claim A into A ∧ B for ~zero cost, in both E001 and E003 settings. Cheapest security improvement identified so far.
5 E002 — dot products Still worth doing, but demoted: its purpose was to extrapolate zkVM cost to model scale, and E001's asymptote already tells us that answer is "no". Run it to get cycles-per-MAC for the record, not as a decision input.
6 Sumcheck/GKR for matmul The honest test of whether cryptography buys anything here: it restores succinctness and weight privacy, which Freivalds cannot provide. Now has a demanding baseline (0.023) to beat.
7 E007/E008 batching + recursion Promoted from "optimisation" to precondition for any ZK use, because E001 showed the cost is a per-proof constant.

What would change our mind about general zkVMs

Any of:

  • a measured R_prove below ~10^2 for a realistic inference workload;
  • GPU proving closing 2+ orders of magnitude (worth one measurement, not a research programme);
  • a requirement for weight privacy or third-party verifiability that Freivalds and interactive proofs cannot meet, at a price someone will actually pay.