NodusLab
E016 · CP-014

The bandwidth bill

what does verification cost in bandwidth, and does that kill it?

Checkpoint: CP-014 · Status: COMPLETE, run 2026-09-01 Question: what does verification cost in bandwidth, and does that kill it?

Answer: yes, then no

E015 measured verification compute and deferred the cost of getting the verifier its data. Measured now, that cost dominates by three to five orders of magnitude.

One layer, s=1024, d=2048:

the job's answer 0.02 MB
intermediates the verifier needs 49.1 MB — 3,010× the answer
compute saved by verifying instead of re-executing 0.097 s
break-even compute price at cloud egress $164/hour
actual price of a rented GPU-hour ~$1

Every configuration loses, including the most favourable (rented GPU + intra- datacentre transfer, still 1.3× underwater). Transmitting only matmul outputs and recomputing the non-linearities saves 82% of the traffic — but 82% off a number 3,000× too large changes nothing.

The assumption we hadn't stated

E003 and E015 aren't wrong; they assumed the verifier already holds what it needs. True for one matmul where the coordinator owns both operands. False for a pipeline, where layer k+1's inputs are layer k's outputs and live only on the provider's machine.

R_verify = 0.023 is a compute ratio. Once data has to move, compute is not the binding constraint.

The rescue: audit one layer, not all of them

Bandwidth for one layer is fixed; compute saved scales with depth.

layers catch probability vs cheap egress vs intra-DC
1 1.000 18× 1.8×
8 0.125 1.6× 0.16× ✓
32 0.031 0.39× ✓ 0.04× ✓
80 0.013 0.15× ✓ 0.02× ✓

At 32 layers on cheap egress, verification costs 39% of re-executing. What you pay is soundness: catch probability falls to 1/L.

And the stake that makes that safe

A provider skipping m of L layers saves m/L of the job's value V and is caught with probability m/L. Cheating is negative-EV when penalty·(m/L) > V·(m/L) — the m and L cancel:

A stake equal to the job's compute value makes layer-skipping negative-expected-value, at any depth and any level of greed.

Implementable with machinery v1 already has: escrow, delayed settlement, integer micro-credits.

What this promotes

Interactive proofs (GKR/sumcheck) verify layered computation with polylogarithmic communication instead of linear-in-the-intermediates — which is exactly the cost measured here. SafetyNets (NeurIPS 2017) applied this to neural inference on an untrusted cloud. That work moves from "interesting" to the known solution to the problem we just hit. We have not measured it.

Files

file contents
implementation/bandwidth.py measures wire sizes (actually compressed, not estimated) and prices them
analysis.md the result, the rescue, and the stake derivation
next_steps.md measure GKR; re-scope the earlier results

Raw records: ../../benchmarks/results/016-verification-bandwidth.jsonl

Status: complete. Result: negative, then rescued. Date: 2026-09-01. Apple M5; NumPy 2.5.2; zlib level 6. Raw records: ../../benchmarks/results/016-verification-bandwidth.jsonl


1. The negative result

E015 measured verification compute and explicitly deferred the cost of getting the verifier the data it needs. Measured now, it does not merely matter — it dominates by three to five orders of magnitude.

For one layer, s=1024, d=2048:

the job's answer 0.02 MB
intermediates the verifier needs (compressed) 49.1 MB
ratio 3,010×
compute saved by verifying instead of re-executing 0.097 s

Priced (compressed bytes, so this is the honest wire size — the 2.4× compression ratio recovers almost exactly the int64-where-int32-would-do waste in our implementation):

compute valued at bandwidth at bandwidth cost / compute saved
local CPU cloud egress $0.09/GB 27,900×
local CPU cheap egress $0.01/GB 3,100×
local CPU intra-datacentre 310×
rented GPU cloud egress 121×
rented GPU cheap egress 13×
rented GPU intra-datacentre 1.3×

Every configuration loses, including the most favourable one. Break-even would require compute worth $164/hour at cloud egress, $18/hour at cheap egress, or $2/hour intra-datacentre. A rented GPU-hour is about $1.

The re-derivation trick from E015's next-steps does work — transmitting only matmul outputs and recomputing every non-linearity from them saves 82% of the traffic. But 82% off a number that is 3,000× too large changes nothing.

2. The assumption we had not stated

E003 and E015 are not wrong, but they rested on something unexamined:

The verifier already holds everything it needs.

That is true for a single matrix product where the coordinator owns A and B. It is false for a pipeline, where the operands of layer k+1 are the outputs of layer k and exist only on the provider's machine.

So the honest scope of the earlier results is narrower than we wrote it: R_verify = 0.023 is a compute ratio, valid when the data is already local. Once the data has to move, compute stops being the binding constraint.

This is why interactive proofs exist. GKR/sumcheck protocols verify a layered computation with communication polylogarithmic in the circuit size rather than linear in the intermediates — which is precisely the cost we just measured. SafetyNets (NeurIPS 2017) applied exactly this to neural network inference on an untrusted cloud. That line of work moves from "interesting alternative" to the known solution to the problem we just hit, and it is now the top of the reading queue.

2a. Correction: the literature had this, and we read it in the wrong order

After running this experiment we read SafetyNets (NeurIPS 2017) in full. Its §4.3 contains the sentence:

"We note SafetyNets has significantly lower bandwidth costs compared to an approach that separately verifies the execution of each layer using only the IP protocol for matrix multiplication."

That approach is precisely the one we built and measured. The qualitative finding of this experiment was published in 2017; we rediscovered it by construction. Research rule 15 says search the literature before concluding something is unknown, and we ran the experiment first. Recorded rather than quietly fixed.

What is genuinely new here: the magnitude in money for a distributed compute economy (3,010× the answer size; break-even at $164/hour against a $1 GPU-hour), the measured compressed wire sizes, and the layer-audit-plus-stake analysis in §3–4. SafetyNets states the direction of the effect; it does not price it for this setting.

For calibration, SafetyNets' measured alternative: prover overhead 5%, communication < 8 KB at batch 2048 (under 2% of the input/output traffic), verifier 8–82× faster than re-executing. Those numbers are the target any scheme of ours has to beat, and we are not close.

2b. Update (E019): the protocol that fixes this exists, and does not fit our contract

E019 implemented the sumcheck interactive proof this section pointed at. On a two-layer chain of 2048x2048 matmuls it replaces a 16.8 MB intermediate with a 792-byte transcript — a 21,183x reduction, improving with scale.

Applied to the E015 layer it saves nothing. All three matmul outputs — 49.1 MB, the exact figure measured here — feed lookup-table non-linearities that the verifier checks by recomputation, so it needs the values regardless. The chain breaks at the first non-linearity and there is one after every matmul.

So the layer-audit-plus-stake design below is not superseded; it remains the cheapest thing that works end to end on a real transformer. See ../019-interactive-proof/analysis.md.

3. The rescue: audit one layer, not every layer

Bandwidth for auditing a single layer is fixed. The compute saved by not re-executing scales with the model's depth. So the ratio improves linearly with the number of layers.

DERIVED from the measurements above, GPU compute at $1.00/hour:

layers L catch probability compute saved vs cloud egress vs cheap egress vs intra-DC
1 1.000 0.10 s 164× 18× 1.8×
8 0.125 1.10 s 14× 1.6× 0.16× ✓
32 0.031 4.54 s 3.5× 0.39× ✓ 0.04× ✓
80 0.013 11.4 s 1.4× 0.15× ✓ 0.02× ✓

At realistic model depths the scheme becomes affordable: for a 32-layer model on cheap egress, verification costs 39% of re-executing; at 80 layers, 15%.

What is paid for it is soundness: the chance of catching a cheat on any one layer falls to 1/L. Verification stops being near-certain and becomes an audit — exactly the transition threat-model.md §2.5 anticipated.

Built in E024

This section's design is no longer arithmetic. ../024-layer-audit/ implements it — Merkle commitment, uniform post-commitment challenge, opening, recomputation — and runs it against Qwen2.5-0.5B. Honest work is accepted on all 24 layers, a single-layer cheat is detected at 4.5% against the predicted 1/L = 4.2%, and a provider that cheats everywhere and then answers the challenge honestly is rejected by the commitment.

Two details here were wrong. The opening does not need to carry the layer's output — the verifier recomputes it and checks its own digest against the root, which halves the traffic and is strictly more secure. And layer 0 is free: its input is the embedding, which the verifier derives itself. Audit traffic comes to 4.2% of verifying every layer.

4. What stake makes cheating unprofitable

If security is now probabilistic, it needs an economic backstop, and the required size falls out cleanly.

A provider skips m of L layers. It saves m/L of the job's compute value V. The coordinator audits one uniformly random layer, so it is caught with probability m/L. Cheating is negative-expected-value when

    penalty · (m/L)  >  V · (m/L)      ⟹      penalty > V

The m and the L cancel. DERIVED result:

A stake equal to the value of the job's compute makes layer-skipping negative-expected-value, regardless of how many layers are skipped and regardless of the model's depth.

That is an unusually convenient number — it is independent of depth and of the attacker's greed, and it is implementable with machinery Nodus v1 already has (escrow, delayed settlement, integer micro-credits).

Assumptions, which are real: the provider is risk-neutral; the penalty is actually collectable; the cheat is detectable once audited (true here, because exact arithmetic has no tolerance); and the audited layer is chosen unpredictably, after the provider commits.

5. Revised position

verify every layer, transmit intermediates economically dead — loses in every configuration measured
verify one random layer + stake ≥ job value viable at L ≥ 8–32, depending on egress price
interactive proofs (GKR/sumcheck) the known way to remove the bandwidth entirely — not yet measured by us
re-execution still the baseline, and still unbeaten when the verifier is not co-located and the model is shallow

Hypothesis H4 — that most practical security comes from cheap mechanisms plus stake, with cryptography reserved for a tail — is now supported by numbers rather than intuition. It arrived from an unexpected direction: not because cryptography was too slow, but because data movement was.

6. What a malicious provider can still get away with

  • Cheating on L-1 layers and being caught with probability 1/L — bounded by the stake, not by the mathematics.
  • Cheating on the layer least likely to be audited, if the audit distribution is predictable or if some layers are more expensive to audit than others. Uniform selection matters and must be enforced.
  • Everything from E013: caching, sub-contracting, being efficient.
  • Corrupting a layer after it has been audited, if the audit is not bound to the same execution as the result.

7. Limitations

  • Compressed with zlib, one codec. Domain-specific coding of int32 activations would do better; how much is unmeasured.
  • Our intermediates are int64 where int32 suffices; compression recovers most of that but a real implementation should just use int32.
  • Bandwidth prices are DECLARED (benchmarks/cost_model.py) and vary widely. The qualitative conclusion survives an order of magnitude in either direction; the L≥8 vs L≥32 threshold does not.
  • The layer-audit analysis is DERIVED arithmetic from measured quantities, not a separate measurement. It assumes uniform per-layer cost.
  • We have not measured GKR. The claim that it removes this cost is from the literature, not from our bench.

1. Measure an interactive proof (GKR / sumcheck) [now the top priority]

This experiment found a cost that GKR-class protocols are specifically designed to remove: they verify a layered computation with communication polylogarithmic in the circuit, instead of linear in the intermediates. Everything we have built pays the linear cost.

Measure, against our own baselines:

  • prover overhead versus native (E001's zkVM was ~7 × 10^4; Freivalds was 0)
  • communication — the quantity that actually decides this
  • verifier time
  • whether it composes with the exact-integer contract (it should; sumcheck is field arithmetic and our operands are already field elements)

Read SafetyNets (Ghodsi, Gu, Garg, NeurIPS 2017) first — it is the original statement of our exact problem and it is already in docs/literature.md.

2. Re-scope the earlier claims

R_verify = 0.023 (E003) and the E015 layer numbers are compute ratios that assume a co-located verifier. That qualifier should appear next to them wherever they are quoted, not only in this analysis. Partly done; audit the docs.

3. Test the layer-audit scheme end to end

The rescue in §3 of the analysis is derived arithmetic, not a running system. Build it: commit to per-layer outputs, have the coordinator choose one layer uniformly after the result is submitted, transmit only that layer's intermediates, verify, settle or slash. Measure real bandwidth and latency rather than modelled bytes.

Watch for: the provider must commit to all layer outputs before learning which is audited, or it can fix up the audited one. That is a Merkle root over layer outputs — cheap, and it is the piece the current design lacks.

4. Better encoding of intermediates

zlib is a generic codec and gave 2.4×. Activations are int32 with a narrow dynamic range and strong structure; domain-specific coding should do considerably better. Worth an afternoon, because bandwidth is now the binding constraint and a 4× here moves the L threshold from 32 to 8.

5. The unpaid bills, unchanged

  • Accuracy cost of integer inference — still UNKNOWN, still the price of the whole direction.
  • Cross-vendor validation — still one Apple machine.