Literature
What already exists, including the results we should have found sooner.
Status: living document, continuously updated. Last revised 2026-09-01.
Conventions:
- [V] = primary source retrieved and read directly by this project.
- [C] = confirmed to exist via the reference list of a [V] source, but not yet read here. Do not cite a [C] entry's numbers — only its existence.
- [U] = unverified/secondary. Never cite for a claim.
- Every performance figure is tagged MEASURED (by the cited authors), DERIVED (computed here from their figures), or DECLARED.
1. The anchor result: ZIP (CCS 2025) [V]
Arman Riasi, Haodi Wang, Rouzbeh Behnia, Viet Vo, Thang Hoang. Zero-Knowledge AI Inference with High Precision. ACM CCS 2025. ePrint 2025/1732 — https://eprint.iacr.org/2025/1732 — code: https://github.com/vt-asaplab/ZIP Local copy:
research/papers/eprint-2025-1732-zip.pdf
What ZIP actually is
A commit-and-prove zkSNARK for ML inference that natively supports IEEE-754 double-precision floating point, rather than the fixed-point quantisation that prior ZKML work relies on. Its core technique is a relative-error-driven method for proving piecewise-polynomial approximations of non-linear layers, plus hardened lookup-argument and range-proof constraints. Backend: PlonK via gnark, with KZG replaced by PolyCommitPed (unconditionally hiding, SDH); lookups via Caulk; approximations generated with NFGen. ~11,000 LOC Python/Go/Rust.
What ZIP proves — in taxonomy terms
Claim A (correctness) at level L5, with model binding (the weights are committed in advance and the commitment is broadcast), plus zero-knowledge for the weights. This is a genuinely strong claim-set — stronger than most ZKML work, because the commitment makes model substitution (threat T3/T4) detectable.
What ZIP does not prove
- B — nothing binds the proof to a fresh execution; a cached (input, output, proof) triple verifies forever.
- C / D — the proof carries no information about resources consumed.
- It does not prove the committed model is a good model, only that it is the committed one.
Measured numbers (MEASURED by the authors)
Hardware: 48-core Intel Xeon Platinum 8360Y @ 2.40 GHz, 512 GB RAM. CPU only.
LeNet-5 / MNIST (~60 K params), Table 5, prover P / verifier V, minutes:
| activation | P (min) | V (min) | constraints |
|---|---|---|---|
| GeLU | 20.67 | 5.89 | 32,451,489 R1CS / 66,133,656 PlonK |
| SeLU | 15.91 | 5.87 | 19,749,473 R1CS / 54,765,423 PlonK |
| ELU | 13.94 | 5.91 | 15,703,492 R1CS / 42,887,728 PlonK |
Without random-linear-combination compression of the linear layers (i.e. at full numerical precision) end-to-end proving is 40 / 42 / 47 min respectively.
mini-BERT / SST-2 (4 layers, hidden 256, ~11 M params), Table 7, hours:
| component | P (hr) | V (hr) |
|---|---|---|
| GeLU | 2.44 | 0.73 |
| Softmax (non-linear part) | 0.09 | 0.03 |
| LayerNorm (non-linear part) | 0.003 | 0.001 |
| all linear layers | 34.53 | 0.09 |
| Total | 37.06 | 0.85 |
Accuracy, Table 6: IEEE-754 double 85.65% vs fixed-point baseline 77.14% — an 8.5-point accuracy recovery, which is ZIP's substantive contribution and should not be understated.
The economic reading (DERIVED by us — the important part)
A mini-BERT forward pass is a millisecond-scale computation. Taking a deliberately conservative native baseline of 10 ms on ordinary CPU:
| quantity | value | DERIVED |
|---|---|---|
proving overhead R_prove |
37.06 hr / 10 ms | ≈ 1.3 × 10^7 × |
verification overhead R_verify |
0.85 hr / 10 ms | ≈ 3.1 × 10^5 × |
The second row is the one that matters for Nodus. ZIP's verifier is roughly five orders of magnitude more expensive than simply re-running the model. A coordinator that can afford 0.85 hr of verification can afford to re-execute the job ~300,000 times. On the cost axis alone, for this model size, ZK is dominated by re-execution by a very wide margin.
That is not a criticism of ZIP, and the authors say as much themselves (§Discussion): "our prototype is just a proof of concept to demonstrate correctness and feasibility, not performance optimization", run on a general-purpose CPU with untuned libraries. They project ~100× from distributed proving (DIZK/Pianist-style) and ≥5× prover / 2× verifier from GPU acceleration. Even granting the full ~500× prover improvement optimistically and compounding the verifier gains, mini-BERT proving lands around 4–5 minutes and verification in the minutes range — still far above re-execution for an 11 M-parameter model.
The conclusion we carry forward is not "ZK is useless". It is:
- ZK buys things re-execution cannot: succinctness for a third party, and privacy of the weights. Where those are required, the comparison changes completely.
- Where they are not required — which is the default case for Nodus's coordinator, who holds the model and could just re-run it — ZK must close ~5 orders of magnitude on verification cost before it is the rational choice.
- Note the shape of Table 7: 93% of the proving cost is the linear layers (34.53 of 37.06 hr), even though those are handled by the cheap method (Freivalds-style RLC, [24]/[20]). The non-linearities that ZIP optimises are no longer the bottleneck once ZIP's own technique is applied. This is a strong hint about where the remaining research value is: linear algebra at scale, not activation functions. It is also a direct endorsement of the choice of matrix multiplication as our first research question.
2. Probabilistic verification of linear algebra
- [C] Rūsiņš Freivalds. Probabilistic Machines Can Use Less Running Time.
IFIP Congress 1977, 839–842. — ZIP ref [24]. Verifies
A·B = Cforn×nmatrices inO(n^2)time with one-sided error≤ 2^-kafterkrandom vectors, vsO(n^ω)to recompute. This is the canonical example ofR << 1and is the baseline for Experiment 003. - [C] Ernstberger, Zhang, Ciprian, Jovanovic, Steinhorst. Zero-Knowledge Location Privacy via Accurate Floating-Point SNARKs. arXiv:2404.14983 (2024). — ZIP ref [20]; the method ZIP uses for all its linear layers.
- [C] Riasi, Guajardo, Hoang. Privacy-Preserving Verifiable Neural Network Inference Service. ACSAC 2024. — ZIP ref [51].
- [C] Cormode, Thaler, Yi. Verifying computations with streaming interactive proofs. arXiv:1109.6882 (2011). — ZIP ref [13].
- [C] Vu, Setty, Blumberg, Walfish — ZIP ref [63]; the "hybrid architecture for interactive verifiable computation" line of work. Directly relevant to our Stage 10 hybrid question.
3. ZKML systems (specialised proofs for neural networks)
-
[C] Liu, Xie, Zhang. zkCNN: Zero-knowledge proofs for convolutional neural network predictions and accuracy. CCS 2021. — ZIP ref [39]. GKR-based.
-
[C] Sun, Li, Zhang. zkLLM: Zero-knowledge proofs for large language models. CCS 2024. — ZIP ref [57]. The closest published work to Nodus's eventual target. ZIP reports its lookup tables need 65,536 entries.
-
[C] Chen, Waiwitlikhit, Stoica, Kang. ZKML: An optimizing system for ML inference in zero-knowledge proofs. EuroSys 2024. — ZIP ref [12].
-
[C] Hao et al. Scalable zero-knowledge proofs for non-linear functions in machine learning. USENIX Security 2024. — ZIP ref [30]. Fixed-point; 0.38 ms/GeLU with 4,096-entry LUT (vs ZIP's 240 ms with 70 entries at double precision — the precision/cost trade-off in one line).
-
[C] Lu et al. An efficient and extensible zero-knowledge proof framework for neural networks. ePrint 2024. — ZIP ref [41]. 0.025 ms/GeLU, fixed-point.
-
[C] Feng, Qin, Zhang, Ding, Chu. ZEN: An optimizing compiler for verifiable, zero-knowledge neural network inferences. ePrint 2021. — ref [23].
-
[C] Lee, Ko, Kim, Oh. vCNN: Verifiable convolutional neural network based on zk-SNARKs. IEEE TDSC 2024. — ref [37].
-
[V] Zahra Ghodsi, Tianyu Gu, Siddharth Garg. SafetyNets: Verifiable Execution of Deep Neural Networks on an Untrusted Cloud. NeurIPS 2017, 4672–4681. arXiv:1706.10268. Local copy:
research/papers/safetynets-nips2017.pdfRead in full 2026-09-01. The most directly relevant prior work in this bibliography, and it anticipates three of our results.Specialised interactive proof (Thaler's IP for matrix multiplication, plus their own IP for quadratic activations) for DNNs expressible as arithmetic circuits over
F_p.MEASURED by the authors (Intel Core i7-4600U @ 2.10 GHz, FcNN-Quad-3, batch 256–2048):
quantity value verifier vs executing the network itself 8× to 82× faster (improves with batch size) prover overhead over unverified execution 5% communication, batch 2048 < 8 KB — under 2% of the input/output traffic soundness error < 2^-30 accuracy MNIST 0.63% test error; MNIST-Back-Rand 4.67%; TIMIT 25.7% Restrictions: quadratic activations instead of ReLU, sum pooling instead of max pooling, and "requires weights and inputs to be integers (in field
F_p)". Scaling factors α, β are capped at 64 because the maximum value must stay belowp = 2^61 − 1(their Table 1 marks infeasible cells in red).Three convergences with our own work, recorded because they matter:
- Their integer/field requirement is our numeric contract (E014, E015,
numeric-contract.md), reached independently and eight years earlier, for a different reason: their proof system needs an arithmetic circuit over a field, ours needs cross-backend determinism. Two different routes to "constrain the arithmetic" is meaningful support for H7. - Their α/β cap is our NC-3 tiling clause — the same accumulation-magnitude constraint, expressed as a bound on scaling rather than on tile length.
- Their §4.3 states our E016 result in one sentence: "SafetyNets has
significantly lower bandwidth costs compared to an approach that separately
verifies the execution of each layer using only the IP protocol for matrix
multiplication." That approach is exactly what we built and measured.
See the note in
../experiments/016-verification-bandwidth/analysis.md§2a.
Where our approach differs and is arguably better: SafetyNets changes the model (quadratic activations are not GELU or ReLU), so accuracy is a negotiation. Our NC-9 lookup-table specification changes only the representation — GELU stays GELU to within the table's resolution. Less invasive, at the cost of not being provable by their IP.
Where it is clearly better than anything we have: prover overhead 5% against our zkVM's 7 × 10^4, and communication < 8 KB against our measured 49 MB per layer. Those two numbers are the case for interactive proofs.
Open question, and the user's own observation on reading it: the results are on 3- and 4-layer networks on MNIST and TIMIT, in 2017. Whether the approach — and specifically the quadratic-activation restriction — survives transformer-scale models with softmax attention is UNKNOWN, and is the obvious next experiment.
- Their integer/field requirement is our numeric contract (E014, E015,
-
[C] Garg, Jain, Jin, Zhang. Succinct zero knowledge for floating point computations. CCS 2022. — ref [27].
4. Proof systems and infrastructure
- [C] Gabizon, Williamson, Ciobotaru. PlonK. ePrint 2019/953. — ref [25].
- [C] Chen, Bünz, Boneh, Zhang. HyperPlonk. EUROCRYPT 2023. — ref [11].
- [C] Kate, Zaverucha, Goldberg. Constant-size commitments to polynomials. ASIACRYPT 2010 (KZG). — ref [35].
- [C] Gennaro, Gentry, Parno, Raykova. Quadratic span programs. EUROCRYPT 2013. — ref [28].
- [C] Parno, Howell, Gentry, Raykova. Pinocchio. CACM 2016. — ref [46].
- [C] Bünz et al. Bulletproofs. IEEE S&P 2018. — ref [7].
- [C] Zapico et al. Caulk — ZIP ref [71]/[72]; position-hiding linkability for vector commitments.
- [C] Botrel et al. gnark v0.12.0, Zenodo. — ref [6].
- [C] Wu et al. DIZK — ref [69]; distributed proof generation.
- [C] Liu, Xie, Zhang, Song, Zhang. Pianist: Scalable zkRollups via fully distributed zero-knowledge proofs. IEEE S&P 2024. — ref [38].
zkVMs (our general-purpose baselines) [V] — installed and running here
- SP1 (Succinct). Repo https://github.com/succinctlabs/sp1, docs
https://docs.succinct.xyz. RISC-V zkVM. Installed version in this repo:
cargo-prove sp1 (92b8eab 2026-08-26), SDK crates 6.5.0, guest targetriscv64im-succinct-zkvm-elf. Used in Experiment 001. - RISC Zero. https://github.com/risc0/risc0. NOT yet installed — planned as the second zkVM baseline so that Experiment 001/002 conclusions are not an artefact of one implementation.
5. Hardware acceleration of proving
- [C] Lu et al. cuZK: Accelerating ZKP with faster parallel MSM on GPUs. TCHES 2023(3). — ref [42].
- [C] Ma et al. GZKP: A GPU accelerated zero-knowledge proof system. ASPLOS 2023. — ref [43].
- [C] Daftardar, Reagen, Garg. SZKP: A scalable accelerator architecture for zero-knowledge proofs. PACT 2024. — ref [16].
- [C] Daftardar et al. Need for zkSpeed: Accelerating HyperPlonk. ISCA 2025. — ref [15].
- [C] Samardzic, Langowski, Devadas, Sanchez. Accelerating zero-knowledge proofs through hardware-algorithm co-design. MICRO 2024. — ref [54].
- [C] Pottier et al. OPTIMSM: FPGA hardware accelerator for ZK MSM. TCHES 2025(2). — ref [48].
- [C] Butt et al. if-ZKP: Intel FPGA-based acceleration of ZKP. arXiv:2412.12481. — ref [8].
These are the primary inputs to Stage 11 / Direction 8 ("verifiable computation
per dollar" rather than "FLOPS per dollar"). Note that all of them accelerate
proving, i.e. they attack R_prove. None of them changes the fact that a
correctness proof says nothing about Claims C or D.
5a. Forcing work: proof-of-work, useful work, sequential work, space
Added 2026-09-01 for Experiment 013 (CP-011). This is the branch that speaks to Claim C/D rather than Claim A, and it is the one most likely to be assumed unsolved when it is not.
Proof of useful work
-
[V] Ilan Komargodski, Omri Weinstein. Proofs of Useful Work from Arbitrary Matrix Multiplication. ePrint 2025/685. https://eprint.iacr.org/2025/685 A PoUW for matrix multiplication with 1 + o(1) multiplicative overhead over naïve matmul — directly on the primitive Nodus cares about, and remarkably cheap. Crucially: it produces a PoW certificate with prescribed hardness, it does not verify correctness of the product, and its security is conjectured, reduced to a new assumption about solving batches of low-rank random linear equations. The authors also note that fast-matmul-style algorithms are "currently impractical", so part of the security rests on practical rather than theoretical barriers. Deserves a full read; a candidate for its own experiment.
-
[V] Pratyush Dikshit, Ashkan Emami, Johannes Sedlmeir, Gilbert Fridgen. SoK: Is Proof-of-Useful-Work Really Useful? ePrint 2025/1814. Local copy:
research/papers/eprint-2025-1814-sok-pouw.pdfSurveys more than 50 PoUW constructions, builds a taxonomy, and adds a formal economic model of miner incentives and security budget. Conclusion, quoted: "Our finding show that PoUW is actually not as useful as expected, since the economic and societal utility do not contribute to the security budget." The negative economic result is the important part for us. -
[C] Marshall Ball, Alon Rosen, Manuel Sabin, Prashant Nalini Vasudevan. Proofs of Useful Work. ePrint 2017/203; and Proofs of Work From Worst-Case Assumptions, CRYPTO 2018. PoWs whose hardness reduces to Orthogonal Vectors, 3SUM and All-Pairs Shortest Paths. The strongest "useful work" foundations available, and still assumption-based.
-
[C] Ofelimos: Combinatorial Optimization via Proof-of-Useful-Work, CRYPTO 2022.
Proof of sequential work / delay — forces expenditure, of the wrong kind
- [C] Dan Boneh, Joseph Bonneau, Benedikt Bünz, Ben Fisch. Verifiable Delay Functions. CRYPTO 2018; ePrint 2018/601.
- [C] Bram Cohen, Krzysztof Pietrzak. Simple Proofs of Sequential Work. EUROCRYPT 2018; ePrint 2018/183. Efficient PoSW without depth-robust graphs.
- [C] Mohammad Mahmoody, Tal Moran, Salil Vadhan. Publicly Verifiable Proofs of Sequential Work. ITCS 2013. The original construction.
These genuinely prove that time was spent — by being deliberately unparallelisable and deliberately useless. Nodus wants work that is massively parallel and useful. Structurally opposed requirements; recorded so we do not re-derive this.
- [C] Stefan Dziembowski, Sebastian Faust, Vladimir Kolmogorov, Krzysztof Pietrzak. Proofs of Space. CRYPTO 2015. Proves storage, not compute.
Proof-of-Learning — the direct ML analogue, and its breaks
- [C] Hengrui Jia et al. Proof-of-Learning: Definitions and Practice. IEEE S&P 2021. Literally "prove you spent the compute to train this model".
- [C] Rui Zhang et al. "Adversarial Examples" for Proof-of-Learning. arXiv:2108.09454. Stochastic spoofing: the adversary generates dummy weights and then adversarial examples that make an arbitrary data point "generate" a given model, producing a proof that verifies cheaply.
- [C] Fang et al. (2023). Systemic vulnerabilities of PoL, argued to require advances in understanding optimisation before they can be addressed.
This is the closest published attempt at proof of computational expenditure for ML, and it was broken. Any Nodus design in this direction must engage with it.
The barrier underneath all of it: no lower bounds
- [V, via secondary] The matrix multiplication exponent ω is unknown and
still falling: 2.371866 (Alman–Vassilevska Williams 2020) → 2.371552 (Jan
2024) → 2.371339 (Duan–Wu–Zhou / Vassilevska Williams et al. 2024) →
< 2.371177 (2026, arXiv:2608.16884, laser method plus ML-driven
optimisation). Nobody can prove an
n×nproduct requires more thann²work.
Consequence, and the reason Experiment 013 refused the "force the provider to do the work" framing: every claim of the form "the provider must have expended this work" rests on an unproven hardness conjecture, not a theorem. There are no unconditional superlinear lower bounds for natural problems; this is the same wall as P vs NP.
5b. Deployed systems on the same thesis
-
[V] David Ribeiro Alves, Vishnu Patankar, Matheus Pereira, Jamie Stephens, Nima Vaziri, Sreeram Kannan. EigenAI: Deterministic Inference, Verifiable Results. arXiv:2602.00182, 30 Jan 2026. Local copy:
research/papers/eigenai-2602.00182.pdfThe closest thing to a production system built on this programme's thesis, and its stated open problem is the one we have measured a solution to.Architecture: deterministic LLM inference + optimistic re-execution + slashing against EigenLayer restaked capital, with TEEs and threshold key release for privacy. Their framing of why determinism is load-bearing matches ours exactly: "inference itself is nondeterministic on modern GPUs. Two identical queries to the same model can yield divergent outputs because of floating-point non-associativity, kernel scheduling, and variable batching. Without reproducibility, verification through re-execution ... is impossible."
How they achieve determinism: by controlling the environment, not the arithmetic — fixed GPU architecture (H100), custom kernels, version-pinned drivers and containers, canonical reduction orders, batch-invariant reduction kernels, GPU clock locking. MEASURED: bit-identical logits and tokens on the same host and across two independent H100 nodes.
Their stated limitation, quoted in full because it is our opening:
"Cross-Architecture Reproducibility. Determinism currently holds only within fixed GPU families. Future work includes portable numeric normalization to enable heterogeneous verifier sets."
That is precisely what E014/E015 measured. Exact integer semantics make CPU and GPU agree bit-for-bit with no supply cost — a portable numeric normalisation, arrived at independently. Nodus's premise is heterogeneous consumer hardware, so a single-architecture policy is a non-starter for us and our contract addresses their gap rather than duplicating their solution.
What they get right that we had underweighted: verification is optimistic — invoked only under dispute — so "the steady-state cost approaches that of normal inference". If verification is rare, making it cheap matters far less than making it unambiguous. This is a genuine challenge to the value of our cheap-verification line and is answered in
architecture.md§5.Third independent arrival at "constrain the arithmetic": SafetyNets (2017, for field arithmetic), EigenAI (2026, for reproducibility), this programme (2026, for cross-backend agreement). H7 now has three witnesses.
6. Trusted execution / attestation
- [C] Costan, Devadas. Intel SGX Explained. ePrint 2016/086. — ref [14].
- [U→V pending] NVIDIA Confidential Computing (Hopper H100 onward), official product/technical documentation: https://www.nvidia.com/en-us/data-center/solutions/confidential-computing/. Attestation is a signed report over firmware measurements plus a workload hash, validated against NVIDIA's remote attestation service (NRAS). Trust set: NVIDIA (device certs, firmware, and the attestation service).
- [U→V pending] Confidential Computing on nVIDIA H100 GPU: A Performance Benchmark Study, arXiv:2409.03992 — a primary measurement of CC-mode overhead; to be read and summarised before Experiment 010.
- Known limitation to carry: current GPU CC implementations do not claim resistance to physical attacks, and TEEs generally have a history of microarchitectural side-channel breaks. A TEE is an L4 mechanism whose trust set includes a vendor and its supply chain.
7. Explicit gaps — searched for, not yet found
Recorded so we do not "discover" something that exists, and do not assume something exists that does not. Each is UNKNOWN until resolved.
| # | Gap | Status |
|---|---|---|
| G1 | A verification scheme that establishes Claim D (physical resource consumption) without trusted hardware | searched (E013); none found. Software can force work under hardness conjectures (§5a), but a proof carries no information about the machine that produced it. H5 survives |
| G2 | Published false-positive rates for output-hash replication across heterogeneous accelerators at temperature 0 | not found. We measured the matmul case ourselves (E009-pre); the decoding case is still open |
| G3 | A cost model comparing ZK / TEE / audit on the same workload with a re-execution baseline | not found; this is the gap this repository is aimed at |
| G4 | Freivalds-style checks applied to transformer inference end-to-end with tolerance analysis for floating point | partially covered by ZIP ref [20]/[24] usage; needs reading |
| G5 | Economic (game-theoretic) analysis of audit-rate vs stake for compute marketplaces | partially closed: the PoUW SoK (2025/1814) contains a formal economic model of security budget; not yet read in full |
| G6 | Whether a PoUW with 1+o(1) overhead (Komargodski–Weinstein) can be composed with a correctness proof, giving Claims A and C together | UNKNOWN — the two papers prove different things and nobody appears to have combined them |
| G7 | Whether SafetyNets-style interactive proofs survive transformer-scale models — specifically whether the quadratic-activation restriction can be replaced by our LUT specification while keeping the IP | UNKNOWN; the obvious next experiment |
| G8 | EigenAI: Deterministic Inference, Verifiable Results (arXiv:2602.00182, 2026) — surfaced while searching; title suggests it is directly on our determinism thread | not yet read |
8. Reading queue (ordered)
SafetyNets [29]— READ 2026-09-01, see §3. Anticipates our numeric contract, our NC-3 magnitude bound, and our E016 bandwidth result.- Freivalds [24] + Ernstberger et al. [20] — the technique doing 93% of the work in ZIP's own pipeline.
- zkLLM [57] — the scaling frontier.
- Thaler, Proofs, Arguments, and Zero-Knowledge — background for GKR/sumcheck.
- arXiv:2409.03992 — H100 confidential computing overhead.
- DIZK [69] / Pianist [38] — distributed proving, since Nodus is a distributed compute network and could prove its own jobs.
Item 6 is worth flagging as a possible original angle: Nodus is a network of provers. Distributed proof generation is normally a deployment detail; for us it is native infrastructure. Whether that changes the economics is UNKNOWN and is a candidate for Stage 7.