Architecture
How a verifiable job class would actually be wired.
Status: design document; the target design is a hypothesis, not a decision. Last revised 2026-09-01.
1. Today (Nodus v1)
consumer ──submit──► coordinator ──assign──► provider node
│ escrow │ llama.cpp inference
│ ▼
│◄────── result + HMAC receipt ──────┘
├─ BasicVerifier: 11 checks → Claim O only (L1)
└─ settle: CU from *attested* token counts → CC
Trust: coordinator fully trusted; provider trusted for correctness and for
metering. See mvp-analysis.md.
2. The insertion points
The useful property of v1 is that the verification layer is already an interface, and there are exactly four places to intervene. Naming them makes every experiment's integration story concrete.
┌──────────────── (1) STATEMENT BINDING ─────────────────┐
│ what, exactly, is the provider being asked to do? │
│ model commitment · input commitment · freshness nonce │
└─────────────────────────┬───────────────────────────────┘
▼
┌──────────────── (2) EVIDENCE PRODUCTION ────────────────┐
│ what does the provider emit alongside the result? │
│ receipt (v1) · trace commitments · proof · TEE quote │
└─────────────────────────┬───────────────────────────────┘
▼
┌──────────────── (3) EVIDENCE CHECKING ──────────────────┐
│ Verifier interface -- already exists in v1 │
│ synchronous check · deferred audit · challenge round │
└─────────────────────────┬───────────────────────────────┘
▼
┌──────────────── (4) ECONOMIC RESPONSE ──────────────────┐
│ escrow release · delayed settlement · stake · slashing │
└─────────────────────────────────────────────────────────┘
v1 implements (2) and (3) at L1 and (4) as immediate settlement. It implements almost none of (1), which is why model substitution and cached answers are invisible to it.
Observation worth acting on early: (1) is nearly free. A commitment to the weights and a freshness nonce folded into the committed input cost microseconds and no cryptographic machinery beyond what v1 already has (SHA-256). They do not by themselves detect anything — but they make every later mechanism able to detect model substitution and replay. Statement binding is a prerequisite for all the expensive machinery, and it is the cheapest thing in the system.
3. Target shape (HYPOTHESIS, tests H6)
Not a decision. The hypothesis is that the answer is tiered by job value, because evidence cost is roughly constant per job while job value is not.
job value ──► low medium high
┌─────────────┐ ┌─────────────┐ ┌──────────────┐
evidence │ bind + audit│ │ bind + TEE │ │ bind + proof │
│ L1 + L3 │ │ L4 │ │ L5 │
R_verify │ ~audit rate │ │ small │ │ large │
claims │ O,A*,B* │ │ O,A,B,C~ │ │ O,A + privacy│
backstop │ stake │ │ vendor │ │ math │
└─────────────┘ └─────────────┘ └──────────────┘
▲ expected to carry the overwhelming majority of jobs
A*/B* denote probabilistic support with a stated error.
Three design commitments that follow from the analysis so far, each falsifiable:
- Statement binding is universal. Every tier gets model commitment, input commitment and a freshness nonce. Cost ≈ 0.
- Metering is derived, never reported. CU comes from the job spec and the
committed I/O. This removes Claim C/D from the critical path entirely — see
research-thesis.md§4. If this holds, Nodus never needs to solve the hardest claim in the taxonomy. - Settlement is delayed and stake-backed. Probabilistic verification has no force without a penalty; escrow-then-settle already gives us the window.
4. What would break this shape
- If heterogeneous honest divergence (E009-pre) is high, the L3 audit tier loses its threshold and the low-value tier collapses into "trust the provider". This is the single biggest risk to the whole design and is why E009-pre was pulled forward.
- If consumers require weight privacy, the L5 tier stops being the expensive option for high-value jobs and becomes the only option for a whole class of them, regardless of cost.
- If Claim D turns out to matter economically (e.g. providers must prove physical capacity to price capacity forwards), the tiering above is insufficient and Stage 9/11 becomes central rather than exploratory.
5. Inline verification vs optimistic verification (added 2026-09-01)
EigenAI (literature.md §5b) ships the architecture this programme's evidence
points at — determinism, re-execution, stake — with one difference that matters:
verification is optimistic, invoked only under dispute. Their claim is that
steady-state verification cost then approaches zero.
That is a real challenge to the value of everything we measured in E003, E015 and E016. If checks are rare, cheap checks are worth little.
Why it does not settle the question for Nodus.
| optimistic | inline | |
|---|---|---|
| steady-state verification cost | ~0 | 4.6%–40% of compute (E003, E015) |
| settlement | delayed by a challenge window | immediate |
| requires publishing data so anyone can challenge | yes (EigenAI uses EigenDA) | no |
| provider capital | bonded for the window | escrow only |
| resolves disputes by | full re-execution | the check itself |
Nodus v1's measured coordinator overhead is 3 ms median and settlement is immediate. Adopting an optimistic scheme replaces that with a challenge window of minutes to days and requires publishing every job for challengers. That is a product change, not only a security change, and it is the reason inline verification remains worth pricing for this network even though a credible system has chosen otherwise.
The honest synthesis, and it is a hypothesis not a result:
- determinism is required by both designs — and is the one thing we have measured a portable solution for (E014/E015), where EigenAI has an architecture-locked one;
- stake is required by both (E016 §4 gives the threshold: a stake equal to the job's compute value);
- the inline-vs-optimistic choice is then a latency and capital question, not a cryptography question.
This reframes CP-010: the hybrid to design is not "which proof system" but "which settlement model, on top of a portable numeric contract".
6. Two candidate architectures (added 2026-09-02, E019 + E020)
The programme now has two coherent designs for a verifiable transformer, with the trade-off between them measured rather than argued.
| keep the model | change the model | |
|---|---|---|
| non-linearities | integer lookup tables (NC-9) | polynomial: x² FFN, squared attention |
| verification | Freivalds per matmul + recomputation | sumcheck over the whole chain |
| communication / layer | 49 MB (E016) | ~800 bytes (E019) |
| quality cost | +1.4% perplexity (E018) | +4.2% perplexity (E020), plus the contract's own if quantised |
| needs an audit + stake? | yes — soundness falls to 1/L | no |
| cross-backend exact? | yes | yes, if also under the contract |
| works on any transformer? | yes | no — requires retraining |
Where the cost sits is the useful part. E020 found that replacing GELU with
x² is free (−0.5%, inside noise). The entire price of provability is in
attention: softmax's exponential is what a sumcheck cannot pass through, and
replacing it is what costs 4.2%.
DECIDED, 2026-09-02 (E020b)
The outstanding measurement has been made, and it is decisive.
| layers | softmax | squared attention | gap |
|---|---|---|---|
| 2 | 5.6240 | 6.0060 | +6.8% |
| 4 | 5.3730 | 5.7282 | +6.6% |
| 8 | 5.2468 | 6.2839 | +19.8% |
Softmax improves with depth; squared attention stops improving and then regresses. The gap is a divergence, not a fixed tax a bigger model amortises.
Decision: the right column is closed. Nodus keeps the model and pays the bandwidth.
NC-0.3 lookup-table contract + one-random-layer audit + a stake equal to the job's compute value. +1.4% perplexity, works on an unmodified transformer, and costs 39% of re-execution for a 32-layer model.
The sumcheck's prover overhead (E019 next-steps §2) is no longer on the critical path, since the architecture it would serve is closed. It remains worth measuring for the record.
Caveat carried forward: all runs used a fixed 1,500 steps, so some of the depth-8 gap could be under-training rather than a capacity limit. The direction is solid — it holds across both depth and context — but the magnitude at depth 8 should be re-measured at convergence before it is quoted as a headline.