# Architecture

**Status:** design document; the target design is a hypothesis, not a decision.
Last revised 2026-09-01.

---

## 1. Today (Nodus v1)

```
consumer ──submit──► coordinator ──assign──► provider node
                      │  escrow                 │ llama.cpp inference
                      │                         ▼
                      │◄────── result + HMAC receipt ──────┘
                      ├─ BasicVerifier: 11 checks  → Claim O only (L1)
                      └─ settle: CU from *attested* token counts → CC
```

Trust: coordinator fully trusted; provider trusted for correctness *and* for
metering. See `mvp-analysis.md`.

---

## 2. The insertion points

The useful property of v1 is that the verification layer is already an
interface, and there are exactly four places to intervene. Naming them makes
every experiment's integration story concrete.

```
      ┌──────────────── (1) STATEMENT BINDING ─────────────────┐
      │  what, exactly, is the provider being asked to do?      │
      │  model commitment · input commitment · freshness nonce  │
      └─────────────────────────┬───────────────────────────────┘
                                ▼
      ┌──────────────── (2) EVIDENCE PRODUCTION ────────────────┐
      │  what does the provider emit alongside the result?      │
      │  receipt (v1) · trace commitments · proof · TEE quote   │
      └─────────────────────────┬───────────────────────────────┘
                                ▼
      ┌──────────────── (3) EVIDENCE CHECKING ──────────────────┐
      │  Verifier interface -- already exists in v1             │
      │  synchronous check · deferred audit · challenge round   │
      └─────────────────────────┬───────────────────────────────┘
                                ▼
      ┌──────────────── (4) ECONOMIC RESPONSE ──────────────────┐
      │  escrow release · delayed settlement · stake · slashing │
      └─────────────────────────────────────────────────────────┘
```

v1 implements (2) and (3) at L1 and (4) as immediate settlement. It implements
almost none of (1), which is why model substitution and cached answers are
invisible to it.

**Observation worth acting on early:** (1) is nearly free. A commitment to the
weights and a freshness nonce folded into the committed input cost microseconds
and no cryptographic machinery beyond what v1 already has (SHA-256). They do not
by themselves *detect* anything — but they make every later mechanism able to
detect model substitution and replay. **Statement binding is a prerequisite for
all the expensive machinery, and it is the cheapest thing in the system.**

---

## 3. Target shape (HYPOTHESIS, tests H6)

Not a decision. The hypothesis is that the answer is *tiered by job value*,
because evidence cost is roughly constant per job while job value is not.

```
 job value ──►  low                    medium                 high
              ┌─────────────┐      ┌─────────────┐      ┌──────────────┐
 evidence     │ bind + audit│      │ bind + TEE  │      │ bind + proof │
              │ L1 + L3     │      │ L4          │      │ L5           │
 R_verify     │ ~audit rate │      │ small       │      │ large        │
 claims       │ O,A*,B*     │      │ O,A,B,C~    │      │ O,A + privacy│
 backstop     │ stake       │      │ vendor      │      │ math         │
              └─────────────┘      └─────────────┘      └──────────────┘
                    ▲ expected to carry the overwhelming majority of jobs
```

`A*`/`B*` denote probabilistic support with a stated error.

Three design commitments that follow from the analysis so far, each falsifiable:

1. **Statement binding is universal.** Every tier gets model commitment, input
   commitment and a freshness nonce. Cost ≈ 0.
2. **Metering is derived, never reported.** CU comes from the job spec and the
   committed I/O. This removes Claim C/D from the critical path entirely — see
   `research-thesis.md` §4. If this holds, Nodus never needs to solve the
   hardest claim in the taxonomy.
3. **Settlement is delayed and stake-backed.** Probabilistic verification has no
   force without a penalty; escrow-then-settle already gives us the window.

---

## 4. What would break this shape

- If heterogeneous honest divergence (E009-pre) is high, the L3 audit tier
  loses its threshold and the low-value tier collapses into "trust the
  provider". This is the single biggest risk to the whole design and is why
  E009-pre was pulled forward.
- If consumers require weight privacy, the L5 tier stops being the expensive
  option for high-value jobs and becomes the *only* option for a whole class of
  them, regardless of cost.
- If Claim D turns out to matter economically (e.g. providers must prove
  physical capacity to price capacity forwards), the tiering above is
  insufficient and Stage 9/11 becomes central rather than exploratory.


---

## 5. Inline verification vs optimistic verification (added 2026-09-01)

EigenAI (`literature.md` §5b) ships the architecture this programme's evidence
points at — determinism, re-execution, stake — with one difference that matters:
**verification is optimistic**, invoked only under dispute. Their claim is that
steady-state verification cost then approaches zero.

That is a real challenge to the value of everything we measured in E003, E015
and E016. If checks are rare, cheap checks are worth little.

**Why it does not settle the question for Nodus.**

| | optimistic | inline |
|---|---|---|
| steady-state verification cost | ~0 | 4.6%–40% of compute (E003, E015) |
| settlement | **delayed by a challenge window** | immediate |
| requires publishing data so anyone can challenge | **yes** (EigenAI uses EigenDA) | no |
| provider capital | bonded for the window | escrow only |
| resolves disputes by | full re-execution | the check itself |

Nodus v1's measured coordinator overhead is **3 ms median** and settlement is
immediate. Adopting an optimistic scheme replaces that with a challenge window
of minutes to days and requires publishing every job for challengers. That is a
**product change**, not only a security change, and it is the reason inline
verification remains worth pricing for this network even though a credible
system has chosen otherwise.

**The honest synthesis, and it is a hypothesis not a result:**

- determinism is required by *both* designs — and is the one thing we have
  measured a *portable* solution for (E014/E015), where EigenAI has an
  architecture-locked one;
- stake is required by both (E016 §4 gives the threshold: a stake equal to the
  job's compute value);
- the inline-vs-optimistic choice is then a **latency and capital** question,
  not a cryptography question.

This reframes CP-010: the hybrid to design is not "which proof system" but
"which settlement model, on top of a portable numeric contract".


---

## 6. Two candidate architectures (added 2026-09-02, E019 + E020)

The programme now has two coherent designs for a verifiable transformer, with
the trade-off between them measured rather than argued.

| | **keep the model** | **change the model** |
|---|---|---|
| non-linearities | integer lookup tables (NC-9) | polynomial: `x²` FFN, squared attention |
| verification | Freivalds per matmul + recomputation | sumcheck over the whole chain |
| communication / layer | **49 MB** (E016) | **~800 bytes** (E019) |
| quality cost | **+1.4% perplexity** (E018) | **+4.2% perplexity** (E020), plus the contract's own if quantised |
| needs an audit + stake? | **yes** — soundness falls to 1/L | no |
| cross-backend exact? | yes | yes, if also under the contract |
| works on any transformer? | **yes** | no — requires retraining |

**Where the cost sits is the useful part.** E020 found that replacing GELU with
`x²` is *free* (−0.5%, inside noise). The entire price of provability is in
**attention**: softmax's exponential is what a sumcheck cannot pass through, and
replacing it is what costs 4.2%.

### DECIDED, 2026-09-02 (E020b)

The outstanding measurement has been made, and it is decisive.

| layers | softmax | squared attention | gap |
|---:|---:|---:|---:|
| 2 | 5.6240 | 6.0060 | +6.8% |
| 4 | 5.3730 | 5.7282 | +6.6% |
| 8 | 5.2468 | 6.2839 | **+19.8%** |

**Softmax improves with depth; squared attention stops improving and then
regresses.** The gap is a divergence, not a fixed tax a bigger model amortises.

> **Decision: the right column is closed. Nodus keeps the model and pays the
> bandwidth.**
>
> NC-0.3 lookup-table contract + one-random-layer audit + a stake equal to the
> job's compute value. +1.4% perplexity, works on an unmodified transformer, and
> costs 39% of re-execution for a 32-layer model.

The sumcheck's prover overhead (E019 next-steps §2) is no longer on the critical
path, since the architecture it would serve is closed. It remains worth
measuring for the record.

Caveat carried forward: all runs used a fixed 1,500 steps, so some of the depth-8
gap could be under-training rather than a capacity limit. The direction is solid
— it holds across both depth and context — but the magnitude at depth 8 should be
re-measured at convergence before it is quoted as a headline.
