# Analysis of Nodus v1 (`../nodus/`)

**Status:** complete for v1 as of commit `1f296df`. Last revised 2026-09-01.

This is the "what do we already have" pass required before any new work. It is
deliberately unflattering where the evidence warrants it and deliberately
generous where v1 is already correct — v1's own documentation is unusually
honest about its limits, and several conclusions below are simply confirmations
of statements v1 already makes about itself.

---

## 1. What v1 is

A single-operator distributed inference network that closes one economic loop:

```
consumer  ->  coordinator (schedule, escrow)  ->  provider node (llama.cpp)
          <-  settle (ledger)                 <-  result + signed receipt
```

~6,900 LOC Python + 264 LOC dashboard; 130 tests (129 pass, 1 gated on real
inference). MEASURED production run: 14 jobs, 2 machines, 1 model
(qwen2.5-0.5b-instruct Q4_K_M), 100% verification pass, 38.165 CU settled,
ledger conserved, coordinator overhead median 3 ms on top of inference.

That last figure is the most useful single number v1 produced for *this*
project: **the economic layer costs ~3 ms per job today.** It is the budget
line that any verification mechanism eats into. A verification scheme costing
300 ms per job is a 100× regression in coordinator overhead even if it is
negligible next to a 500 ms inference.

## 2. Units

- **CU-v0**: 1 CU = 100 tokens of a fixed reference workload. Node capacity is
  MEASURED (219.25 tok/s on an M5, ±1% over n=5). Job cost is
  `(prompt_tokens × 0.1 + completion_tokens) × model_weight / 100`.
- **CC**: transferable credit, integer micro-credits, database ledger.
- `model_weight` is **DECLARED** in `models.yaml`, not measured. v1 says so.

## 3. What v1's verification actually establishes

`coordinator/verification/verifier.py::BasicVerifier` runs 11 checks. Mapped to
the taxonomy (`verification-taxonomy.md`):

| v1 check | Claim | Level |
|---|---|---|
| signature valid under node token | **O** | L1 |
| job_id / node_id match assignment | O | L1 |
| input_hash matches prompt sent | O (binds statement) | L1 |
| output_hash matches returned text | O (binds statement) | L1 |
| model / model_version match request | **none** — self-attested string | L0 |
| `completion_tokens <= max_tokens` | none (a bound, not evidence) | L0 |
| output/token-count mutual consistency | none | L0 |
| no prior result for this job | O (replay) | L1 |
| optional replication, output-hash equality | A, weakly | L2 |

**Conclusion: v1's `VERIFIED` status means Claim O and nothing more.** v1's own
docstrings say exactly this ("This is deliberately *not* a proof of
computation"), so this is a confirmation, not a criticism. The value of writing
it down in taxonomy terms is that it makes the *gap* precise: v1 has L1/Claim-O,
and every experiment here is measured by what it adds beyond that point.

The replication path is the only Claim-A evidence v1 has, and v1 correctly
records a mismatch as a **signal, not proof of fraud** because floating-point
non-determinism across backends can produce legitimate divergence. This is the
right call and it is also the reason naive replication will not scale to a
heterogeneous network: the false-positive rate is unquantified. Quantifying it
is a concrete, cheap, high-value experiment (see §6, E009-pre).

## 4. Finding: settlement trusts a recomputable quantity

**MEASURED (code read):** `coordinator/jobs.py:419-421`

```python
prompt_tokens = int(att.get("prompt_tokens") or 0)
completion_tokens = int(att.get("completion_tokens") or 0)
cost = cu_for_tokens(prompt_tokens, completion_tokens, spec.cu_weight)
```

Settlement multiplies the provider's **self-reported** token counts by a
declared weight, capped only by the escrow reserve. This is attack T11/T12 in
`threat-model.md`, and v1's own limitations section names it ("a malicious
provider can ... inflate self-reported prefill token counts up to the escrow
cap").

What makes this interesting rather than merely a bug is that **both quantities
are recomputable by the coordinator at near-zero cost**:

- `prompt_tokens` is a function of the prompt *the coordinator itself sent*.
- `completion_tokens` is a function of the returned text, which the coordinator
  already has and already hashes.

Both require one tokeniser pass over text the coordinator holds — O(length),
microseconds. The coordinator already contains a heuristic
`estimate_prompt_tokens` (`jobs.py:66`) but does not use it for settlement.

This is the thesis of the whole programme in miniature, at the smallest possible
scale:

> A quantity that the verifier can derive from data it already holds should
> never be an input the prover is trusted for.

It also demonstrates the §4 point of `research-thesis.md`: **once metering is
derived from the job specification and the committed I/O rather than from
provider self-report, Claim C collapses into Claim A.** The billing-fraud
surface disappears and the only remaining question is whether the output is
real. That is a strictly easier research problem, and it is reached by deleting
code rather than by adding cryptography.

**Recommendation to the v1 codebase (not yet applied, out of scope for this
repo's first commit):** derive both token counts coordinator-side with the
model's real tokeniser; keep the attested values only as a *consistency signal*
(a provider whose self-report diverges from the truth is flagged, not paid
differently). Estimated effort: small. Estimated cost: one tokeniser pass.
Estimated benefit: closes T11, T12 and most of T5.

## 5. Structural properties that help us

v1 got several things right that this research depends on:

1. **`Verifier` is an abstract base class** with a stable `JobFacts` /
   `VerificationResult` interface. New mechanisms plug in without touching the
   job engine. This is the integration point for everything in `experiments/`.
2. **Input and output are already hashed and bound into the receipt.** A
   commit-and-prove scheme needs exactly these commitments; v1 has them (SHA-256
   over text — adequate as a binding commitment, though not hiding).
3. **Escrow-then-settle** means there is already a natural place to hold funds
   pending *delayed* verification. Probabilistic audit schemes need exactly this
   (a settlement window in which a job can still be challenged).
4. **Audit re-runs are funded by the treasury, not the consumer.** The economic
   plumbing for a network-funded audit budget exists.
5. **Integer micro-credit accounting**, so penalties/slashing can be expressed
   exactly.

## 6. Gaps that define the research

| Gap | Consequence | Where addressed |
|---|---|---|
| No Claim-A evidence in the default path | fabricated output is paid | E001–E009 |
| No model binding | model substitution undetectable | E006, E012 |
| No freshness binding | cached/precomputed answers undetectable | E009, E012 |
| Metering trusts the prover | billing fraud up to escrow cap | §4, immediately fixable |
| Coordinator fully trusted | it could forge everything | out of scope until Stage 10+ |
| Replication false-positive rate UNKNOWN | replication cannot be enforced | **E009-pre — cheapest useful experiment available** |
| No stake / penalty | probabilistic schemes have no teeth | E009, E012 |

## 7. Immediate cheap experiment suggested by this analysis

> **ANSWERED (E018).** Greedy decoding, 10 prompts x 300 tokens, honest CPU vs
> honest GPU: **9/10 identical, 1 diverged at token 72.** A ~10% false-positive
> rate means v1's `REPLICATED_MISMATCH` **cannot be used for enforcement** on
> float inference — it would slash one honest provider in ten. Under the NC-0.3
> contract the rate is 0/10. This is the number this section asked for.

**E009-pre — "How often do two honest heterogeneous nodes disagree?"**
Run the same prompt at temperature 0 across the available backends
(Metal/llama.cpp on M5, CPU-only llama.cpp, simulated node) and measure the
distribution of output divergence: exact-match rate, first-divergent-token
index, and logit-level distance. Without this number, replication and audit
cannot be turned into enforcement, because we cannot set a threshold. v1
observed exactly one mismatch instance (against a simulated node) — n=1.

This costs hours, not weeks, requires no cryptography, and gates the entire
L2/L3 branch of the taxonomy. It has been added to the roadmap as a Stage-0
deliverable rather than waiting for Stage 8.
