MVP analysis
What Nodus v1 establishes today, which is Claim O and nothing else.
Status: complete for v1 as of commit 1f296df. Last revised 2026-09-01.
This is the "what do we already have" pass required before any new work. It is deliberately unflattering where the evidence warrants it and deliberately generous where v1 is already correct — v1's own documentation is unusually honest about its limits, and several conclusions below are simply confirmations of statements v1 already makes about itself.
1. What v1 is
A single-operator distributed inference network that closes one economic loop:
consumer -> coordinator (schedule, escrow) -> provider node (llama.cpp)
<- settle (ledger) <- result + signed receipt
~6,900 LOC Python + 264 LOC dashboard; 130 tests (129 pass, 1 gated on real inference). MEASURED production run: 14 jobs, 2 machines, 1 model (qwen2.5-0.5b-instruct Q4_K_M), 100% verification pass, 38.165 CU settled, ledger conserved, coordinator overhead median 3 ms on top of inference.
That last figure is the most useful single number v1 produced for this project: the economic layer costs ~3 ms per job today. It is the budget line that any verification mechanism eats into. A verification scheme costing 300 ms per job is a 100× regression in coordinator overhead even if it is negligible next to a 500 ms inference.
2. Units
- CU-v0: 1 CU = 100 tokens of a fixed reference workload. Node capacity is
MEASURED (219.25 tok/s on an M5, ±1% over n=5). Job cost is
(prompt_tokens × 0.1 + completion_tokens) × model_weight / 100. - CC: transferable credit, integer micro-credits, database ledger.
model_weightis DECLARED inmodels.yaml, not measured. v1 says so.
3. What v1's verification actually establishes
coordinator/verification/verifier.py::BasicVerifier runs 11 checks. Mapped to
the taxonomy (verification-taxonomy.md):
| v1 check | Claim | Level |
|---|---|---|
| signature valid under node token | O | L1 |
| job_id / node_id match assignment | O | L1 |
| input_hash matches prompt sent | O (binds statement) | L1 |
| output_hash matches returned text | O (binds statement) | L1 |
| model / model_version match request | none — self-attested string | L0 |
completion_tokens <= max_tokens |
none (a bound, not evidence) | L0 |
| output/token-count mutual consistency | none | L0 |
| no prior result for this job | O (replay) | L1 |
| optional replication, output-hash equality | A, weakly | L2 |
Conclusion: v1's VERIFIED status means Claim O and nothing more. v1's own
docstrings say exactly this ("This is deliberately not a proof of
computation"), so this is a confirmation, not a criticism. The value of writing
it down in taxonomy terms is that it makes the gap precise: v1 has L1/Claim-O,
and every experiment here is measured by what it adds beyond that point.
The replication path is the only Claim-A evidence v1 has, and v1 correctly records a mismatch as a signal, not proof of fraud because floating-point non-determinism across backends can produce legitimate divergence. This is the right call and it is also the reason naive replication will not scale to a heterogeneous network: the false-positive rate is unquantified. Quantifying it is a concrete, cheap, high-value experiment (see §6, E009-pre).
4. Finding: settlement trusts a recomputable quantity
MEASURED (code read): coordinator/jobs.py:419-421
prompt_tokens = int(att.get("prompt_tokens") or 0)
completion_tokens = int(att.get("completion_tokens") or 0)
cost = cu_for_tokens(prompt_tokens, completion_tokens, spec.cu_weight)
Settlement multiplies the provider's self-reported token counts by a
declared weight, capped only by the escrow reserve. This is attack T11/T12 in
threat-model.md, and v1's own limitations section names it ("a malicious
provider can ... inflate self-reported prefill token counts up to the escrow
cap").
What makes this interesting rather than merely a bug is that both quantities are recomputable by the coordinator at near-zero cost:
prompt_tokensis a function of the prompt the coordinator itself sent.completion_tokensis a function of the returned text, which the coordinator already has and already hashes.
Both require one tokeniser pass over text the coordinator holds — O(length),
microseconds. The coordinator already contains a heuristic
estimate_prompt_tokens (jobs.py:66) but does not use it for settlement.
This is the thesis of the whole programme in miniature, at the smallest possible scale:
A quantity that the verifier can derive from data it already holds should never be an input the prover is trusted for.
It also demonstrates the §4 point of research-thesis.md: once metering is
derived from the job specification and the committed I/O rather than from
provider self-report, Claim C collapses into Claim A. The billing-fraud
surface disappears and the only remaining question is whether the output is
real. That is a strictly easier research problem, and it is reached by deleting
code rather than by adding cryptography.
Recommendation to the v1 codebase (not yet applied, out of scope for this repo's first commit): derive both token counts coordinator-side with the model's real tokeniser; keep the attested values only as a consistency signal (a provider whose self-report diverges from the truth is flagged, not paid differently). Estimated effort: small. Estimated cost: one tokeniser pass. Estimated benefit: closes T11, T12 and most of T5.
5. Structural properties that help us
v1 got several things right that this research depends on:
Verifieris an abstract base class with a stableJobFacts/VerificationResultinterface. New mechanisms plug in without touching the job engine. This is the integration point for everything inexperiments/.- Input and output are already hashed and bound into the receipt. A commit-and-prove scheme needs exactly these commitments; v1 has them (SHA-256 over text — adequate as a binding commitment, though not hiding).
- Escrow-then-settle means there is already a natural place to hold funds pending delayed verification. Probabilistic audit schemes need exactly this (a settlement window in which a job can still be challenged).
- Audit re-runs are funded by the treasury, not the consumer. The economic plumbing for a network-funded audit budget exists.
- Integer micro-credit accounting, so penalties/slashing can be expressed exactly.
6. Gaps that define the research
| Gap | Consequence | Where addressed |
|---|---|---|
| No Claim-A evidence in the default path | fabricated output is paid | E001–E009 |
| No model binding | model substitution undetectable | E006, E012 |
| No freshness binding | cached/precomputed answers undetectable | E009, E012 |
| Metering trusts the prover | billing fraud up to escrow cap | §4, immediately fixable |
| Coordinator fully trusted | it could forge everything | out of scope until Stage 10+ |
| Replication false-positive rate UNKNOWN | replication cannot be enforced | E009-pre — cheapest useful experiment available |
| No stake / penalty | probabilistic schemes have no teeth | E009, E012 |
7. Immediate cheap experiment suggested by this analysis
ANSWERED (E018). Greedy decoding, 10 prompts x 300 tokens, honest CPU vs honest GPU: 9/10 identical, 1 diverged at token 72. A ~10% false-positive rate means v1's
REPLICATED_MISMATCHcannot be used for enforcement on float inference — it would slash one honest provider in ten. Under the NC-0.3 contract the rate is 0/10. This is the number this section asked for.
E009-pre — "How often do two honest heterogeneous nodes disagree?" Run the same prompt at temperature 0 across the available backends (Metal/llama.cpp on M5, CPU-only llama.cpp, simulated node) and measure the distribution of output divergence: exact-match rate, first-divergent-token index, and logit-level distance. Without this number, replication and audit cannot be turned into enforcement, because we cannot set a threshold. v1 observed exactly one mismatch instance (against a simulated node) — n=1.
This costs hours, not weeks, requires no cryptography, and gates the entire L2/L3 branch of the taxonomy. It has been added to the roadmap as a Stage-0 deliverable rather than waiting for Stage 8.