NodusLab

MVP analysis

What Nodus v1 establishes today, which is Claim O and nothing else.

Status: complete for v1 as of commit 1f296df. Last revised 2026-09-01.

This is the "what do we already have" pass required before any new work. It is deliberately unflattering where the evidence warrants it and deliberately generous where v1 is already correct — v1's own documentation is unusually honest about its limits, and several conclusions below are simply confirmations of statements v1 already makes about itself.


1. What v1 is

A single-operator distributed inference network that closes one economic loop:

consumer  ->  coordinator (schedule, escrow)  ->  provider node (llama.cpp)
          <-  settle (ledger)                 <-  result + signed receipt

~6,900 LOC Python + 264 LOC dashboard; 130 tests (129 pass, 1 gated on real inference). MEASURED production run: 14 jobs, 2 machines, 1 model (qwen2.5-0.5b-instruct Q4_K_M), 100% verification pass, 38.165 CU settled, ledger conserved, coordinator overhead median 3 ms on top of inference.

That last figure is the most useful single number v1 produced for this project: the economic layer costs ~3 ms per job today. It is the budget line that any verification mechanism eats into. A verification scheme costing 300 ms per job is a 100× regression in coordinator overhead even if it is negligible next to a 500 ms inference.

2. Units

  • CU-v0: 1 CU = 100 tokens of a fixed reference workload. Node capacity is MEASURED (219.25 tok/s on an M5, ±1% over n=5). Job cost is (prompt_tokens × 0.1 + completion_tokens) × model_weight / 100.
  • CC: transferable credit, integer micro-credits, database ledger.
  • model_weight is DECLARED in models.yaml, not measured. v1 says so.

3. What v1's verification actually establishes

coordinator/verification/verifier.py::BasicVerifier runs 11 checks. Mapped to the taxonomy (verification-taxonomy.md):

v1 check Claim Level
signature valid under node token O L1
job_id / node_id match assignment O L1
input_hash matches prompt sent O (binds statement) L1
output_hash matches returned text O (binds statement) L1
model / model_version match request none — self-attested string L0
completion_tokens <= max_tokens none (a bound, not evidence) L0
output/token-count mutual consistency none L0
no prior result for this job O (replay) L1
optional replication, output-hash equality A, weakly L2

Conclusion: v1's VERIFIED status means Claim O and nothing more. v1's own docstrings say exactly this ("This is deliberately not a proof of computation"), so this is a confirmation, not a criticism. The value of writing it down in taxonomy terms is that it makes the gap precise: v1 has L1/Claim-O, and every experiment here is measured by what it adds beyond that point.

The replication path is the only Claim-A evidence v1 has, and v1 correctly records a mismatch as a signal, not proof of fraud because floating-point non-determinism across backends can produce legitimate divergence. This is the right call and it is also the reason naive replication will not scale to a heterogeneous network: the false-positive rate is unquantified. Quantifying it is a concrete, cheap, high-value experiment (see §6, E009-pre).

4. Finding: settlement trusts a recomputable quantity

MEASURED (code read): coordinator/jobs.py:419-421

prompt_tokens = int(att.get("prompt_tokens") or 0)
completion_tokens = int(att.get("completion_tokens") or 0)
cost = cu_for_tokens(prompt_tokens, completion_tokens, spec.cu_weight)

Settlement multiplies the provider's self-reported token counts by a declared weight, capped only by the escrow reserve. This is attack T11/T12 in threat-model.md, and v1's own limitations section names it ("a malicious provider can ... inflate self-reported prefill token counts up to the escrow cap").

What makes this interesting rather than merely a bug is that both quantities are recomputable by the coordinator at near-zero cost:

  • prompt_tokens is a function of the prompt the coordinator itself sent.
  • completion_tokens is a function of the returned text, which the coordinator already has and already hashes.

Both require one tokeniser pass over text the coordinator holds — O(length), microseconds. The coordinator already contains a heuristic estimate_prompt_tokens (jobs.py:66) but does not use it for settlement.

This is the thesis of the whole programme in miniature, at the smallest possible scale:

A quantity that the verifier can derive from data it already holds should never be an input the prover is trusted for.

It also demonstrates the §4 point of research-thesis.md: once metering is derived from the job specification and the committed I/O rather than from provider self-report, Claim C collapses into Claim A. The billing-fraud surface disappears and the only remaining question is whether the output is real. That is a strictly easier research problem, and it is reached by deleting code rather than by adding cryptography.

Recommendation to the v1 codebase (not yet applied, out of scope for this repo's first commit): derive both token counts coordinator-side with the model's real tokeniser; keep the attested values only as a consistency signal (a provider whose self-report diverges from the truth is flagged, not paid differently). Estimated effort: small. Estimated cost: one tokeniser pass. Estimated benefit: closes T11, T12 and most of T5.

5. Structural properties that help us

v1 got several things right that this research depends on:

  1. Verifier is an abstract base class with a stable JobFacts / VerificationResult interface. New mechanisms plug in without touching the job engine. This is the integration point for everything in experiments/.
  2. Input and output are already hashed and bound into the receipt. A commit-and-prove scheme needs exactly these commitments; v1 has them (SHA-256 over text — adequate as a binding commitment, though not hiding).
  3. Escrow-then-settle means there is already a natural place to hold funds pending delayed verification. Probabilistic audit schemes need exactly this (a settlement window in which a job can still be challenged).
  4. Audit re-runs are funded by the treasury, not the consumer. The economic plumbing for a network-funded audit budget exists.
  5. Integer micro-credit accounting, so penalties/slashing can be expressed exactly.

6. Gaps that define the research

Gap Consequence Where addressed
No Claim-A evidence in the default path fabricated output is paid E001–E009
No model binding model substitution undetectable E006, E012
No freshness binding cached/precomputed answers undetectable E009, E012
Metering trusts the prover billing fraud up to escrow cap §4, immediately fixable
Coordinator fully trusted it could forge everything out of scope until Stage 10+
Replication false-positive rate UNKNOWN replication cannot be enforced E009-pre — cheapest useful experiment available
No stake / penalty probabilistic schemes have no teeth E009, E012

7. Immediate cheap experiment suggested by this analysis

ANSWERED (E018). Greedy decoding, 10 prompts x 300 tokens, honest CPU vs honest GPU: 9/10 identical, 1 diverged at token 72. A ~10% false-positive rate means v1's REPLICATED_MISMATCH cannot be used for enforcement on float inference — it would slash one honest provider in ten. Under the NC-0.3 contract the rate is 0/10. This is the number this section asked for.

E009-pre — "How often do two honest heterogeneous nodes disagree?" Run the same prompt at temperature 0 across the available backends (Metal/llama.cpp on M5, CPU-only llama.cpp, simulated node) and measure the distribution of output divergence: exact-match rate, first-divergent-token index, and logit-level distance. Without this number, replication and audit cannot be turned into enforcement, because we cannot set a threshold. v1 observed exactly one mismatch instance (against a simulated node) — n=1.

This costs hours, not weeks, requires no cryptography, and gates the entire L2/L3 branch of the taxonomy. It has been added to the roadmap as a Stage-0 deliverable rather than waiting for Stage 8.