NodusLab
E009 · CP-008 (main), **CP-008-pre (complete)**

Probabilistic verification

Is the tolerance required to accept every honest backend still tight enough to catch a provider who cheats on precision?

Partly completeE009-pre COMPLETE, run 2026-09-01. Main CP-008 (audit rates,experiments/009-probabilistic-verification/ ↗

Checkpoint: CP-008 (main), CP-008-pre (complete) Status: E009-pre COMPLETE, run 2026-09-01. Main CP-008 (audit rates, stake, expected-loss analysis) not started.

E009-pre — how far do two honest backends diverge?

Pulled forward out of Stage 8 on day 1 because it gates everything else: a probabilistic check needs a threshold, and a threshold cannot be set without knowing how far honest providers disagree.

Question

Is the tolerance required to accept every honest backend still tight enough to catch a provider who cheats on precision?

Answer: no.

network composition honest worst-case error bf16 cheat error separation
CPU-class only 1.0 × 10^-6 2.84 × 10^-3 ≈ 2,800×
CPU + Metal GPU 7.7 × 10^-4 2.84 × 10^-3 ≈ 3.7×

A 2,800× separation is a threshold anyone can set. A 3.7× separation is not. Hypothesis H9b is falsified, and with it E003's claim that precision cheating would be detectable in a heterogeneous network.

The tolerance must inflate 191–381× to accept an honest GPU provider, and at that width the cheat passes at every size tested.

The finding underneath it

The divergence is caused by the device, not the library or the summation order. mlx-cpu-f32 and mlx-gpu-f32 — identical library, identical dtype, identical call, one line of difference — diverge by 765–3,069×. The GPU path's relative error is constant at 7.7 × 10^-4 across n, and it is bit-exact on small integers: the signature of operand rounding, not accumulation.

A backend advertising float32 may not be computing in float32. Threat T7 is not only an attack — it is the default behaviour of a mainstream accelerator path, performed by providers with no intent to cheat.

What survives

E003's cost result is untouched: R_verify = 0.023, 43.7× cheaper than recomputing. And the check still detects gross cheating — fabricated output, a substituted model, skipped layers all produce 10^-2–10^0 relative errors, still 10–1,000× above even the inflated tolerance. It keeps the property that matters economically: it catches the cheats that actually save the provider money.

Files

file contents
hypothesis.md H9a–H9c and predictions recorded before running
methodology.md six backends, three tolerances, and why the 4× multiplier is not the result
implementation/divergence.py the experiment
analysis.md the result, and what it does to E003 and to v1's replication path
next_steps.md what to measure next

Raw records: ../../benchmarks/results/009-pre-backend-divergence.jsonl

Research question (CP-008-pre)

How far do two honest backends diverge on the same computation, and is the tolerance required to accommodate that divergence still tight enough to catch a provider who cheats on precision?

Why this had to be run before anything else

Experiment 003 showed that Freivalds verifies a matmul 43.7× more cheaply than recomputing it — but only after setting a tolerance τ, because floating-point arithmetic makes the honest residual nonzero. E003 derived τ from one library on one device, and flagged that as the result's central weakness:

"Cross-backend honest divergence will be larger — possibly much larger — than the single-machine residual, and τ must be raised to cover it, which raises the corruption an attacker can hide by the same factor. We do not know that factor."

Everything in the L2/L3 branch of the taxonomy — Freivalds checking, replication, random audit, and v1's existing REPLICATED_MISMATCH signal — needs a threshold, and a threshold cannot be set without this number.

Hypotheses

H9a. Honest cross-backend divergence exceeds honest same-backend divergence by at least an order of magnitude.

H9b. Even so, a separation remains: the tolerance wide enough to accept all honest backends is still tight enough to reject a provider who computes in bfloat16 while claiming float32 (threat T7). This is the optimistic reading carried in E003's analysis, which claimed such cheating would be "detected, but not by much".

  • If H9b is false, E003's security conclusion is falsified for heterogeneous networks, and probabilistic float-domain checking cannot police the attacks Nodus actually cares about.

H9c. The dominant source of divergence is the device, not the library or the summation order.

Predictions (recorded before running)

quantity prediction
CPU-vs-CPU relative divergence (different summation order) ~10^-7, i.e. float32 noise
CPU-vs-GPU relative divergence 10^-6–10^-5
tolerance inflation for heterogeneity 10–100×
bf16 cheat still detected under the inflated tolerance yes (H9b)

What a negative result would mean

If the honest CPU–GPU gap is comparable to the gap caused by an actual precision cheat, then no threshold separates them. A heterogeneous network could not enforce numerical correctness probabilistically at all, and would have to either (a) mandate exact/quantised arithmetic so the clean Freivalds guarantee returns, or (b) specify numerical semantics per job class and treat "wrong backend" as a protocol violation rather than something to be detected statistically.

Implementation: implementation/divergence.py.

Backends

All compute the same C = A·B from the same float32 A, B:

backend what it varies honest?
cpu-blas-f32 numpy + Accelerate — this is the verifier's own arithmetic honest
cpu-blocked-f32 128-column chunked K accumulation — different summation order honest
mlx-cpu-f32 MLX, CPU stream — different library, same device honest
mlx-gpu-f32 MLX, GPU stream — same library and dtype, only the device differs honest
mlx-gpu-bf16 MLX, GPU stream, bfloat16 operands CHEAT (threat T7)
cpu-f64-ref float64 then cast down reference / ground truth

The mlx-cpu-f32 / mlx-gpu-f32 pair is the important one: it isolates device from framework, because everything else about the two calls is identical.

The check

The verifier always uses its own arithmetic (cpu-blas-f32) regardless of which backend produced C:

residual(C) = max | A·(B·R) − C·R |        R ~ N(0,1), n×k, k = 40

This is exactly the E003 check, with the prover's backend varied.

Tolerances compared

  • τ_homo = 4 × the verifier's own residual — what E003 used.
  • τ_cpu = 4 × worst residual over the CPU-class backends.
  • τ_hetero = 4 × worst residual over all honest backends, GPU included.

A backend is accepted iff its residual ≤ τ. The experiment then asks whether the cheating backend is accepted under each τ, and reports the margin cheat_residual / τ so the verdict can be re-derived under a different safety multiplier.

The 4× multiplier is arbitrary and the result is sensitive to it

Stated prominently because it matters: 4× is a chosen safety margin, not a derived one. The verdicts below therefore should not be read as "the cheat is missed" full stop, but through the separation ratio — how far apart honest divergence and cheating divergence actually are — which is multiplier-independent and is the figure the analysis leads with.

Metrics

  • Freivalds residual per backend, computed by the verifier.
  • Max absolute and mean relative error against the float64 reference.
  • Tolerance inflation τ_hetero / τ_homo.
  • Cheat margin against each τ.

Sizes n ∈ {1024, 2048, 4096}, k = 40, seed fixed. MLX warmed up before timing.

Limitations

  • One machine. The CPU and GPU here share a vendor, a memory system and a compiler. Genuinely different vendors (NVIDIA, AMD, other CPU BLAS) will diverge at least as much, so these figures understate the problem.
  • One framework for the GPU path (MLX 0.32.2). Not a claim about Metal in general, still less about GPUs in general.
  • float64 is treated as ground truth. It is not exact either.
  • Single run per configuration.

Status: complete. Result: H9b falsified. Date: 2026-09-01. Machine: Apple M5, 10 cores, 25.8 GB. NumPy 2.5.2 + Accelerate; MLX 0.32.2 (Metal). Raw records: ../../benchmarks/results/009-pre-backend-divergence.jsonl


1. The headline

On this machine, the divergence between two honest backends is within a factor of 3.7 of the divergence caused by an actual precision cheat.

That leaves essentially no room to set a threshold. The consequence is direct and unwelcome:

Experiment 003's security conclusion does not hold for a heterogeneous network. E003 stated that a provider computing in bfloat16 while claiming float32 would be "detected, but not by much". Measured against a real GPU backend, it is not detected at all under any tolerance wide enough to accept that GPU as honest.

2. What was measured

Mean relative error against a float64 reference, and the Freivalds residual as computed by the verifier's own arithmetic:

backend honest? rel. err vs f64 residual (n=4096)
cpu-blas-f32 (the verifier) honest 1.01 × 10^-6 0.074
cpu-blocked-f32 honest 2.25 × 10^-7 0.058
mlx-cpu-f32 honest 1.01 × 10^-6 0.074
mlx-gpu-f32 honest 7.70 × 10^-4 14.16
mlx-gpu-bf16 CHEAT (T7) 2.84 × 10^-3 55.21

Tolerances and verdicts, all three sizes:

n τ_homo (what E003 used) τ_cpu τ_hetero inflation cheat vs τ_cpu cheat vs τ_hetero
1024 0.0342 0.0371 13.03 381× 305× → detected 0.87× → missed
2048 0.127 0.127 30.2 238× 224× → detected 0.94× → missed
4096 0.297 0.297 56.7 191× 186× → detected 0.97× → missed

3. The separation ratio — the multiplier-independent statement

The "missed" verdicts above depend on the arbitrary 4× safety multiplier, so they are not the result. The result is the separation between honest divergence and cheating divergence:

network composition honest worst-case rel. err cheat rel. err separation
CPU-class only 1.0 × 10^-6 2.84 × 10^-3 ≈ 2,800×
CPU + this GPU 7.7 × 10^-4 2.84 × 10^-3 ≈ 3.7×

A 2,800× separation is a threshold anyone can set. A 3.7× separation is not. Any tolerance tight enough to catch the cheat will reject honest GPU providers, and any tolerance loose enough to accept them will pass the cheat. Note also that the cheat margin rises with n (0.87 → 0.94 → 0.97): the two distributions are converging, not separating, as problems get larger.

4. H9c confirmed — and the finding underneath it

The dominant source of divergence is the device, not the library and not the summation order:

  • cpu-blas-f32 and mlx-cpu-f32 agree to the last bit measured (both 1.01 × 10^-6). Different framework, same device → no divergence.
  • cpu-blocked-f32, a deliberately different summation order, is better than either (2.25 × 10^-7). Summation order is not the problem.
  • mlx-cpu-f32 vs mlx-gpu-f32 — identical library, identical dtype, identical call, one line of difference — diverge by 765–3,069×.

Two additional measurements characterise the GPU path:

  1. Its relative error is constant at 7.70 × 10^-4 across n = 256, 1024, 4096. Accumulation error would grow with n; a constant relative error is the signature of operand rounding, not accumulation.
  2. On small-integer operands whose products are exactly representable, the GPU path is bit-exact (0.0 error) — consistent with rounding inputs to a reduced-precision format and then accumulating in float32.

The magnitude (7.7 × 10^-4) sits close to fp16 machine epsilon (9.8 × 10^-4).

MEASURED: MLX 0.32.2's matmul on the Metal GPU, called with float32 inputs and returning float32, delivers roughly 11 bits of effective mantissa precision, ~1,500× worse than the same call on the CPU stream. INFERRED (not verified against MLX internals): the operands are rounded to a reduced-precision format before multiplication. mx.matmul exposes no precision option.

This is worth stating carefully because it generalises beyond MLX: a backend advertising float32 may not be computing in float32. For a compute network this means threat T7 (quantise more aggressively than specified) is not only an attack — it is the default behaviour of at least one mainstream accelerator path, performed by providers with no intent to cheat.

5. What this does to the rest of the programme

E003's cost result is untouched. R_verify = 0.023 and the 43.7× advantage are properties of the algorithm and stand. What is falsified is E003's claim about what the float-domain check can enforce in a heterogeneous setting.

Updated position, replacing the optimistic paragraph in E003 §3:

network Freivalds float check verdict
homogeneous, CPU-class separation ~2,800× usable — detects gross cheating and precision cheating
heterogeneous, CPU + GPU separation ~3.7× not usable for precision cheating; still detects gross cheating (10^-2–10^0 errors)
any, exact/quantised integer arithmetic soundness ≤ 2^-k, no tolerance usable, with a real guarantee

The gross-cheating case survives and matters: fabricated output, a substituted model, or skipped layers produce relative errors of 10^-2 to 10^0, which is still 10–1,000× above even τ_hetero. The mechanism keeps its most economically important property — it catches the cheats that actually save the provider money — and loses the fine-grained one.

5a. Update (E014): the problem is dissolved, not traded against

E014 re-ran this comparison on integer operands instead of real-valued ones:

operands worst honest backend the cheat window
real-valued (this experiment) 7.70 × 10^-4 2.84 × 10^-3 3.7×
integer (E014) 0.0 — exact, every backend 1.41 × 10^-3 unbounded

The GPU's imprecision was never about the GPU. It was about representing real numbers. Given integers in range, CPU and GPU return the bit-exact product and the 3.7× window becomes unbounded. The heterogeneity problem this experiment identified does not need a cleverer threshold — it needs different arithmetic. See ../014-numeric-contract/ and docs/numeric-contract.md.

6. Consequences for Nodus's design

  1. Numerical semantics must be part of the job specification. Today v1 specifies a model and a model_version. That is not enough to define what a correct answer is: the same model on the same input on CPU and GPU gives answers 1,500× further apart than float32 rounding would suggest. Without a declared numeric contract, "correct" is undefined and no verifier can be right.
  2. This reframes v1's REPLICATED_MISMATCH as unfixable-as-designed. v1 compares output hashes for exact equality at temperature 0 and correctly treats a mismatch as a signal rather than proof. This experiment explains why, and shows the problem is worse than assumed: with a 7.7 × 10^-4 per-matmul divergence feeding an autoregressive decoder, honest CPU and GPU providers will diverge in sampled tokens, routinely. Exact-match replication cannot work across devices. Prediction to test next: honest CPU/GPU greedy decoding of the same prompt diverges within tens of tokens.
  3. The strongest argument yet for quantised integer inference. Exact integer arithmetic restores the clean 2^-40 Freivalds guarantee and removes the tolerance entirely. That is a network-wide architectural choice driven purely by a verification requirement — exactly the kind of consequence this programme exists to surface.
  4. Providers should be classed by numeric behaviour, not just by hardware. A cheap partition — "CPU-class exact-ish" vs "GPU-class reduced-precision" — lets the coordinator apply a tight tolerance within a class and use replication within a class rather than across classes.

7. What a malicious provider can still get away with

  • Computing in reduced precision on a GPU and being indistinguishable from an honest GPU provider. Under a resource-priced economy this is profitable: bfloat16 matmul is materially cheaper than float32.
  • Everything E003 already listed: caching, no Claim B/C/D.
  • Choosing to be a "GPU-class" provider precisely because that class must be granted the loose tolerance.

That last one is a genuine incentive problem: the tolerance a network must grant its least precise honest hardware becomes the budget available to every attacker.

8. Limitations

  • One machine; CPU and GPU share vendor and memory system. Cross-vendor divergence should be larger, making the separation worse, not better.
  • One GPU framework. Not a claim about Metal generally or GPUs generally.
  • The reduced-precision inference in §4 is inferred from external behaviour, not confirmed against MLX's kernels.
  • Single run per configuration; matmul only, not end-to-end inference. §6.2 is a prediction, not a measurement.

1. The same measurement on real inference [highest value]

This experiment measured one matmul. Nodus runs autoregressive decoding, where a 7.7 × 10^-4 per-matmul divergence compounds through layers and then through a sampling step that is a discontinuous function of the logits. The prediction in analysis.md §6.2 — that honest CPU and GPU greedy decoding of the same prompt diverge within tens of tokens — is currently untested.

Measure, using v1's own stack (../../../nodus, llama.cpp, qwen2.5-0.5b):

  • exact output-match rate between Metal and CPU backends at temperature 0;
  • the index of the first divergent token, as a distribution;
  • whether divergence correlates with prompt length or sampling entropy.

This gives v1 the number its REPLICATED_MISMATCH path needs and has never had (v1 observed exactly one mismatch, against a simulated node, n=1). It also decides whether cross-device replication is viable at all.

2. Cross-vendor divergence

CPU and GPU here share a vendor, a memory system and a compiler, so these figures understate the problem. Repeat on NVIDIA (CUDA/cuBLAS, and TF32 on/off) and on a different CPU BLAS (OpenBLAS, MKL). TF32 is the direct analogue of what we found: NVIDIA's default for float32 matmul on Ampere and later is a 19-bit-mantissa format, so the same "float32 that is not float32" issue should appear there. Confirming it would show this is a property of accelerators, not of MLX.

3. Confirm the operand-rounding inference

§4 infers reduced-precision operand rounding from external behaviour. Confirm it directly: sweep operand magnitudes to locate the mantissa cutoff, and check MLX's Metal kernels. Worth doing because "a float32 API that is not float32" is a claim we would want to be exactly right about before repeating it.

4. Quantised-integer Freivalds [the fix]

Now clearly the most promising direction. Exact integer arithmetic removes the tolerance and restores the 2^-40 guarantee, which this experiment shows is the only way to police precision cheating in a heterogeneous network. Measure the cost of the check in modular arithmetic, and the accuracy cost of quantised inference, and decide whether the trade is worth it network-wide.

5. Numeric-class provider partitioning

A cheap engineering response, testable now: classify providers by measured numeric behaviour (run a fixed reference matmul at registration, cluster the residuals), then apply a tight tolerance within a class and replicate only within a class. Measure how cleanly the classes separate and whether a provider can lie about its class — it can, so the class must be measured by the coordinator, not declared by the provider.

6. Then, the actual CP-008

Audit rates, detection probability, and the stake that makes cheating negative expected value. That work was always downstream of this measurement; it now has the input it needed, plus a sharper question: the tolerance a network must grant its least precise honest hardware is the budget available to every attacker, so the economic analysis has to price that budget.