Probabilistic verification
Is the tolerance required to accept every honest backend still tight enough to catch a provider who cheats on precision?
Overview
README.md ↗Checkpoint: CP-008 (main), CP-008-pre (complete) Status: E009-pre COMPLETE, run 2026-09-01. Main CP-008 (audit rates, stake, expected-loss analysis) not started.
E009-pre — how far do two honest backends diverge?
Pulled forward out of Stage 8 on day 1 because it gates everything else: a probabilistic check needs a threshold, and a threshold cannot be set without knowing how far honest providers disagree.
Question
Is the tolerance required to accept every honest backend still tight enough to catch a provider who cheats on precision?
Answer: no.
| network composition | honest worst-case error | bf16 cheat error | separation |
|---|---|---|---|
| CPU-class only | 1.0 × 10^-6 | 2.84 × 10^-3 | ≈ 2,800× |
| CPU + Metal GPU | 7.7 × 10^-4 | 2.84 × 10^-3 | ≈ 3.7× |
A 2,800× separation is a threshold anyone can set. A 3.7× separation is not. Hypothesis H9b is falsified, and with it E003's claim that precision cheating would be detectable in a heterogeneous network.
The tolerance must inflate 191–381× to accept an honest GPU provider, and at that width the cheat passes at every size tested.
The finding underneath it
The divergence is caused by the device, not the library or the summation
order. mlx-cpu-f32 and mlx-gpu-f32 — identical library, identical dtype,
identical call, one line of difference — diverge by 765–3,069×. The GPU path's
relative error is constant at 7.7 × 10^-4 across n, and it is bit-exact on
small integers: the signature of operand rounding, not accumulation.
A backend advertising float32 may not be computing in float32. Threat T7 is not only an attack — it is the default behaviour of a mainstream accelerator path, performed by providers with no intent to cheat.
What survives
E003's cost result is untouched: R_verify = 0.023, 43.7× cheaper than
recomputing. And the check still detects gross cheating — fabricated output, a
substituted model, skipped layers all produce 10^-2–10^0 relative errors, still
10–1,000× above even the inflated tolerance. It keeps the property that matters
economically: it catches the cheats that actually save the provider money.
Files
| file | contents |
|---|---|
hypothesis.md |
H9a–H9c and predictions recorded before running |
methodology.md |
six backends, three tolerances, and why the 4× multiplier is not the result |
implementation/divergence.py |
the experiment |
analysis.md |
the result, and what it does to E003 and to v1's replication path |
next_steps.md |
what to measure next |
Raw records: ../../benchmarks/results/009-pre-backend-divergence.jsonl
Hypothesis
hypothesis.md ↗Research question (CP-008-pre)
How far do two honest backends diverge on the same computation, and is the tolerance required to accommodate that divergence still tight enough to catch a provider who cheats on precision?
Why this had to be run before anything else
Experiment 003 showed that Freivalds verifies a matmul 43.7× more cheaply than
recomputing it — but only after setting a tolerance τ, because floating-point
arithmetic makes the honest residual nonzero. E003 derived τ from one
library on one device, and flagged that as the result's central weakness:
"Cross-backend honest divergence will be larger — possibly much larger — than the single-machine residual, and
τmust be raised to cover it, which raises the corruption an attacker can hide by the same factor. We do not know that factor."
Everything in the L2/L3 branch of the taxonomy — Freivalds checking, replication,
random audit, and v1's existing REPLICATED_MISMATCH signal — needs a threshold,
and a threshold cannot be set without this number.
Hypotheses
H9a. Honest cross-backend divergence exceeds honest same-backend divergence by at least an order of magnitude.
H9b. Even so, a separation remains: the tolerance wide enough to accept all honest backends is still tight enough to reject a provider who computes in bfloat16 while claiming float32 (threat T7). This is the optimistic reading carried in E003's analysis, which claimed such cheating would be "detected, but not by much".
- If H9b is false, E003's security conclusion is falsified for heterogeneous networks, and probabilistic float-domain checking cannot police the attacks Nodus actually cares about.
H9c. The dominant source of divergence is the device, not the library or the summation order.
Predictions (recorded before running)
| quantity | prediction |
|---|---|
| CPU-vs-CPU relative divergence (different summation order) | ~10^-7, i.e. float32 noise |
| CPU-vs-GPU relative divergence | 10^-6–10^-5 |
| tolerance inflation for heterogeneity | 10–100× |
| bf16 cheat still detected under the inflated tolerance | yes (H9b) |
What a negative result would mean
If the honest CPU–GPU gap is comparable to the gap caused by an actual precision cheat, then no threshold separates them. A heterogeneous network could not enforce numerical correctness probabilistically at all, and would have to either (a) mandate exact/quantised arithmetic so the clean Freivalds guarantee returns, or (b) specify numerical semantics per job class and treat "wrong backend" as a protocol violation rather than something to be detected statistically.
Methodology
methodology.md ↗Implementation: implementation/divergence.py.
Backends
All compute the same C = A·B from the same float32 A, B:
| backend | what it varies | honest? |
|---|---|---|
cpu-blas-f32 |
numpy + Accelerate — this is the verifier's own arithmetic | honest |
cpu-blocked-f32 |
128-column chunked K accumulation — different summation order | honest |
mlx-cpu-f32 |
MLX, CPU stream — different library, same device | honest |
mlx-gpu-f32 |
MLX, GPU stream — same library and dtype, only the device differs | honest |
mlx-gpu-bf16 |
MLX, GPU stream, bfloat16 operands | CHEAT (threat T7) |
cpu-f64-ref |
float64 then cast down | reference / ground truth |
The mlx-cpu-f32 / mlx-gpu-f32 pair is the important one: it isolates device
from framework, because everything else about the two calls is identical.
The check
The verifier always uses its own arithmetic (cpu-blas-f32) regardless of which
backend produced C:
residual(C) = max | A·(B·R) − C·R | R ~ N(0,1), n×k, k = 40
This is exactly the E003 check, with the prover's backend varied.
Tolerances compared
τ_homo= 4 × the verifier's own residual — what E003 used.τ_cpu= 4 × worst residual over the CPU-class backends.τ_hetero= 4 × worst residual over all honest backends, GPU included.
A backend is accepted iff its residual ≤ τ. The experiment then asks whether the
cheating backend is accepted under each τ, and reports the margin
cheat_residual / τ so the verdict can be re-derived under a different safety
multiplier.
The 4× multiplier is arbitrary and the result is sensitive to it
Stated prominently because it matters: 4× is a chosen safety margin, not a
derived one. The verdicts below therefore should not be read as "the cheat is
missed" full stop, but through the separation ratio — how far apart honest
divergence and cheating divergence actually are — which is multiplier-independent
and is the figure the analysis leads with.
Metrics
- Freivalds residual per backend, computed by the verifier.
- Max absolute and mean relative error against the float64 reference.
- Tolerance inflation
τ_hetero / τ_homo. - Cheat margin against each τ.
Sizes n ∈ {1024, 2048, 4096}, k = 40, seed fixed. MLX warmed up before timing.
Limitations
- One machine. The CPU and GPU here share a vendor, a memory system and a compiler. Genuinely different vendors (NVIDIA, AMD, other CPU BLAS) will diverge at least as much, so these figures understate the problem.
- One framework for the GPU path (MLX 0.32.2). Not a claim about Metal in general, still less about GPUs in general.
- float64 is treated as ground truth. It is not exact either.
- Single run per configuration.
Analysis
analysis.md ↗Status: complete. Result: H9b falsified.
Date: 2026-09-01. Machine: Apple M5, 10 cores, 25.8 GB. NumPy 2.5.2 +
Accelerate; MLX 0.32.2 (Metal).
Raw records: ../../benchmarks/results/009-pre-backend-divergence.jsonl
1. The headline
On this machine, the divergence between two honest backends is within a factor of 3.7 of the divergence caused by an actual precision cheat.
That leaves essentially no room to set a threshold. The consequence is direct and unwelcome:
Experiment 003's security conclusion does not hold for a heterogeneous network. E003 stated that a provider computing in bfloat16 while claiming float32 would be "detected, but not by much". Measured against a real GPU backend, it is not detected at all under any tolerance wide enough to accept that GPU as honest.
2. What was measured
Mean relative error against a float64 reference, and the Freivalds residual as computed by the verifier's own arithmetic:
| backend | honest? | rel. err vs f64 | residual (n=4096) |
|---|---|---|---|
cpu-blas-f32 (the verifier) |
honest | 1.01 × 10^-6 | 0.074 |
cpu-blocked-f32 |
honest | 2.25 × 10^-7 | 0.058 |
mlx-cpu-f32 |
honest | 1.01 × 10^-6 | 0.074 |
mlx-gpu-f32 |
honest | 7.70 × 10^-4 | 14.16 |
mlx-gpu-bf16 |
CHEAT (T7) | 2.84 × 10^-3 | 55.21 |
Tolerances and verdicts, all three sizes:
| n | τ_homo (what E003 used) |
τ_cpu |
τ_hetero |
inflation | cheat vs τ_cpu |
cheat vs τ_hetero |
|---|---|---|---|---|---|---|
| 1024 | 0.0342 | 0.0371 | 13.03 | 381× | 305× → detected | 0.87× → missed |
| 2048 | 0.127 | 0.127 | 30.2 | 238× | 224× → detected | 0.94× → missed |
| 4096 | 0.297 | 0.297 | 56.7 | 191× | 186× → detected | 0.97× → missed |
3. The separation ratio — the multiplier-independent statement
The "missed" verdicts above depend on the arbitrary 4× safety multiplier, so they are not the result. The result is the separation between honest divergence and cheating divergence:
| network composition | honest worst-case rel. err | cheat rel. err | separation |
|---|---|---|---|
| CPU-class only | 1.0 × 10^-6 | 2.84 × 10^-3 | ≈ 2,800× |
| CPU + this GPU | 7.7 × 10^-4 | 2.84 × 10^-3 | ≈ 3.7× |
A 2,800× separation is a threshold anyone can set. A 3.7× separation is not. Any tolerance tight enough to catch the cheat will reject honest GPU providers, and any tolerance loose enough to accept them will pass the cheat. Note also that the cheat margin rises with n (0.87 → 0.94 → 0.97): the two distributions are converging, not separating, as problems get larger.
4. H9c confirmed — and the finding underneath it
The dominant source of divergence is the device, not the library and not the summation order:
cpu-blas-f32andmlx-cpu-f32agree to the last bit measured (both 1.01 × 10^-6). Different framework, same device → no divergence.cpu-blocked-f32, a deliberately different summation order, is better than either (2.25 × 10^-7). Summation order is not the problem.mlx-cpu-f32vsmlx-gpu-f32— identical library, identical dtype, identical call, one line of difference — diverge by 765–3,069×.
Two additional measurements characterise the GPU path:
- Its relative error is constant at 7.70 × 10^-4 across n = 256, 1024,
4096. Accumulation error would grow with
n; a constant relative error is the signature of operand rounding, not accumulation. - On small-integer operands whose products are exactly representable, the GPU path is bit-exact (0.0 error) — consistent with rounding inputs to a reduced-precision format and then accumulating in float32.
The magnitude (7.7 × 10^-4) sits close to fp16 machine epsilon (9.8 × 10^-4).
MEASURED: MLX 0.32.2's
matmulon the Metal GPU, called with float32 inputs and returning float32, delivers roughly 11 bits of effective mantissa precision, ~1,500× worse than the same call on the CPU stream. INFERRED (not verified against MLX internals): the operands are rounded to a reduced-precision format before multiplication.mx.matmulexposes no precision option.
This is worth stating carefully because it generalises beyond MLX: a backend advertising float32 may not be computing in float32. For a compute network this means threat T7 (quantise more aggressively than specified) is not only an attack — it is the default behaviour of at least one mainstream accelerator path, performed by providers with no intent to cheat.
5. What this does to the rest of the programme
E003's cost result is untouched. R_verify = 0.023 and the 43.7× advantage
are properties of the algorithm and stand. What is falsified is E003's claim
about what the float-domain check can enforce in a heterogeneous setting.
Updated position, replacing the optimistic paragraph in E003 §3:
| network | Freivalds float check | verdict |
|---|---|---|
| homogeneous, CPU-class | separation ~2,800× | usable — detects gross cheating and precision cheating |
| heterogeneous, CPU + GPU | separation ~3.7× | not usable for precision cheating; still detects gross cheating (10^-2–10^0 errors) |
| any, exact/quantised integer arithmetic | soundness ≤ 2^-k, no tolerance | usable, with a real guarantee |
The gross-cheating case survives and matters: fabricated output, a substituted
model, or skipped layers produce relative errors of 10^-2 to 10^0, which is
still 10–1,000× above even τ_hetero. The mechanism keeps its most
economically important property — it catches the cheats that actually save the
provider money — and loses the fine-grained one.
5a. Update (E014): the problem is dissolved, not traded against
E014 re-ran this comparison on integer operands instead of real-valued ones:
| operands | worst honest backend | the cheat | window |
|---|---|---|---|
| real-valued (this experiment) | 7.70 × 10^-4 | 2.84 × 10^-3 | 3.7× |
| integer (E014) | 0.0 — exact, every backend | 1.41 × 10^-3 | unbounded |
The GPU's imprecision was never about the GPU. It was about representing real
numbers. Given integers in range, CPU and GPU return the bit-exact product and
the 3.7× window becomes unbounded. The heterogeneity problem this experiment
identified does not need a cleverer threshold — it needs different arithmetic.
See ../014-numeric-contract/ and docs/numeric-contract.md.
6. Consequences for Nodus's design
- Numerical semantics must be part of the job specification. Today v1
specifies a model and a
model_version. That is not enough to define what a correct answer is: the same model on the same input on CPU and GPU gives answers 1,500× further apart than float32 rounding would suggest. Without a declared numeric contract, "correct" is undefined and no verifier can be right. - This reframes v1's
REPLICATED_MISMATCHas unfixable-as-designed. v1 compares output hashes for exact equality at temperature 0 and correctly treats a mismatch as a signal rather than proof. This experiment explains why, and shows the problem is worse than assumed: with a 7.7 × 10^-4 per-matmul divergence feeding an autoregressive decoder, honest CPU and GPU providers will diverge in sampled tokens, routinely. Exact-match replication cannot work across devices. Prediction to test next: honest CPU/GPU greedy decoding of the same prompt diverges within tens of tokens. - The strongest argument yet for quantised integer inference. Exact integer arithmetic restores the clean 2^-40 Freivalds guarantee and removes the tolerance entirely. That is a network-wide architectural choice driven purely by a verification requirement — exactly the kind of consequence this programme exists to surface.
- Providers should be classed by numeric behaviour, not just by hardware. A cheap partition — "CPU-class exact-ish" vs "GPU-class reduced-precision" — lets the coordinator apply a tight tolerance within a class and use replication within a class rather than across classes.
7. What a malicious provider can still get away with
- Computing in reduced precision on a GPU and being indistinguishable from an honest GPU provider. Under a resource-priced economy this is profitable: bfloat16 matmul is materially cheaper than float32.
- Everything E003 already listed: caching, no Claim B/C/D.
- Choosing to be a "GPU-class" provider precisely because that class must be granted the loose tolerance.
That last one is a genuine incentive problem: the tolerance a network must grant its least precise honest hardware becomes the budget available to every attacker.
8. Limitations
- One machine; CPU and GPU share vendor and memory system. Cross-vendor divergence should be larger, making the separation worse, not better.
- One GPU framework. Not a claim about Metal generally or GPUs generally.
- The reduced-precision inference in §4 is inferred from external behaviour, not confirmed against MLX's kernels.
- Single run per configuration; matmul only, not end-to-end inference. §6.2 is a prediction, not a measurement.
Next steps
next_steps.md ↗1. The same measurement on real inference [highest value]
This experiment measured one matmul. Nodus runs autoregressive decoding, where a
7.7 × 10^-4 per-matmul divergence compounds through layers and then through a
sampling step that is a discontinuous function of the logits. The prediction
in analysis.md §6.2 — that honest CPU and GPU greedy decoding of the same
prompt diverge within tens of tokens — is currently untested.
Measure, using v1's own stack (../../../nodus, llama.cpp, qwen2.5-0.5b):
- exact output-match rate between Metal and CPU backends at temperature 0;
- the index of the first divergent token, as a distribution;
- whether divergence correlates with prompt length or sampling entropy.
This gives v1 the number its REPLICATED_MISMATCH path needs and has never had
(v1 observed exactly one mismatch, against a simulated node, n=1). It also
decides whether cross-device replication is viable at all.
2. Cross-vendor divergence
CPU and GPU here share a vendor, a memory system and a compiler, so these figures understate the problem. Repeat on NVIDIA (CUDA/cuBLAS, and TF32 on/off) and on a different CPU BLAS (OpenBLAS, MKL). TF32 is the direct analogue of what we found: NVIDIA's default for float32 matmul on Ampere and later is a 19-bit-mantissa format, so the same "float32 that is not float32" issue should appear there. Confirming it would show this is a property of accelerators, not of MLX.
3. Confirm the operand-rounding inference
§4 infers reduced-precision operand rounding from external behaviour. Confirm it directly: sweep operand magnitudes to locate the mantissa cutoff, and check MLX's Metal kernels. Worth doing because "a float32 API that is not float32" is a claim we would want to be exactly right about before repeating it.
4. Quantised-integer Freivalds [the fix]
Now clearly the most promising direction. Exact integer arithmetic removes the tolerance and restores the 2^-40 guarantee, which this experiment shows is the only way to police precision cheating in a heterogeneous network. Measure the cost of the check in modular arithmetic, and the accuracy cost of quantised inference, and decide whether the trade is worth it network-wide.
5. Numeric-class provider partitioning
A cheap engineering response, testable now: classify providers by measured numeric behaviour (run a fixed reference matmul at registration, cluster the residuals), then apply a tight tolerance within a class and replicate only within a class. Measure how cleanly the classes separate and whether a provider can lie about its class — it can, so the class must be measured by the coordinator, not declared by the provider.
6. Then, the actual CP-008
Audit rates, detection probability, and the stake that makes cheating negative expected value. That work was always downstream of this measurement; it now has the input it needed, plus a sharper question: the tolerance a network must grant its least precise honest hardware is the budget available to every attacker, so the economic analysis has to price that budget.