The numeric contract
what does a numeric contract cost in supply, and what does it buy in security?
Overview
README.md ↗Checkpoint: CP-012 · Status: COMPLETE, run 2026-09-01 Question: what does a numeric contract cost in supply, and what does it buy in security?
Result — the expected trade-off does not exist
| operands | worst honest backend | the cheat | usable window |
|---|---|---|---|
| real-valued | 7.70 × 10^-4 (GPU) | 2.84 × 10^-3 | 3.7× |
| integer | 0.0 — exact, every backend | 1.41 × 10^-3 | unbounded |
On integer operands, CPU BLAS, MLX-CPU and MLX-GPU all return the bit-exact product.
The GPU's imprecision was never about the GPU — it was about representing real numbers. Given integers in range, the same hardware that diverged by 1,500× in E009-pre agrees perfectly.
So the strictest contract is not the most exclusive one. H14c was predicted and is falsified: the exact contract costs no supply (5/5 honest backends admitted on every realistic operand regime) while separating cheats without bound. E009-pre's heterogeneity problem does not need a cleverer threshold; it needs different arithmetic.
Contract conformance
| contract | supply | security |
|---|---|---|
C0-none (Nodus v1 today) |
5/5 | 0/2 |
C1-float-1e-2 |
5/5 | 1/2 |
C2-float-1e-3 |
5/5 | 2/2 — but on a 3.7× knife-edge |
C4-exact-unbounded |
5/5 (3/5 on spike) |
2/2 |
C5-exact-tiled |
5/5 by construction | 2/2 |
C0 is the row that matters: v1 imposes no numeric requirement and admits both
cheats at every size and regime.
The finding random testing would have hidden
On magnitude-stressed input, exact conformance drops to 3/5 — and the GPU
passes while CPU BLAS fails. Worst-case |C| is M²·n but random data only
reaches M²·√n, a 64× gap at n = 4096, and float32 is exact on integers only to
2^24.
A numeric contract validated on random inputs is not validated.
Hence the tiling clause, with a derived number: for int8 into float32, accumulation length ≤ 1040. With it, every backend is exact by construction on every input.
Deliverable
docs/numeric-contract.md — draft NC-0.1,
eight clauses, a worked example, and an explicit statement of what it does not
cover (non-linearities, the accuracy cost of quantisation, cross-vendor
validation).
Files
| file | contents |
|---|---|
hypothesis.md |
H14a–c, including the prediction that was backwards |
methodology.md |
six contracts, seven backends, four operand regimes |
implementation/contracts.py |
the experiment |
analysis.md |
the result, and why exact contracts also win on checking cost |
next_steps.md |
non-linearities are where this breaks next |
Raw records: ../../benchmarks/results/014-numeric-contract.jsonl
A method error worth reading
The first version of this experiment used an all-127 matrix as its "worst case". That matrix is rank-1 and exactly representable in bfloat16, so both cheating backends returned the exactly correct product and every tolerance contract appeared to admit every cheat. The conclusion was wrong and was a property of the test data. Operand structure and magnitude are varied separately for this reason.
Hypothesis
hypothesis.md ↗Checkpoint CP-012
What does a numeric contract cost in supply, and what does it buy in security?
Why this is the next experiment
E009-pre and E013 reached the same residual risk from opposite directions:
A provider can be perfectly correct with respect to a cheaper specification than the one intended.
A job that names a model but not its arithmetic does not determine a correct answer, so no verifier can be right — and no combination of mechanisms reaches the problem, because each would faithfully verify the wrong thing. A ZK proof of the cheap circuit is a valid proof. An attestation of the cheap binary is a valid attestation. The specification sits upstream of every mechanism.
Stating the contract costs nothing. The question is what it costs to enforce.
The expected trade-off
A numeric contract is a clause in the job specification saying what arithmetic the provider must use, evaluated as a predicate on the output. The assumed shape of the problem:
strict contract -> excludes cheats -> but also excludes honest hardware
-> smaller provider pool
-> less supply, higher prices
loose contract -> admits everyone -> including the cheats
So we expect a curve, and the deliverable is where on it to sit.
Hypotheses
H14a. There is a genuine supply/security trade-off: contracts strict enough to exclude the E009-pre precision cheat also exclude a material fraction of honest backends.
H14b. A float-tolerance contract can be placed in the gap E009-pre measured (honest GPU at 7.7 × 10^-4, cheat at 2.84 × 10^-3), but the gap is only 3.7× wide, so the contract will have almost no engineering margin and will be fragile to a change of hardware, size or workload.
H14c. An exact-integer contract will be the most exclusive of all — it is the strictest thing one can write — and will therefore cost the most supply.
Prediction that turned out to be wrong
H14c is recorded here because it was the research lead's expectation going in and
it is exactly backwards. See analysis.md.
Method note recorded before the fact
Operand structure and operand magnitude must be varied separately. A first version of this experiment used an all-127 matrix as its "worst case", which is simultaneously rank-1 and exactly representable in bfloat16 — so both cheating backends returned the exactly correct product and every tolerance contract appeared to fail. That was a property of the test data, not of the contracts, and it is the reason the regimes below are split.
Methodology
methodology.md ↗Implementation: implementation/contracts.py.
What a contract is here
A predicate on the provider's output, evaluated by the coordinator. Six candidates, from nothing to strictest:
| contract | clause |
|---|---|
C0-none |
no numeric requirement — this is Nodus v1 today |
C1-float-1e-2 |
float32 operands; mean relative deviation ≤ 10^-2 |
C2-float-1e-3 |
… ≤ 10^-3 |
C3-float-1e-5 |
… ≤ 10^-5 |
C4-exact-unbounded |
integer operands; product must be exactly equal |
C5-exact-tiled |
as C4, and the job must be tiled so that `max |
C5's extra clause is a constraint on how the job may be posed, not on whether a backend computed it correctly. The two are reported separately, so that "this job must be split into smaller tiles" is never confused with "no backend conforms".
Backends
Five honest (cpu-blas-f32, cpu-blocked-f32, cpu-f64, mlx-cpu-f32,
mlx-gpu-f32) and two cheating (mlx-gpu-bf16 — the E009-pre precision cheat;
cheat-low-rank — a rank-32 approximation).
Operand regimes — structure and magnitude varied separately
| regime | what it stresses |
|---|---|
random |
full-rank uniform int8 |
lowrank-16, lowrank-128 |
approximately rank r, quantised — the shape real activations have, and the case that makes a low-rank shortcut nearly correct |
spike |
full rank, but one row of A and one column of B set to +127, forcing ` |
The magnitude axis matters because float32 represents integers exactly only up
to 2^24. Worst-case |C| is max|A| · max|B| · n, while random data gives
only about max² · √n. A contract validated on random inputs is not
validated.
Metrics
Per (contract, size, regime):
- supply — honest backends conforming / honest backends total
- security — cheats excluded / cheats total
- job admissibility — whether the job as posed satisfies the contract's own tiling clause, and the maximum accumulation length that would
- worst-case and observed
|C|
Sizes n ∈ {1024, 2048, 4096}, four regimes, seven backends, six contracts.
Limitations
- Two cheating strategies, five honest backends, one machine. Cross-vendor hardware would be the real test of the supply claim.
- Contract conformance is evaluated against a float64 reference product, which a coordinator could not afford to compute per job. The cost of checking conformance cheaply is a separate question, addressed in the analysis.
- Matmul only.
Analysis
analysis.md ↗Status: complete. H14c falsified — the prediction was exactly backwards.
Date: 2026-09-01. Apple M5, NumPy 2.5.2 + Accelerate, MLX 0.32.2.
Raw records: ../../benchmarks/results/014-numeric-contract.jsonl
1. The headline
There is no supply/security trade-off, because the strictest contract is not
the most exclusive one. Measured at n = 2048:
| operands | worst honest backend | the cheat | usable window |
|---|---|---|---|
| real-valued (standard normal) | 7.70 × 10^-4 (GPU) | 2.84 × 10^-3 | 3.7× |
| integer (int8 range) | 0.0 — exact, every backend | 1.41 × 10^-3 | unbounded |
Read that second row carefully. On integer operands, cpu-blas-f32,
mlx-cpu-f32 and mlx-gpu-f32 all return the bit-exact product. The CPU and
the GPU agree perfectly.
The GPU's imprecision was never about the GPU. It was about representing real numbers. Given integers in range, the same hardware that diverged by 1,500× in E009-pre is exact.
This dissolves the problem E009-pre raised rather than trading against it. The 3.7× window that made a heterogeneous network unpoliceable was an artefact of floating-point operands, not a property of heterogeneous hardware.
2. Contract conformance
Across n ∈ {1024, 2048, 4096} and all four operand regimes:
| contract | supply (honest admitted) | security (cheats excluded) |
|---|---|---|
C0-none (Nodus v1 today) |
5/5 | 0/2 |
C1-float-1e-2 |
5/5 | 1/2 |
C2-float-1e-3 |
5/5 | 2/2 |
C3-float-1e-5 |
5/5 | 2/2 |
C4-exact-unbounded |
5/5 (3/5 on spike, n ≥ 2048) |
2/2 |
C5-exact-tiled |
5/5 by construction | 2/2 |
H14a and H14c falsified. The exact contract costs no supply on realistic
operands. H14b confirmed: a float-tolerance contract can be placed in the
gap — C2 at 10^-3 works — but it sits 1.3× above the honest GPU and 2.8× below
the cheat. That is not engineering margin; it is a knife-edge, and a different
GPU, size or workload moves either number.
C0 is the important row. Nodus v1 imposes no numeric requirement, so it
admits both cheats at every size and every regime. This is not a hypothetical
gap.
3. The magnitude finding, and why random testing hides it
On the spike regime at n ≥ 2048, exact conformance drops to 3/5 — and
the backends that fail are not the ones anyone would guess:
| backend | conforms on spike? |
|---|---|
cpu-blas-f32 |
no |
mlx-cpu-f32 |
no |
cpu-blocked-f32 |
yes |
cpu-f64 |
yes |
mlx-gpu-f32 |
yes |
On magnitude-stressed input, the GPU is exact and the CPU BLAS is not. The
cause is accumulator width: |C| = 127² · 2048 = 3.3 × 10^7 exceeds float32's
exact-integer limit of 2^24 = 1.68 × 10^7. The blocked CPU version survives
because its partial sums stay small; float64 survives trivially; the GPU
survives because its accumulation is evidently wider than its (reduced-precision)
operand path.
And random data never reveals this. On random, lowrank-16 and
lowrank-128, C4 admits 5/5 at every size — everything looks fine. Only the
spike regime exposes it, and spike is a perfectly legitimate input, not an
attack. Random |C| grows like max²·√n; worst-case grows like max²·n. At
n = 4096 that is a 64× difference, and it is the difference between passing
and failing.
A numeric contract validated on random inputs is not validated.
This is what C5's tiling clause exists for, and the clause has a derived
number: for int8 operands into a float32 accumulator,
max accumulation length k ≤ 2^24 / 127² = 1040
With that clause, every backend is exact by construction, on every input, regardless of hardware. It converts a hardware-dependent hope into a specification.
4. Why exact contracts also win on checking cost
A tolerance contract has a practical problem this experiment's method quietly exposes: to evaluate "mean relative deviation ≤ 10^-3" we computed a float64 reference product — i.e. we did the whole job. A coordinator cannot afford that, which is the entire premise of the programme.
So a tolerance contract must in practice be checked by a proxy — the Freivalds residual — and that is precisely the metric E009-pre measured at a 3.7× window. The tolerance contract's apparent comfort in §2 comes from a metric the verifier cannot cheaply compute.
An exact contract has no such gap: equality is checkable by exact Freivalds over
F_p at 4.6% of the computation (E003 variant D), with a 9.1 × 10^-13
soundness bound and no tuning parameter.
| tolerance contract | exact contract | |
|---|---|---|
| separation from cheats | 3.7× | unbounded |
| honest backends admitted | 5/5 | 5/5 |
| cheaply checkable? | no — proxy metric only | yes, 4.6% (E003-D) |
| tuning parameters | τ, per hardware generation |
none |
| survives new hardware | unknown | yes, if the tiling clause holds |
5. What this means for Nodus
The programme's recommendation, now with the mechanism attached:
Define verifiable job classes by their arithmetic. Integer operands, a declared accumulator width, and a tiling bound derived from the two. Then honest heterogeneous hardware agrees exactly, cheating is unboundedly separated, verification costs 4.6% of the computation, and there is nothing to tune.
A draft of the clause is in docs/numeric-contract.md.
Three consequences worth stating plainly:
- This is the cheapest thing in the system and it is upstream of everything else. It is a specification change, not a mechanism.
- It removes the hardest problem rather than solving it. E009-pre's heterogeneity problem does not need a cleverer threshold; it needs different arithmetic.
- It is a constraint on the workload, arrived at from a verification requirement — the strongest evidence yet for hypothesis H7.
6. What a malicious provider can still get away with
- Everything from E013: caching (a pricing problem), sub-contracting, and being legitimately efficient.
- Operating outside the contract entirely — a contract defines what correct means, it does not detect violations. Detection is still E003-D's job. The two are complementary and neither substitutes for the other.
- Exploiting an under-specified clause. We pinned dtype, accumulator width and tiling. A real contract must also pin the quantisation scheme, and probably operation order for non-associative reductions. Each unpinned degree of freedom is a cheaper specification someone can be correct with respect to.
7. Limitations
- The accuracy cost of integer inference is still UNKNOWN, and it is the price of this entire recommendation. Nothing here measures it.
- Matmul only. Non-linearities (softmax, layernorm) are not obviously expressible in exact integer arithmetic, and they are where this proposal is most likely to break. That is the next experiment.
- One machine, two vendors' worth of arithmetic at most (Apple CPU and Apple GPU). The supply claim needs NVIDIA and a non-Apple CPU before it can be trusted.
- Two cheating strategies.
mlx-gpu-f32's exactness on integers is measured, not explained; we infer a wider accumulator than its operand path but have not confirmed it.
Next steps
next_steps.md ↗1. Non-linearities — where this proposal most likely breaks [next]
NC-0.1 covers linear layers only, and says so. Softmax, layernorm and GELU are not obviously expressible in exact integer arithmetic: they involve division, exponentials and inverse square roots. If they cannot be pinned exactly, then a full inference pipeline is partly contractable and partly not, and the question becomes whether a hybrid contract — exact linear layers, tolerance-bound non-linearities — retains any of the security we just measured.
This is open clause O1, and it is the single thing that decides whether the whole numeric-contract direction scales to real inference. It should be the next experiment.
Note the literature already has a relevant result: ZIP (CCS 2025) exists precisely because non-linearities are the hard part for exact-semantics proofs, and its own numbers show linear layers still dominate cost. Read before building.
2. The accuracy cost of integer inference
Still UNKNOWN, and it is the price of everything above. Measure end-to-end task accuracy for a real quantised model against its float baseline. Without this number the recommendation "mandate integer semantics" is an engineering argument with no counterweight.
3. Cross-vendor validation of the supply claim
The claim "an exact contract costs no supply" rests on one Apple M5 — an Apple CPU and an Apple GPU. NVIDIA (with TF32 both on and off) and a non-Apple CPU BLAS are needed before it can be trusted. Blocked on hardware.
The specific thing to test: does an NVIDIA GPU return bit-exact integer products, or does its tensor-core path round operands the way MLX's does for reals? If it does not, the supply claim weakens and the tiling bound gets tighter.
4. Test NC-3 against real transformer shapes (O2)
Reduction dimensions in real models are 4096–16384, against a bound of 1040 for
int8-into-float32. That forces 4–16 tiles. Is that free (it is roughly what a
blocked GEMM does anyway) or a real cost? Measure. If int32 accumulators are
available the bound rises to 2^31/127² = 133,000 and the clause becomes
vacuous — worth checking which accumulators real backends actually expose.
5. Adversarial conformance testing (O3)
NC-8 asks providers to reproduce a reference vector. A provider that can recognise the test vector can pass it and cheat on real work. The test must be indistinguishable from real traffic — which probably means auditing real jobs at random rather than issuing synthetic ones. This connects directly to the main CP-008 audit-rate work.
6. Then fold into CP-010
The hybrid architecture question now has a concrete bottom layer: contract + exact Freivalds at 4.6%. What sits above it (stake, attestation, ZK for the privacy tail) can finally be evaluated against a real baseline instead of against re-execution.