NodusLab
E014 · CP-012

The numeric contract

what does a numeric contract cost in supply, and what does it buy in security?

Checkpoint: CP-012 · Status: COMPLETE, run 2026-09-01 Question: what does a numeric contract cost in supply, and what does it buy in security?

Result — the expected trade-off does not exist

operands worst honest backend the cheat usable window
real-valued 7.70 × 10^-4 (GPU) 2.84 × 10^-3 3.7×
integer 0.0 — exact, every backend 1.41 × 10^-3 unbounded

On integer operands, CPU BLAS, MLX-CPU and MLX-GPU all return the bit-exact product.

The GPU's imprecision was never about the GPU — it was about representing real numbers. Given integers in range, the same hardware that diverged by 1,500× in E009-pre agrees perfectly.

So the strictest contract is not the most exclusive one. H14c was predicted and is falsified: the exact contract costs no supply (5/5 honest backends admitted on every realistic operand regime) while separating cheats without bound. E009-pre's heterogeneity problem does not need a cleverer threshold; it needs different arithmetic.

Contract conformance

contract supply security
C0-none (Nodus v1 today) 5/5 0/2
C1-float-1e-2 5/5 1/2
C2-float-1e-3 5/5 2/2 — but on a 3.7× knife-edge
C4-exact-unbounded 5/5 (3/5 on spike) 2/2
C5-exact-tiled 5/5 by construction 2/2

C0 is the row that matters: v1 imposes no numeric requirement and admits both cheats at every size and regime.

The finding random testing would have hidden

On magnitude-stressed input, exact conformance drops to 3/5 — and the GPU passes while CPU BLAS fails. Worst-case |C| is M²·n but random data only reaches M²·√n, a 64× gap at n = 4096, and float32 is exact on integers only to 2^24.

A numeric contract validated on random inputs is not validated.

Hence the tiling clause, with a derived number: for int8 into float32, accumulation length ≤ 1040. With it, every backend is exact by construction on every input.

Deliverable

docs/numeric-contract.md — draft NC-0.1, eight clauses, a worked example, and an explicit statement of what it does not cover (non-linearities, the accuracy cost of quantisation, cross-vendor validation).

Files

file contents
hypothesis.md H14a–c, including the prediction that was backwards
methodology.md six contracts, seven backends, four operand regimes
implementation/contracts.py the experiment
analysis.md the result, and why exact contracts also win on checking cost
next_steps.md non-linearities are where this breaks next

Raw records: ../../benchmarks/results/014-numeric-contract.jsonl

A method error worth reading

The first version of this experiment used an all-127 matrix as its "worst case". That matrix is rank-1 and exactly representable in bfloat16, so both cheating backends returned the exactly correct product and every tolerance contract appeared to admit every cheat. The conclusion was wrong and was a property of the test data. Operand structure and magnitude are varied separately for this reason.

Checkpoint CP-012

What does a numeric contract cost in supply, and what does it buy in security?

Why this is the next experiment

E009-pre and E013 reached the same residual risk from opposite directions:

A provider can be perfectly correct with respect to a cheaper specification than the one intended.

A job that names a model but not its arithmetic does not determine a correct answer, so no verifier can be right — and no combination of mechanisms reaches the problem, because each would faithfully verify the wrong thing. A ZK proof of the cheap circuit is a valid proof. An attestation of the cheap binary is a valid attestation. The specification sits upstream of every mechanism.

Stating the contract costs nothing. The question is what it costs to enforce.

The expected trade-off

A numeric contract is a clause in the job specification saying what arithmetic the provider must use, evaluated as a predicate on the output. The assumed shape of the problem:

  strict contract  ->  excludes cheats  ->  but also excludes honest hardware
                                            -> smaller provider pool
                                            -> less supply, higher prices
  loose contract   ->  admits everyone  ->  including the cheats

So we expect a curve, and the deliverable is where on it to sit.

Hypotheses

H14a. There is a genuine supply/security trade-off: contracts strict enough to exclude the E009-pre precision cheat also exclude a material fraction of honest backends.

H14b. A float-tolerance contract can be placed in the gap E009-pre measured (honest GPU at 7.7 × 10^-4, cheat at 2.84 × 10^-3), but the gap is only 3.7× wide, so the contract will have almost no engineering margin and will be fragile to a change of hardware, size or workload.

H14c. An exact-integer contract will be the most exclusive of all — it is the strictest thing one can write — and will therefore cost the most supply.

Prediction that turned out to be wrong

H14c is recorded here because it was the research lead's expectation going in and it is exactly backwards. See analysis.md.

Method note recorded before the fact

Operand structure and operand magnitude must be varied separately. A first version of this experiment used an all-127 matrix as its "worst case", which is simultaneously rank-1 and exactly representable in bfloat16 — so both cheating backends returned the exactly correct product and every tolerance contract appeared to fail. That was a property of the test data, not of the contracts, and it is the reason the regimes below are split.

Implementation: implementation/contracts.py.

What a contract is here

A predicate on the provider's output, evaluated by the coordinator. Six candidates, from nothing to strictest:

contract clause
C0-none no numeric requirement — this is Nodus v1 today
C1-float-1e-2 float32 operands; mean relative deviation ≤ 10^-2
C2-float-1e-3 … ≤ 10^-3
C3-float-1e-5 … ≤ 10^-5
C4-exact-unbounded integer operands; product must be exactly equal
C5-exact-tiled as C4, and the job must be tiled so that `max

C5's extra clause is a constraint on how the job may be posed, not on whether a backend computed it correctly. The two are reported separately, so that "this job must be split into smaller tiles" is never confused with "no backend conforms".

Backends

Five honest (cpu-blas-f32, cpu-blocked-f32, cpu-f64, mlx-cpu-f32, mlx-gpu-f32) and two cheating (mlx-gpu-bf16 — the E009-pre precision cheat; cheat-low-rank — a rank-32 approximation).

Operand regimes — structure and magnitude varied separately

regime what it stresses
random full-rank uniform int8
lowrank-16, lowrank-128 approximately rank r, quantised — the shape real activations have, and the case that makes a low-rank shortcut nearly correct
spike full rank, but one row of A and one column of B set to +127, forcing `

The magnitude axis matters because float32 represents integers exactly only up to 2^24. Worst-case |C| is max|A| · max|B| · n, while random data gives only about max² · √n. A contract validated on random inputs is not validated.

Metrics

Per (contract, size, regime):

  • supply — honest backends conforming / honest backends total
  • security — cheats excluded / cheats total
  • job admissibility — whether the job as posed satisfies the contract's own tiling clause, and the maximum accumulation length that would
  • worst-case and observed |C|

Sizes n ∈ {1024, 2048, 4096}, four regimes, seven backends, six contracts.

Limitations

  • Two cheating strategies, five honest backends, one machine. Cross-vendor hardware would be the real test of the supply claim.
  • Contract conformance is evaluated against a float64 reference product, which a coordinator could not afford to compute per job. The cost of checking conformance cheaply is a separate question, addressed in the analysis.
  • Matmul only.

Status: complete. H14c falsified — the prediction was exactly backwards. Date: 2026-09-01. Apple M5, NumPy 2.5.2 + Accelerate, MLX 0.32.2. Raw records: ../../benchmarks/results/014-numeric-contract.jsonl


1. The headline

There is no supply/security trade-off, because the strictest contract is not the most exclusive one. Measured at n = 2048:

operands worst honest backend the cheat usable window
real-valued (standard normal) 7.70 × 10^-4 (GPU) 2.84 × 10^-3 3.7×
integer (int8 range) 0.0 — exact, every backend 1.41 × 10^-3 unbounded

Read that second row carefully. On integer operands, cpu-blas-f32, mlx-cpu-f32 and mlx-gpu-f32 all return the bit-exact product. The CPU and the GPU agree perfectly.

The GPU's imprecision was never about the GPU. It was about representing real numbers. Given integers in range, the same hardware that diverged by 1,500× in E009-pre is exact.

This dissolves the problem E009-pre raised rather than trading against it. The 3.7× window that made a heterogeneous network unpoliceable was an artefact of floating-point operands, not a property of heterogeneous hardware.

2. Contract conformance

Across n ∈ {1024, 2048, 4096} and all four operand regimes:

contract supply (honest admitted) security (cheats excluded)
C0-none (Nodus v1 today) 5/5 0/2
C1-float-1e-2 5/5 1/2
C2-float-1e-3 5/5 2/2
C3-float-1e-5 5/5 2/2
C4-exact-unbounded 5/5 (3/5 on spike, n ≥ 2048) 2/2
C5-exact-tiled 5/5 by construction 2/2

H14a and H14c falsified. The exact contract costs no supply on realistic operands. H14b confirmed: a float-tolerance contract can be placed in the gap — C2 at 10^-3 works — but it sits 1.3× above the honest GPU and 2.8× below the cheat. That is not engineering margin; it is a knife-edge, and a different GPU, size or workload moves either number.

C0 is the important row. Nodus v1 imposes no numeric requirement, so it admits both cheats at every size and every regime. This is not a hypothetical gap.

3. The magnitude finding, and why random testing hides it

On the spike regime at n ≥ 2048, exact conformance drops to 3/5 — and the backends that fail are not the ones anyone would guess:

backend conforms on spike?
cpu-blas-f32 no
mlx-cpu-f32 no
cpu-blocked-f32 yes
cpu-f64 yes
mlx-gpu-f32 yes

On magnitude-stressed input, the GPU is exact and the CPU BLAS is not. The cause is accumulator width: |C| = 127² · 2048 = 3.3 × 10^7 exceeds float32's exact-integer limit of 2^24 = 1.68 × 10^7. The blocked CPU version survives because its partial sums stay small; float64 survives trivially; the GPU survives because its accumulation is evidently wider than its (reduced-precision) operand path.

And random data never reveals this. On random, lowrank-16 and lowrank-128, C4 admits 5/5 at every size — everything looks fine. Only the spike regime exposes it, and spike is a perfectly legitimate input, not an attack. Random |C| grows like max²·√n; worst-case grows like max²·n. At n = 4096 that is a 64× difference, and it is the difference between passing and failing.

A numeric contract validated on random inputs is not validated.

This is what C5's tiling clause exists for, and the clause has a derived number: for int8 operands into a float32 accumulator,

    max accumulation length  k ≤ 2^24 / 127²  =  1040

With that clause, every backend is exact by construction, on every input, regardless of hardware. It converts a hardware-dependent hope into a specification.

4. Why exact contracts also win on checking cost

A tolerance contract has a practical problem this experiment's method quietly exposes: to evaluate "mean relative deviation ≤ 10^-3" we computed a float64 reference product — i.e. we did the whole job. A coordinator cannot afford that, which is the entire premise of the programme.

So a tolerance contract must in practice be checked by a proxy — the Freivalds residual — and that is precisely the metric E009-pre measured at a 3.7× window. The tolerance contract's apparent comfort in §2 comes from a metric the verifier cannot cheaply compute.

An exact contract has no such gap: equality is checkable by exact Freivalds over F_p at 4.6% of the computation (E003 variant D), with a 9.1 × 10^-13 soundness bound and no tuning parameter.

tolerance contract exact contract
separation from cheats 3.7× unbounded
honest backends admitted 5/5 5/5
cheaply checkable? no — proxy metric only yes, 4.6% (E003-D)
tuning parameters τ, per hardware generation none
survives new hardware unknown yes, if the tiling clause holds

5. What this means for Nodus

The programme's recommendation, now with the mechanism attached:

Define verifiable job classes by their arithmetic. Integer operands, a declared accumulator width, and a tiling bound derived from the two. Then honest heterogeneous hardware agrees exactly, cheating is unboundedly separated, verification costs 4.6% of the computation, and there is nothing to tune.

A draft of the clause is in docs/numeric-contract.md.

Three consequences worth stating plainly:

  1. This is the cheapest thing in the system and it is upstream of everything else. It is a specification change, not a mechanism.
  2. It removes the hardest problem rather than solving it. E009-pre's heterogeneity problem does not need a cleverer threshold; it needs different arithmetic.
  3. It is a constraint on the workload, arrived at from a verification requirement — the strongest evidence yet for hypothesis H7.

6. What a malicious provider can still get away with

  • Everything from E013: caching (a pricing problem), sub-contracting, and being legitimately efficient.
  • Operating outside the contract entirely — a contract defines what correct means, it does not detect violations. Detection is still E003-D's job. The two are complementary and neither substitutes for the other.
  • Exploiting an under-specified clause. We pinned dtype, accumulator width and tiling. A real contract must also pin the quantisation scheme, and probably operation order for non-associative reductions. Each unpinned degree of freedom is a cheaper specification someone can be correct with respect to.

7. Limitations

  • The accuracy cost of integer inference is still UNKNOWN, and it is the price of this entire recommendation. Nothing here measures it.
  • Matmul only. Non-linearities (softmax, layernorm) are not obviously expressible in exact integer arithmetic, and they are where this proposal is most likely to break. That is the next experiment.
  • One machine, two vendors' worth of arithmetic at most (Apple CPU and Apple GPU). The supply claim needs NVIDIA and a non-Apple CPU before it can be trusted.
  • Two cheating strategies.
  • mlx-gpu-f32's exactness on integers is measured, not explained; we infer a wider accumulator than its operand path but have not confirmed it.

1. Non-linearities — where this proposal most likely breaks [next]

NC-0.1 covers linear layers only, and says so. Softmax, layernorm and GELU are not obviously expressible in exact integer arithmetic: they involve division, exponentials and inverse square roots. If they cannot be pinned exactly, then a full inference pipeline is partly contractable and partly not, and the question becomes whether a hybrid contract — exact linear layers, tolerance-bound non-linearities — retains any of the security we just measured.

This is open clause O1, and it is the single thing that decides whether the whole numeric-contract direction scales to real inference. It should be the next experiment.

Note the literature already has a relevant result: ZIP (CCS 2025) exists precisely because non-linearities are the hard part for exact-semantics proofs, and its own numbers show linear layers still dominate cost. Read before building.

2. The accuracy cost of integer inference

Still UNKNOWN, and it is the price of everything above. Measure end-to-end task accuracy for a real quantised model against its float baseline. Without this number the recommendation "mandate integer semantics" is an engineering argument with no counterweight.

3. Cross-vendor validation of the supply claim

The claim "an exact contract costs no supply" rests on one Apple M5 — an Apple CPU and an Apple GPU. NVIDIA (with TF32 both on and off) and a non-Apple CPU BLAS are needed before it can be trusted. Blocked on hardware.

The specific thing to test: does an NVIDIA GPU return bit-exact integer products, or does its tensor-core path round operands the way MLX's does for reals? If it does not, the supply claim weakens and the tiling bound gets tighter.

4. Test NC-3 against real transformer shapes (O2)

Reduction dimensions in real models are 4096–16384, against a bound of 1040 for int8-into-float32. That forces 4–16 tiles. Is that free (it is roughly what a blocked GEMM does anyway) or a real cost? Measure. If int32 accumulators are available the bound rises to 2^31/127² = 133,000 and the clause becomes vacuous — worth checking which accumulators real backends actually expose.

5. Adversarial conformance testing (O3)

NC-8 asks providers to reproduce a reference vector. A provider that can recognise the test vector can pass it and cheat on real work. The test must be indistinguishable from real traffic — which probably means auditing real jobs at random rather than issuing synthetic ones. This connects directly to the main CP-008 audit-rate work.

6. Then fold into CP-010

The hybrid architecture question now has a concrete bottom layer: contract + exact Freivalds at 4.6%. What sits above it (stake, attestation, ZK for the privacy tail) can finally be evaluated against a real baseline instead of against re-execution.