NodusLab
E017 · CP-015

What the numeric contract costs in accuracy

we have told the network to switch to exact integer arithmetic. What does that cost in answer quality?

Checkpoint: CP-015 · Status: COMPLETE, run 2026-09-01 Question: we have told the network to switch to exact integer arithmetic. What does that cost in answer quality?

Answer: nothing measurable at 8 bits

MNIST MLP (784-256-256-10, LayerNorm + GELU), three seeds, 10,000 test images:

bits quantisation cost extra cost of determinism total vs float32
8 +0.0002 +0.0001 — sign flips across seeds +0.0003
6 +0.0006 −0.0001 +0.0005
4 +0.0086 +0.0038 +0.0124

The number this programme owed is the middle column: not "integer vs float" (ordinary quantisation is already integer and known to work) but the extra cost of being integer end to end — integer accumulation, integer layernorm, lookup-table GELU — with no float step anywhere to reintroduce cross-backend divergence.

At 8 bits that cost is one test image in ten thousand, and it changes sign between seeds. Indistinguishable from zero.

At 4 bits it becomes real (+0.38 points, same sign in all three seeds). The contract is free at 8 bits, cheap at 6, and starts to bite at 4 — a usable design boundary.

Determinism, on a real model

CPU and Metal GPU produced bit-identical logits across all 10,000 test images, at every bit width and seed. E015 showed this on random matrices; it now holds on a trained network doing a real task.

Why it matters

EigenAI achieves determinism by restricting the network to identical H100s, and lists "portable numeric normalization to enable heterogeneous verifier sets" as future work. We have measured a portable alternative and shown it costs no accuracy at 8 bits. For a network built on heterogeneous consumer hardware, that is the difference between a workable design and an unworkable one.

This also closes the last blocking unknown on NC-0.2.

Files

file contents
hypothesis.md H17a–c and why B − C is the number that matters
methodology.md three configurations, and the bug that made the first run score at chance
implementation/accuracy.py training, all three inference paths, cross-backend check
analysis.md the result
next_steps.md a transformer, and a harder task

Raw records: ../../benchmarks/results/017-accuracy-cost.jsonl

Caveat we are not hiding

MNIST is easy. 98% accuracy leaves headroom a harder task would not. Read this as "no measurable cost on this task", not "no cost".

Checkpoint CP-015

What does the numeric contract cost in answer quality?

Why this is the bill we owe

Everything since E014 rests on one recommendation: verifiable job classes should use exact integer arithmetic, because that is what makes honest heterogeneous hardware agree bit-for-bit. It is the thing that distinguishes our design from EigenAI's (which achieves determinism by restricting everyone to identical H100s), and it is required by both settlement models — optimistic and inline — because re-execution only resolves a dispute if honest machines agree.

We have never measured what it costs. Every experiment since E014 has carried "the accuracy cost of integer inference is UNKNOWN" as a limitation. This experiment closes it.

The measurement that actually matters

Not "integer vs float". Ordinary quantised inference is already integer-ish and is known to work — llama.cpp, GPTQ and AWQ all ship it, and nobody needs convincing.

What is unusual about NC-0.2 is that it is integer end to end. Standard int8 inference still accumulates in float32 and computes softmax and layernorm in float. Ours cannot: any float step reintroduces exactly the cross-backend divergence E009-pre measured.

So three configurations on the same trained weights:

accumulation layernorm / GELU deterministic across hardware?
A float32 reference float float no
B standard quantisation float float no
C NC-0.2 contract integer, tiled per NC-3 integer, LUT per NC-9 yes
  • A − B is the cost of quantisation, which the industry already accepts.
  • B − C is the extra cost of determinism — the number this programme owes.

Hypotheses

H17a. A − B is small at 8 bits (well under one accuracy point). Not novel; included as a sanity check that our quantisation is competently implemented.

H17b (the one that matters). B − C is also small at 8 bits — the extra cost of making everything integer, on top of already being quantised, is near zero.

H17c. The extra cost of determinism grows as precision shrinks, because integer layernorm and a LUT GELU have less headroom to absorb rounding at 4 bits than at 8.

Prediction recorded before running

B − C under 0.5 accuracy points at 8 bits.

What a bad result would mean

If determinism costs a meaningful amount of accuracy, then the portable numeric contract — our main differentiator against architecture-locked determinism — is a worse trade than it looks, and the honest recommendation becomes EigenAI's: fix the hardware instead of fixing the arithmetic.

Implementation: implementation/accuracy.py.

Task and model

MNIST, 60,000 train / 10,000 test. MLP 784 → 256 → 256 → 10, LayerNorm and GELU at each hidden layer. Trained in float32 with MLX on the GPU, Adam, 10 epochs. The same trained weights are used for all three configurations — no quantisation-aware training, no fine-tuning, so the comparison isolates inference arithmetic.

Three seeds (0, 1, 2), each a fresh training run.

The three configurations

A — float32 reference. Plain numpy float64 evaluation of the trained weights.

B — standard quantisation. Per-tensor symmetric quantisation of weights and activations to b bits, round-half-to-even (NC-5). Matmuls are computed on integers but accumulated and rescaled in float, and layernorm and GELU are float. This is what ordinary quantised inference does.

C — NC-0.2 contract. Integer end to end:

  • matmuls tiled per NC-3 so no partial sum leaves fp32's exact-integer range;
  • layernorm exact in integers — mean and variance by integer division, an exact integer square root, and division by that root directly (see below);
  • GELU as an integer lookup table (NC-9) over a fixed-point domain of ±8.0 at 64 units per 1.0;
  • every tensor carried as an integer plus a known scale; nothing returns to float at any point.

An implementation bug worth recording

The first version computed a fixed-point reciprocal inv = (1<<12) // isqrt(var) for layernorm. Post-matmul variances reach ~5 × 10^13, so isqrt(var) ≈ 7 × 10^6 and the reciprocal underflowed to zero, collapsing accuracy to 10.3% — exactly chance for ten classes.

The fix is to divide by the integer square root directly, (d * LN_SCALE) // root, which needs no reciprocal and cannot underflow. Recorded because "the integer version scores at chance" looks like a damning result about the contract and was in fact a two-line arithmetic error on our side.

Stage-by-stage validation against the float path was added and is how it was found: correlations of 1.0000 (pre-layernorm), 0.9999 (post-layernorm) and 0.9992 (post-GELU) confirm the integer pipeline tracks the float one.

Cross-backend check

Configuration C is run twice — matmuls on numpy/BLAS and on the Metal GPU via MLX — and the full logit tensors are compared for bit-identity on all 10,000 test images. This repeats E015's determinism result on a real trained model rather than random matrices.

Metrics

Test accuracy for A, B and C at 8, 6 and 4 bits, plus the two differences A − B and B − C. Reported across three seeds, because at 8 bits the differences amount to a handful of test images and a single run cannot distinguish them from noise.

Limitations

  • MNIST is easy. 98% accuracy leaves large headroom; a harder task may be less forgiving.
  • One architecture, one dataset, three seeds.
  • No attention. The model has no softmax on the accuracy-critical path (argmax is invariant to it anyway), so E015's integer softmax is not exercised here.
  • Post-training quantisation only. Quantisation-aware training would likely improve the 4-bit numbers for both B and C.

Status: complete. H17a, H17b and H17c all confirmed. Date: 2026-09-01. Apple M5; MLX 0.32.2 (training); NumPy 2.5.2 (inference). Raw records: ../../benchmarks/results/017-accuracy-cost.jsonl


1. The answer

At 8 bits, making inference fully deterministic costs nothing measurable.

Three seeds, MNIST, 10,000 test images:

bits seed float32 standard quant NC-0.2 contract quantisation cost determinism cost
8 0 0.9822 0.9820 0.9816 +0.0002 +0.0004
8 1 0.9816 0.9812 0.9815 +0.0004 −0.0003
8 2 0.9795 0.9796 0.9793 −0.0001 +0.0003
mean +0.0002 +0.0001

The extra cost of full determinism at 8 bits is +0.0001 accuracy — one test image in ten thousand — and it changes sign across seeds. It is indistinguishable from zero.

Total cost against the float32 baseline is +0.0003, i.e. three images in ten thousand.

H17b confirmed, and comfortably inside the predicted 0.5-point bound.

2. Determinism holds on a real model

Configuration C was run with its matmuls on numpy/BLAS and on the Metal GPU, and the full 10,000 × 10 logit tensors compared for bit-identity:

bits 8 6 4
CPU vs GPU bit-identical yes yes yes

At every bit width and every seed. E015 established this on random matrices; it now holds on a trained network doing a real task.

3. The trend: determinism is free at 8 bits, not at 4

bits quantisation cost (A−B) determinism cost (B−C)
8 +0.0002 +0.0001 (sign flips — zero)
6 +0.0006 −0.0001 (zero)
4 +0.0086 +0.0038 (consistent across seeds)

H17c confirmed. At 4 bits the extra cost of determinism is +0.38 accuracy points, on top of quantisation's own +0.86 — and unlike the 8-bit figure it is the same sign in all three seeds, so it is real.

The mechanism is headroom: integer layernorm and a lookup-table GELU have to round into a 4-bit codebook, and there is not enough of it. The contract is free at 8 bits, cheap at 6, and starts to bite at 4. That is a usable design boundary rather than a blanket claim.

4. What this means for the programme

The recommendation made in E014 and carried through E015 and E016 — mandate exact integer semantics for verifiable job classes — now has its price tag, and the price at 8 bits is zero within measurement noise.

That matters most for the comparison with EigenAI (literature.md §5b). They achieve determinism by restricting the network to one GPU architecture, and list "portable numeric normalization to enable heterogeneous verifier sets" as future work. We have now measured a portable alternative and shown it costs no accuracy at 8 bits. For a network whose premise is heterogeneous consumer hardware, that is the difference between a workable design and an unworkable one.

It also removes the last blocking unknown from the numeric contract. NC-0.2 is specified (E014, E015), its determinism is measured on random matrices and on a real model (E015, E017), its verification cost is measured (E003-D, E015), its bandwidth cost is measured and its architecture consequence understood (E016), and its accuracy cost is now measured here.

5. An error on our side, recorded

The first implementation scored 10.32% — exactly chance. That reads as a devastating result about integer inference and was a two-line bug: a fixed-point reciprocal (1 << 12) // isqrt(var) underflowed to zero, because post-matmul variances reach ~5 × 10^13 and isqrt(var) ≈ 7 × 10^6.

Dividing by the integer square root directly removes the reciprocal and cannot underflow. Stage-by-stage correlation against the float path (1.0000 / 0.9999 / 0.9992) is what located it, and that check is now part of the implementation.

The general lesson, which applies to every integer-arithmetic result in this repository: a broken integer pipeline fails silently and looks like evidence against the approach. Validate stage by stage against a float reference, always.

6. Limitations

  • MNIST is easy, and 98% leaves headroom that a harder task would not. This is the single largest caveat: the result should be read as "no measurable cost on this task", not "no cost".
  • One architecture, one dataset, three seeds.
  • No attention. The model has no softmax on the accuracy-critical path, so E015's integer softmax is specified and determinism-tested but not accuracy-tested. A transformer is the obvious next step.
  • Post-training quantisation only; quantisation-aware training would likely improve both B and C at 4 bits and might change the 4-bit gap.
  • Accuracy is the only quality metric. Generation tasks would need perplexity or a task score, and small per-token deviations could compound differently.

1. A harder task, and a transformer

MNIST at 98% has headroom that hides small regressions. Two things need repeating on something less forgiving:

  • A harder classification task — CIFAR-10, or MNIST-Back-Rand, which SafetyNets used precisely because it is harder.
  • A transformer with real attention. E015 specified and determinism-tested integer softmax, but E017 never exercised it for accuracy, because argmax over logits is invariant to softmax. Attention puts softmax on the critical path, where its rounding does affect the output.

Until then the claim is "no measurable cost on an easy task", and should be written that way everywhere it appears.

2. Generation, not classification

Accuracy on a classifier is forgiving: a small logit perturbation usually does not flip the argmax. Autoregressive generation is not — a single flipped token changes everything downstream. Measure perplexity, and the distribution of the first divergent token, on a small LM under the contract.

This is also the deferred cross-backend decoding experiment from E009-pre's next-steps, and the two should be done together.

3. Quantisation-aware training

Everything here is post-training quantisation. QAT would likely close much of the 4-bit gap for both standard quantisation and the contract, and would tell us whether the 4-bit determinism penalty is intrinsic or just an artefact of not training for it.

4. Feed the result back into the contract

NC-0.2 should state a recommended precision floor: 8 bits free, 6 bits cheap, 4 bits carries a measured penalty. That is a concrete clause a job class can reference, and it is now backed by data.