What the numeric contract costs in accuracy
we have told the network to switch to exact integer arithmetic. What does that cost in answer quality?
Overview
README.md ↗Checkpoint: CP-015 · Status: COMPLETE, run 2026-09-01 Question: we have told the network to switch to exact integer arithmetic. What does that cost in answer quality?
Answer: nothing measurable at 8 bits
MNIST MLP (784-256-256-10, LayerNorm + GELU), three seeds, 10,000 test images:
| bits | quantisation cost | extra cost of determinism | total vs float32 |
|---|---|---|---|
| 8 | +0.0002 | +0.0001 — sign flips across seeds | +0.0003 |
| 6 | +0.0006 | −0.0001 | +0.0005 |
| 4 | +0.0086 | +0.0038 | +0.0124 |
The number this programme owed is the middle column: not "integer vs float" (ordinary quantisation is already integer and known to work) but the extra cost of being integer end to end — integer accumulation, integer layernorm, lookup-table GELU — with no float step anywhere to reintroduce cross-backend divergence.
At 8 bits that cost is one test image in ten thousand, and it changes sign between seeds. Indistinguishable from zero.
At 4 bits it becomes real (+0.38 points, same sign in all three seeds). The contract is free at 8 bits, cheap at 6, and starts to bite at 4 — a usable design boundary.
Determinism, on a real model
CPU and Metal GPU produced bit-identical logits across all 10,000 test images, at every bit width and seed. E015 showed this on random matrices; it now holds on a trained network doing a real task.
Why it matters
EigenAI achieves determinism by restricting the network to identical H100s, and lists "portable numeric normalization to enable heterogeneous verifier sets" as future work. We have measured a portable alternative and shown it costs no accuracy at 8 bits. For a network built on heterogeneous consumer hardware, that is the difference between a workable design and an unworkable one.
This also closes the last blocking unknown on NC-0.2.
Files
| file | contents |
|---|---|
hypothesis.md |
H17a–c and why B − C is the number that matters |
methodology.md |
three configurations, and the bug that made the first run score at chance |
implementation/accuracy.py |
training, all three inference paths, cross-backend check |
analysis.md |
the result |
next_steps.md |
a transformer, and a harder task |
Raw records: ../../benchmarks/results/017-accuracy-cost.jsonl
Caveat we are not hiding
MNIST is easy. 98% accuracy leaves headroom a harder task would not. Read this as "no measurable cost on this task", not "no cost".
Hypothesis
hypothesis.md ↗Checkpoint CP-015
What does the numeric contract cost in answer quality?
Why this is the bill we owe
Everything since E014 rests on one recommendation: verifiable job classes should use exact integer arithmetic, because that is what makes honest heterogeneous hardware agree bit-for-bit. It is the thing that distinguishes our design from EigenAI's (which achieves determinism by restricting everyone to identical H100s), and it is required by both settlement models — optimistic and inline — because re-execution only resolves a dispute if honest machines agree.
We have never measured what it costs. Every experiment since E014 has carried "the accuracy cost of integer inference is UNKNOWN" as a limitation. This experiment closes it.
The measurement that actually matters
Not "integer vs float". Ordinary quantised inference is already integer-ish and is known to work — llama.cpp, GPTQ and AWQ all ship it, and nobody needs convincing.
What is unusual about NC-0.2 is that it is integer end to end. Standard int8 inference still accumulates in float32 and computes softmax and layernorm in float. Ours cannot: any float step reintroduces exactly the cross-backend divergence E009-pre measured.
So three configurations on the same trained weights:
| accumulation | layernorm / GELU | deterministic across hardware? | |
|---|---|---|---|
| A float32 reference | float | float | no |
| B standard quantisation | float | float | no |
| C NC-0.2 contract | integer, tiled per NC-3 | integer, LUT per NC-9 | yes |
A − Bis the cost of quantisation, which the industry already accepts.B − Cis the extra cost of determinism — the number this programme owes.
Hypotheses
H17a. A − B is small at 8 bits (well under one accuracy point). Not novel;
included as a sanity check that our quantisation is competently implemented.
H17b (the one that matters). B − C is also small at 8 bits — the extra
cost of making everything integer, on top of already being quantised, is near
zero.
H17c. The extra cost of determinism grows as precision shrinks, because integer layernorm and a LUT GELU have less headroom to absorb rounding at 4 bits than at 8.
Prediction recorded before running
B − C under 0.5 accuracy points at 8 bits.
What a bad result would mean
If determinism costs a meaningful amount of accuracy, then the portable numeric contract — our main differentiator against architecture-locked determinism — is a worse trade than it looks, and the honest recommendation becomes EigenAI's: fix the hardware instead of fixing the arithmetic.
Methodology
methodology.md ↗Implementation: implementation/accuracy.py.
Task and model
MNIST, 60,000 train / 10,000 test. MLP 784 → 256 → 256 → 10, LayerNorm and
GELU at each hidden layer. Trained in float32 with MLX on the GPU, Adam, 10
epochs. The same trained weights are used for all three configurations — no
quantisation-aware training, no fine-tuning, so the comparison isolates
inference arithmetic.
Three seeds (0, 1, 2), each a fresh training run.
The three configurations
A — float32 reference. Plain numpy float64 evaluation of the trained weights.
B — standard quantisation. Per-tensor symmetric quantisation of weights and
activations to b bits, round-half-to-even (NC-5). Matmuls are computed on
integers but accumulated and rescaled in float, and layernorm and GELU are
float. This is what ordinary quantised inference does.
C — NC-0.2 contract. Integer end to end:
- matmuls tiled per NC-3 so no partial sum leaves fp32's exact-integer range;
- layernorm exact in integers — mean and variance by integer division, an exact integer square root, and division by that root directly (see below);
- GELU as an integer lookup table (NC-9) over a fixed-point domain of ±8.0 at 64 units per 1.0;
- every tensor carried as an integer plus a known scale; nothing returns to float at any point.
An implementation bug worth recording
The first version computed a fixed-point reciprocal inv = (1<<12) // isqrt(var)
for layernorm. Post-matmul variances reach ~5 × 10^13, so isqrt(var) ≈ 7 × 10^6
and the reciprocal underflowed to zero, collapsing accuracy to 10.3% —
exactly chance for ten classes.
The fix is to divide by the integer square root directly, (d * LN_SCALE) // root, which needs no reciprocal and cannot underflow. Recorded because "the
integer version scores at chance" looks like a damning result about the contract
and was in fact a two-line arithmetic error on our side.
Stage-by-stage validation against the float path was added and is how it was found: correlations of 1.0000 (pre-layernorm), 0.9999 (post-layernorm) and 0.9992 (post-GELU) confirm the integer pipeline tracks the float one.
Cross-backend check
Configuration C is run twice — matmuls on numpy/BLAS and on the Metal GPU via MLX — and the full logit tensors are compared for bit-identity on all 10,000 test images. This repeats E015's determinism result on a real trained model rather than random matrices.
Metrics
Test accuracy for A, B and C at 8, 6 and 4 bits, plus the two differences
A − B and B − C. Reported across three seeds, because at 8 bits the
differences amount to a handful of test images and a single run cannot
distinguish them from noise.
Limitations
- MNIST is easy. 98% accuracy leaves large headroom; a harder task may be less forgiving.
- One architecture, one dataset, three seeds.
- No attention. The model has no softmax on the accuracy-critical path (argmax is invariant to it anyway), so E015's integer softmax is not exercised here.
- Post-training quantisation only. Quantisation-aware training would likely improve the 4-bit numbers for both B and C.
Analysis
analysis.md ↗Status: complete. H17a, H17b and H17c all confirmed.
Date: 2026-09-01. Apple M5; MLX 0.32.2 (training); NumPy 2.5.2 (inference).
Raw records: ../../benchmarks/results/017-accuracy-cost.jsonl
1. The answer
At 8 bits, making inference fully deterministic costs nothing measurable.
Three seeds, MNIST, 10,000 test images:
| bits | seed | float32 | standard quant | NC-0.2 contract | quantisation cost | determinism cost |
|---|---|---|---|---|---|---|
| 8 | 0 | 0.9822 | 0.9820 | 0.9816 | +0.0002 | +0.0004 |
| 8 | 1 | 0.9816 | 0.9812 | 0.9815 | +0.0004 | −0.0003 |
| 8 | 2 | 0.9795 | 0.9796 | 0.9793 | −0.0001 | +0.0003 |
| mean | +0.0002 | +0.0001 |
The extra cost of full determinism at 8 bits is +0.0001 accuracy — one test image in ten thousand — and it changes sign across seeds. It is indistinguishable from zero.
Total cost against the float32 baseline is +0.0003, i.e. three images in ten thousand.
H17b confirmed, and comfortably inside the predicted 0.5-point bound.
2. Determinism holds on a real model
Configuration C was run with its matmuls on numpy/BLAS and on the Metal GPU, and the full 10,000 × 10 logit tensors compared for bit-identity:
| bits | 8 | 6 | 4 |
|---|---|---|---|
| CPU vs GPU bit-identical | yes | yes | yes |
At every bit width and every seed. E015 established this on random matrices; it now holds on a trained network doing a real task.
3. The trend: determinism is free at 8 bits, not at 4
| bits | quantisation cost (A−B) |
determinism cost (B−C) |
|---|---|---|
| 8 | +0.0002 | +0.0001 (sign flips — zero) |
| 6 | +0.0006 | −0.0001 (zero) |
| 4 | +0.0086 | +0.0038 (consistent across seeds) |
H17c confirmed. At 4 bits the extra cost of determinism is +0.38 accuracy points, on top of quantisation's own +0.86 — and unlike the 8-bit figure it is the same sign in all three seeds, so it is real.
The mechanism is headroom: integer layernorm and a lookup-table GELU have to round into a 4-bit codebook, and there is not enough of it. The contract is free at 8 bits, cheap at 6, and starts to bite at 4. That is a usable design boundary rather than a blanket claim.
4. What this means for the programme
The recommendation made in E014 and carried through E015 and E016 — mandate exact integer semantics for verifiable job classes — now has its price tag, and the price at 8 bits is zero within measurement noise.
That matters most for the comparison with EigenAI (literature.md §5b). They
achieve determinism by restricting the network to one GPU architecture, and list
"portable numeric normalization to enable heterogeneous verifier sets" as
future work. We have now measured a portable alternative and shown it costs no
accuracy at 8 bits. For a network whose premise is heterogeneous consumer
hardware, that is the difference between a workable design and an unworkable one.
It also removes the last blocking unknown from the numeric contract. NC-0.2 is specified (E014, E015), its determinism is measured on random matrices and on a real model (E015, E017), its verification cost is measured (E003-D, E015), its bandwidth cost is measured and its architecture consequence understood (E016), and its accuracy cost is now measured here.
5. An error on our side, recorded
The first implementation scored 10.32% — exactly chance. That reads as a
devastating result about integer inference and was a two-line bug: a fixed-point
reciprocal (1 << 12) // isqrt(var) underflowed to zero, because post-matmul
variances reach ~5 × 10^13 and isqrt(var) ≈ 7 × 10^6.
Dividing by the integer square root directly removes the reciprocal and cannot underflow. Stage-by-stage correlation against the float path (1.0000 / 0.9999 / 0.9992) is what located it, and that check is now part of the implementation.
The general lesson, which applies to every integer-arithmetic result in this repository: a broken integer pipeline fails silently and looks like evidence against the approach. Validate stage by stage against a float reference, always.
6. Limitations
- MNIST is easy, and 98% leaves headroom that a harder task would not. This is the single largest caveat: the result should be read as "no measurable cost on this task", not "no cost".
- One architecture, one dataset, three seeds.
- No attention. The model has no softmax on the accuracy-critical path, so E015's integer softmax is specified and determinism-tested but not accuracy-tested. A transformer is the obvious next step.
- Post-training quantisation only; quantisation-aware training would likely improve both B and C at 4 bits and might change the 4-bit gap.
- Accuracy is the only quality metric. Generation tasks would need perplexity or a task score, and small per-token deviations could compound differently.
Next steps
next_steps.md ↗1. A harder task, and a transformer
MNIST at 98% has headroom that hides small regressions. Two things need repeating on something less forgiving:
- A harder classification task — CIFAR-10, or MNIST-Back-Rand, which SafetyNets used precisely because it is harder.
- A transformer with real attention. E015 specified and determinism-tested integer softmax, but E017 never exercised it for accuracy, because argmax over logits is invariant to softmax. Attention puts softmax on the critical path, where its rounding does affect the output.
Until then the claim is "no measurable cost on an easy task", and should be written that way everywhere it appears.
2. Generation, not classification
Accuracy on a classifier is forgiving: a small logit perturbation usually does not flip the argmax. Autoregressive generation is not — a single flipped token changes everything downstream. Measure perplexity, and the distribution of the first divergent token, on a small LM under the contract.
This is also the deferred cross-backend decoding experiment from E009-pre's next-steps, and the two should be done together.
3. Quantisation-aware training
Everything here is post-training quantisation. QAT would likely close much of the 4-bit gap for both standard quantisation and the contract, and would tell us whether the 4-bit determinism penalty is intrinsic or just an artefact of not training for it.
4. Feed the result back into the contract
NC-0.2 should state a recommended precision floor: 8 bits free, 6 bits cheap, 4 bits carries a measured penalty. That is a concrete clause a job class can reference, and it is now backed by data.