NodusLab
E025

Error bars on the accuracy cost

Pre-registeredhypothesis recorded before the run. Written while the measurement wasexperiments/025-accuracy-error-bars/ ↗

Status: hypothesis recorded before the run. Written while the measurement was still executing; nothing below was edited afterwards.

Why

E023 reported three perplexities from 4–6 chunks of 128 tokens as bare point estimates:

path perplexity
float64 23.270
per-channel int8, float accumulation 24.883
NC-0.5 contract 25.762

and the programme's headline accuracy claim — the contract costs +3.5% over the best achievable int8 quantisation — is the difference between the last two.

A difference of two point estimates with no stated uncertainty is not a result. It is the first thing a reviewer will attack, and correctly: from 4–6 chunks that difference could plausibly be noise. paper/paper.md §11 already concedes this. E025 closes it.

The measurement is also not reproducible as it stands, because the token ids were derived ad hoc rather than committed. implementation/tokenise.py fixes that; it reproduces the E023 ids exactly (verified against conformance/qwen_ids96.json).

Design

120 non-overlapping 128-token chunks, evenly spaced across the whole wikitext-2 test set rather than taken from the front — the first N chunks are all one Wikipedia article, and perplexity varies far more between articles than between arithmetic pipelines.

All three pipelines see identical chunks, so the comparison is paired. The float and int8 paths share one code path with a single switch at the linear layers, so nothing but the arithmetic differs.

Uncertainty by bootstrap over chunks (20,000 resamples), computed on the same resamples for all three pipelines so the pairing survives into the intervals.

Hypotheses, recorded before looking

H1. The determinism gap (contract − int8) is real: its 95% CI excludes zero. Confidence: high. Every mechanism we understand predicts it — a fully integer pipeline has strictly more rounding boundaries than one that keeps accumulators exact between steps.

H2. The point estimate lands in 2–5%, i.e. near E023's +3.5% but not necessarily on it.

H3. The absolute perplexities will not reproduce 23.270 / 24.883 / 25.762, because those came from a handful of chunks at the front of the corpus and these are spread across all of it. This is a change of corpus sample, not a correction — but it means the paper's absolute numbers must be restated against a stated sample, and E023's should be marked superseded rather than quietly replaced.

H4. The interval on the gaps will be much tighter than the interval on any individual perplexity, because pairing cancels chunk difficulty. If this fails, the pairing is not working and the analysis is wrong.

H5. A 5-chunk sample — what E023 actually had — will give an interval wide enough to have been consistent with a gap of zero. This is the specific claim that the original measurement could not support its conclusion.

What would falsify the programme's claim

If the determinism gap's CI includes zero at 120 chunks, then "+3.5%" was noise and the honest statement becomes "we cannot distinguish the cost of determinism from zero at this sample size" — which would be a better result for the contract, not a worse one, and must be reported as readily.

If the gap is much larger than 5%, the contract is more expensive than the paper claims and §8.2 needs rewriting.

Files

implementation/tokenise.py corpus → token ids, cached and committed
implementation/ppl.py the three pipelines, per-chunk NLL
implementation/analyse.py bootstrap intervals, paired
implementation/records.jsonl one record per chunk, raw

All five pre-registered hypotheses confirmed, including H5, which says the earlier measurement could not support its own conclusion. The headline number moves: +3.5% → +2.43%, and the old estimate sits outside the new interval.

The result

120 chunks × 128 tokens, 15,240 scored positions, bootstrap over chunks (20,000 resamples). Raw records in implementation/records.jsonl.

pipeline perplexity 95% CI
float64 30.089 [27.676, 32.690]
per-channel int8, float accumulation 32.096 [29.495, 34.915]
NC-0.5 contract, integer end to end 32.876 [30.264, 35.683]

Those intervals are wide — about ±8% — and they are the uninteresting half of the result. Perplexity depends overwhelmingly on which chunks you draw: per-chunk values in this run range from 8.6 to 43.3, a 5× spread against pipeline differences of a few percent.

The quantity the contract is actually priced on is paired, and pairing cancels chunk difficulty:

gap estimate 95% CI sign holds
quantisation (int8 vs float) +6.67% [+5.95%, +7.40%] 113/120
determinism (contract vs int8) +2.43% [+1.63%, +3.23%] 83/120
both (contract vs float) +9.26% [+8.34%, +10.20%] 117/120

The cost of full integer determinism, over and above ordinary quantisation, is +2.43% perplexity [95% CI 1.63–3.23].

The interval on the gap is 1.6 points wide against ~5 points on either endpoint. That is H4 confirmed, and it is the whole reason the design is paired: an unpaired study would have needed vastly more chunks to say anything.

H5: the original measurement could not support its conclusion

Resampling at the sample size E023 actually used:

chunks 95% CI on the determinism gap width
5 [−1.36%, +6.47%] 7.8 pts
20 [+0.50%, +4.44%] 3.9 pts
120 [+1.64%, +3.23%] 1.6 pts

At five chunks the interval contains zero. E023 reported "+3.5%" from that evidence, and the honest statement available at the time was "we cannot distinguish this cost from zero." The conclusion happened to be right — the gap is real — but the evidence did not establish it. That is a methodological error independent of whether the answer was correct, and it is the kind that only looks harmless in hindsight.

Corrections to E023

1. The determinism cost was overstated: +3.5% → +2.43% [1.63, 3.23]. The old point estimate lies outside the new interval. The contract is cheaper than the paper claimed.

2. The absolute perplexities do not reproduce, and were never meant to be portable. E023 reported 23.270 / 24.883 / 25.762 from a handful of chunks at the front of the corpus; spread across the whole test set the same pipelines give 30.089 / 32.096 / 32.876. This is a change of corpus sample, not an arithmetic error — the gaps agree closely (E023: +6.9% / +3.5%; here: +6.67% / +2.43%), which is what one expects when the sample changes and the effect does not.

Absolute perplexity for a 0.5B model on wikitext-2 is a property of the chunk selection as much as of the model, so it should always have been quoted with the sample. It now is.

3. The token ids are committed. E023's were derived ad hoc, which is why none of this could be re-run. implementation/tokenise.py reproduces them exactly (verified against conformance/qwen_ids96.json).

The nuance worth keeping

The determinism gap holds its expected sign in only 83 of 120 chunks (69%), against 113/120 for quantisation. The mean gap is solid and its interval clean, but on any individual chunk the contract beats plain int8 nearly a third of the time.

That is what a small rounding difference looks like: not a systematic degradation, but a slight adverse shift in a distribution that is mostly noise. It supports the mechanism we proposed — more requantisation boundaries, so more rounding — and it argues against reading +2.43% as "the contract is worse on every input". It is not. It is worse on average, by a little.

What this does not settle

  • One model, one corpus. Qwen2.5-0.5B on wikitext-2. Nothing here says the gap is +2.43% for a 7B model, or on code, or on a different language.
  • One context length. 128 tokens throughout, for comparability with E023. Longer contexts stress attention quantisation harder and were not tested.
  • The bootstrap assumes chunks are exchangeable. They are disjoint and spread across the corpus, but wikitext is not i.i.d. — adjacent chunks share an article. Evenly spacing the draws reduces this; it does not eliminate it.
  • The int8 ceiling is our implementation of the standard W8A8 recipe, not a tuned production quantiser. A better ceiling would make the contract look worse, not better, so this is not a self-serving choice — but it is not the best int8 anyone could build either.