NodusLab
E023

the contract on Qwen2.5-0.5B

CompleteDeterminism holds on a real model. Accuracy is close butexperiments/023-real-model/ ↗

Status: complete. Determinism holds on a real model. Accuracy is close but not at parity, and the residual gap may be intrinsic. Date: 2026-09-02. Qwen2.5-0.5B — 24 layers, hidden 896, GQA (14 query heads over 2 KV), RMSNorm, SwiGLU, RoPE θ=1e6, vocab 151,936. wikitext-2 test. Raw records: ../../benchmarks/results/023-real-model.jsonl


1. The headline: determinism holds at real scale, across vendors

machine stack logits hash
Apple M5, CPU ARM · Accelerate · macOS 26.2 · Py 3.13.9 · NumPy 2.5.2 202f51a312e5c0e685354355
Apple M5, GPU Metal · MLX 0.32.2 202f51a312e5c0e685354355
Intel Xeon x86-64 · OpenBLAS · Debian 13 · Py 3.13.5 · NumPy 2.5.2 202f51a312e5c0e685354355
Intel Xeon x86-64 · OpenBLAS · Ubuntu 22.04 · Py 3.10.12 · NumPy 2.2.6 202f51a312e5c0e685354355
NVIDIA L4 CUDA 12.4 · torch 2.6 · TF32 ENABLED 202f51a312e5c0e685354355

Bit-identical on all five, with every next-token prediction matching — across two CPU vendors and two GPU vendors, three operating systems, three Python versions and three NumPy versions.

Note the last row. TF32 was deliberately left enabled on the NVIDIA run. E022 measured TF32's operand bound at 2^11; int8 operands are ≤ 127, four bits inside it, so the reduced-precision tensor-core path cannot lose anything the contract depends on. The contract does not require anyone to disable TF32 — NC-3b makes it irrelevant. That is the clause working as designed rather than by luck, and it is the difference between a contract a provider can actually comply with and one that demands they cripple their hardware.

Previous determinism results were 2-layer character models and random matrices. This is the first on something the network actually serves, and the first at real scale across silicon vendors.

The conformance kit used for this (../../conformance/check_qwen.py) is numpy-only and has the token ids hardcoded, so the input cannot differ between machines and no tokenizer or model-loading library sits in the trust path. It parses safetensors directly, so even the bfloat16 widening is an explicit 16-bit shift rather than a library call that might differ by platform.

2. The defect only a real model could reveal

Measured with float accumulation — no integer arithmetic involved, so this is about quantisation granularity alone:

configuration perplexity
float64 reference 16.77
int8, per-tensor scales — what NC-5 specified 70.39
int8, per-channel / per-token scales 17.36

Per-tensor int8 destroys the model — 4× worse. NC-5 was written from a 2-layer character model where it was harmless. A 24-layer model with a 151,936-token vocabulary has dynamic range a toy does not.

The contract was wrong, and only scale could have shown it. The fix costs nothing in principle: per-channel scales are equally specifiable and equally deterministic. It is more clauses, not weaker ones.

3. Six implementation lessons, each worth a contract clause

Getting the integer path from unusable to close took six rounds. Each was a real defect, and each is the kind of thing an implementer will otherwise rediscover:

# fix perplexity vs float
— starting point (per-tensor, int8 intermediates, truncating) +47.6%
1 per-channel weights, per-token activations —
2 carry intermediates at the 2^11 operand budget, not int8 +21.0%
3 round on requantisation instead of truncating +15.4%
4 RMSNorm gamma at 16 bits instead of 7; rounding division +15.0%
5 stop requantising twice in the MLP +11.8%
6 LUT outputs also get the operand budget +10.3%

Three of those are the same mistake in different places: int8 is a storage convention, not the contract's limit. NC-3b permits 2^11, and every place we capped at 127 — activations, LUT outputs — was discarding four bits for free.

And one rule that cost three separate bugs to learn:

A scale may not vary along an axis that is contracted over. Per-channel weight scales do not factor through q @ kᵀ; a per-token scale on v does not factor through attn @ v. Both must be folded into the integers before the contraction, by an integer multiply-and-shift derived from committed weights.

And one that is embarrassing but instructive: >> floors. Every requantisation was biased downward by ~0.5 LSB — a systematic error, so it compounds through depth instead of averaging out. NC-5 specifies round-to-nearest, so our implementation was violating our own contract. Fixing it was worth 8 perplexity points.

4. Where accuracy stands — corrected

An earlier version of this analysis reported an "unexplained ~7 point gap". That was a measurement error on our side: the +3.5% ceiling figure came from a 256-token context and the contract figure from a 128-token one. Comparing across configurations is exactly the mistake this repository's house rules exist to prevent, and it produced a phantom problem that several hours went into chasing.

Measured like for like — same context, same chunks, same run:

path perplexity vs float
float64 reference 23.270 —
per-channel int8, float accumulation (the achievable ceiling) 24.883 +6.9%
NC-0.5 contract, integer end to end 25.762 +10.7%

The contract costs +3.5% over the best achievable int8 quantisation. Quantisation itself costs +6.9%; full integer determinism adds +3.5% on top.

⚠ SUPERSEDED by E025

The numbers above are point estimates from 4–6 chunks with no stated uncertainty, and resampling at that size gives a 95% CI of [−1.36%, +6.47%] — which contains zero. The evidence here did not establish the conclusion drawn from it.

Re-measured over 120 paired chunks: quantisation +6.67% [5.95, 7.40], determinism +2.43% [1.63, 3.23]. The determinism cost was overstated; the contract is cheaper than this page claims.

The absolute perplexities do not reproduce either (30.089 / 32.096 / 32.876 across the full test set) because these chunks came from the front of the corpus. That is a change of sample, not an arithmetic error — the gaps agree. Absolute perplexity here should always have been quoted with its sample.

Kept unedited rather than corrected in place, because the point of a research log is that it records what was believed when.

That +3.5% is the honest price of determinism at this scale, and it is plausibly close to intrinsic. A float-accumulation pipeline quantises matmul inputs and keeps accumulators exact between steps; a fully-integer pipeline cannot, because there is no wide float to fall back on. More requantisation boundaries mean more rounding, and that is a consequence of the requirement rather than a defect.

What was ruled out while chasing the phantom

Each of these was exactified in turn and measured; none moved the number:

component effect
softmax LUT and normalisation none (33.2 → 33.5, noise)
RMSNorm none (33.2 → 34.6, noise)
RoPE fixed-point precision none (+10.7% → +10.6%)
per-channel scale folding none (+10.7% → +11.9%)
softmax score grid, swept 1/256 → 1/16384 none (+9.4% to +11.2% across a 64× range)
residual grid finer than 2^20 overflows — (x·x).sum() exceeds int64

The pre-softmax scores correlate with the float path at 0.9993, so the arithmetic was never in question. The remaining cost is distributed across the requantisation boundaries rather than concentrated in any one operation.

5. What this changes

  • NC-5 is amended from per-tensor to per-channel/per-token, with the contracted-axis rule stated explicitly.
  • NC-3b guidance added: use the operand budget, do not default to int8.
  • NC-5 rounding is now enforced in the reference implementation rather than merely specified.
  • The programme's determinism claim now holds on a real model, closing the most serious form of the "toy models" objection.

6. Limitations

  • 4–6 chunks of 128 tokens. Both the float64 reference and the integer path are numpy, and slow; the sample is small and the perplexity figures carry no error bars.
  • Cross-backend check is Apple CPU vs Apple GPU. The Intel and NVIDIA machines from E021/E022 were released before this experiment ran, so the real model has not been checked against them. The conformance kit would do it in minutes.
  • The residual ~7 points is unexplained and may be intrinsic (§4).
  • No generation test at this scale — perplexity only.