the contract on Qwen2.5-0.5B
Analysis
analysis.md ↗Status: complete. Determinism holds on a real model. Accuracy is close but
not at parity, and the residual gap may be intrinsic.
Date: 2026-09-02. Qwen2.5-0.5B — 24 layers, hidden 896, GQA (14 query heads
over 2 KV), RMSNorm, SwiGLU, RoPE θ=1e6, vocab 151,936. wikitext-2 test.
Raw records: ../../benchmarks/results/023-real-model.jsonl
1. The headline: determinism holds at real scale, across vendors
| machine | stack | logits hash |
|---|---|---|
| Apple M5, CPU | ARM · Accelerate · macOS 26.2 · Py 3.13.9 · NumPy 2.5.2 | 202f51a312e5c0e685354355 |
| Apple M5, GPU | Metal · MLX 0.32.2 | 202f51a312e5c0e685354355 |
| Intel Xeon | x86-64 · OpenBLAS · Debian 13 · Py 3.13.5 · NumPy 2.5.2 | 202f51a312e5c0e685354355 |
| Intel Xeon | x86-64 · OpenBLAS · Ubuntu 22.04 · Py 3.10.12 · NumPy 2.2.6 | 202f51a312e5c0e685354355 |
| NVIDIA L4 | CUDA 12.4 · torch 2.6 · TF32 ENABLED | 202f51a312e5c0e685354355 |
Bit-identical on all five, with every next-token prediction matching — across two CPU vendors and two GPU vendors, three operating systems, three Python versions and three NumPy versions.
Note the last row. TF32 was deliberately left enabled on the NVIDIA run. E022 measured TF32's operand bound at 2^11; int8 operands are ≤ 127, four bits inside it, so the reduced-precision tensor-core path cannot lose anything the contract depends on. The contract does not require anyone to disable TF32 — NC-3b makes it irrelevant. That is the clause working as designed rather than by luck, and it is the difference between a contract a provider can actually comply with and one that demands they cripple their hardware.
Previous determinism results were 2-layer character models and random matrices. This is the first on something the network actually serves, and the first at real scale across silicon vendors.
The conformance kit used for this (../../conformance/check_qwen.py) is
numpy-only and has the token ids hardcoded, so the input cannot differ
between machines and no tokenizer or model-loading library sits in the trust
path. It parses safetensors directly, so even the bfloat16 widening is an
explicit 16-bit shift rather than a library call that might differ by platform.
2. The defect only a real model could reveal
Measured with float accumulation — no integer arithmetic involved, so this is about quantisation granularity alone:
| configuration | perplexity |
|---|---|
| float64 reference | 16.77 |
| int8, per-tensor scales — what NC-5 specified | 70.39 |
| int8, per-channel / per-token scales | 17.36 |
Per-tensor int8 destroys the model — 4× worse. NC-5 was written from a 2-layer character model where it was harmless. A 24-layer model with a 151,936-token vocabulary has dynamic range a toy does not.
The contract was wrong, and only scale could have shown it. The fix costs nothing in principle: per-channel scales are equally specifiable and equally deterministic. It is more clauses, not weaker ones.
3. Six implementation lessons, each worth a contract clause
Getting the integer path from unusable to close took six rounds. Each was a real defect, and each is the kind of thing an implementer will otherwise rediscover:
| # | fix | perplexity vs float |
|---|---|---|
| — | starting point (per-tensor, int8 intermediates, truncating) | +47.6% |
| 1 | per-channel weights, per-token activations | — |
| 2 | carry intermediates at the 2^11 operand budget, not int8 | +21.0% |
| 3 | round on requantisation instead of truncating | +15.4% |
| 4 | RMSNorm gamma at 16 bits instead of 7; rounding division | +15.0% |
| 5 | stop requantising twice in the MLP | +11.8% |
| 6 | LUT outputs also get the operand budget | +10.3% |
Three of those are the same mistake in different places: int8 is a storage convention, not the contract's limit. NC-3b permits 2^11, and every place we capped at 127 — activations, LUT outputs — was discarding four bits for free.
And one rule that cost three separate bugs to learn:
A scale may not vary along an axis that is contracted over. Per-channel weight scales do not factor through
q @ kᵀ; a per-token scale onvdoes not factor throughattn @ v. Both must be folded into the integers before the contraction, by an integer multiply-and-shift derived from committed weights.
And one that is embarrassing but instructive: >> floors. Every
requantisation was biased downward by ~0.5 LSB — a systematic error, so it
compounds through depth instead of averaging out. NC-5 specifies
round-to-nearest, so our implementation was violating our own contract. Fixing
it was worth 8 perplexity points.
4. Where accuracy stands — corrected
An earlier version of this analysis reported an "unexplained ~7 point gap". That was a measurement error on our side: the +3.5% ceiling figure came from a 256-token context and the contract figure from a 128-token one. Comparing across configurations is exactly the mistake this repository's house rules exist to prevent, and it produced a phantom problem that several hours went into chasing.
Measured like for like — same context, same chunks, same run:
| path | perplexity | vs float |
|---|---|---|
| float64 reference | 23.270 | — |
| per-channel int8, float accumulation (the achievable ceiling) | 24.883 | +6.9% |
| NC-0.5 contract, integer end to end | 25.762 | +10.7% |
The contract costs +3.5% over the best achievable int8 quantisation. Quantisation itself costs +6.9%; full integer determinism adds +3.5% on top.
⚠ SUPERSEDED by E025
The numbers above are point estimates from 4–6 chunks with no stated uncertainty, and resampling at that size gives a 95% CI of [−1.36%, +6.47%] — which contains zero. The evidence here did not establish the conclusion drawn from it.
Re-measured over 120 paired chunks: quantisation +6.67% [5.95, 7.40], determinism +2.43% [1.63, 3.23]. The determinism cost was overstated; the contract is cheaper than this page claims.
The absolute perplexities do not reproduce either (30.089 / 32.096 / 32.876 across the full test set) because these chunks came from the front of the corpus. That is a change of sample, not an arithmetic error — the gaps agree. Absolute perplexity here should always have been quoted with its sample.
Kept unedited rather than corrected in place, because the point of a research log is that it records what was believed when.
That +3.5% is the honest price of determinism at this scale, and it is plausibly close to intrinsic. A float-accumulation pipeline quantises matmul inputs and keeps accumulators exact between steps; a fully-integer pipeline cannot, because there is no wide float to fall back on. More requantisation boundaries mean more rounding, and that is a consequence of the requirement rather than a defect.
What was ruled out while chasing the phantom
Each of these was exactified in turn and measured; none moved the number:
| component | effect |
|---|---|
| softmax LUT and normalisation | none (33.2 → 33.5, noise) |
| RMSNorm | none (33.2 → 34.6, noise) |
| RoPE fixed-point precision | none (+10.7% → +10.6%) |
| per-channel scale folding | none (+10.7% → +11.9%) |
| softmax score grid, swept 1/256 → 1/16384 | none (+9.4% to +11.2% across a 64× range) |
| residual grid finer than 2^20 | overflows — (x·x).sum() exceeds int64 |
The pre-softmax scores correlate with the float path at 0.9993, so the arithmetic was never in question. The remaining cost is distributed across the requantisation boundaries rather than concentrated in any one operation.
5. What this changes
- NC-5 is amended from per-tensor to per-channel/per-token, with the contracted-axis rule stated explicitly.
- NC-3b guidance added: use the operand budget, do not default to int8.
- NC-5 rounding is now enforced in the reference implementation rather than merely specified.
- The programme's determinism claim now holds on a real model, closing the most serious form of the "toy models" objection.
6. Limitations
- 4–6 chunks of 128 tokens. Both the float64 reference and the integer path are numpy, and slow; the sample is small and the perplexity figures carry no error bars.
- Cross-backend check is Apple CPU vs Apple GPU. The Intel and NVIDIA machines from E021/E022 were released before this experiment ran, so the real model has not been checked against them. The conformance kit would do it in minutes.
- The residual ~7 points is unexplained and may be intrinsic (§4).
- No generation test at this scale — perplexity only.