NodusLab
E018 · CP-016

The contract on a transformer, doing generation

does NC-0.2 survive a transformer, where softmax is on the critical path and a single flipped token compounds?

Checkpoint: CP-016 · Status: COMPLETE, run 2026-09-01 Question: does NC-0.2 survive a transformer, where softmax is on the critical path and a single flipped token compounds?

Answer: yes — and it delivers what v1 needs

Greedy decoding, 10 prompts × 300 generated tokens:

comparison runs identical
NC-0.2 contract, CPU vs GPU 10 / 10
float, CPU vs GPU (both honest) 9 / 10

One honest float pair in ten diverged. That is the number Nodus v1 has needed since it was built: its REPLICATED_MISMATCH path compares output hashes for exact equality on evidence of n = 1, against a simulated node. A 10% false-positive rate means output-hash replication cannot be used for enforcement — you would slash one honest provider in ten.

Under the contract the rate is zero, by construction rather than by luck.

Quality cost

path perplexity
float64 5.1708
standard quantisation (int8) 5.1639
NC-0.2 contract 5.2370

Extra cost of determinism over ordinary quantisation: +1.4% perplexity. Materially more than the ~0 E017 found on MNIST — because a classifier's argmax absorbs small perturbations and a language model's perplexity does not.

Two corrections to our own claims

1. Our decoding-divergence prediction was wrong. E009-pre predicted honest CPU/GPU greedy decoding would "diverge within tens of tokens". Measured: 9 of 10 prompts were identical over 300 tokens. Argmax has a margin, so divergence in the arithmetic does not automatically become divergence in the tokens. The conclusion survives — 90% agreement is still unusable — but we asserted a mechanism we had not measured.

2. NC-3 bounded the wrong quantity. The contract diverged at token 102, which should be impossible. MEASURED: MLX's Metal matmul is exact on integer operands up to 2^11 − 1 = 2047 and inexact from 2^12, regardless of accumulator size. NC-3's 2^24 accumulator bound was never the binding one on this backend. Every earlier experiment satisfied the real bound by accident, because int8 operands are ≤ 127; attention probabilities at 2^15 were the first to exceed it. Fixed by carrying probabilities at 2^10 and adding NC-3b, which matmul_exact now asserts rather than assumes.

A numeric contract is only as good as the quantity it bounds, and the binding constraint is backend-specific.

Files

file contents
implementation/model.py the float transformer, trained with MLX
implementation/paths.py all three inference paths; the contract is integer end to end
implementation/run.py perplexity, generation, cross-backend checks
implementation/divergence_dist.py the 10-prompt divergence distribution
analysis.md the result and both corrections

Raw records: ../../benchmarks/results/018-transformer-generation.jsonl, 018b-generation-divergence.jsonl

Caveat

2 layers, d=128, vocab 65. Deeper models compound more and larger vocabularies have tighter logit margins, so the 10% float divergence rate is likely an underestimate for real models.

Status: complete. Two corrections to our own prior claims, and one decisive result for Nodus v1. Date: 2026-09-01. Apple M5; MLX 0.32.2; NumPy 2.5.2. Model: 2-layer char transformer, d=128, 4 heads, ctx 64, vocab 65, tiny-shakespeare. Raw records: ../../benchmarks/results/018-transformer-generation.jsonl, 018b-generation-divergence.jsonl


1. The headline: replication works under the contract and not without it

Greedy decoding, 10 prompts × 300 generated tokens:

comparison runs identical verdict
NC-0.2 contract, CPU vs GPU 10 / 10 usable for enforcement
float, CPU vs GPU (both honest) 9 / 10 not usable
float vs contract 0 / 10 different model, as expected

One honest float pair in ten diverged, at generated token 72.

That single number is the one Nodus v1 has needed since it was built. v1's REPLICATED_MISMATCH path compares output hashes for exact equality and treats a mismatch as a fraud signal; its evidence base is n = 1, against a simulated node. A 10% false-positive rate between honest heterogeneous backends means output-hash replication cannot be turned into enforcement — you would slash one honest provider in ten.

Under the contract that rate is zero out of ten, and it is zero by construction rather than by luck.

2. Correction: our own prediction about decoding divergence was wrong

E009-pre measured a 7.7 × 10^-4 relative divergence between honest CPU and GPU on a single matmul, and predicted:

"honest CPU and GPU greedy decoding of the same prompt diverge within tens of tokens."

Measured: 9 of 10 prompts produced identical 300-token outputs. The prediction was wrong in direction.

The reason is instructive. Greedy decoding takes an argmax, and argmax has a margin: the leading logit is usually far enough ahead that a 10^-4 perturbation cannot flip it. Divergence in the arithmetic does not automatically become divergence in the tokens.

This matters beyond the correction, because it cuts both ways: replication looks far more reliable than E009-pre implied, and is still not reliable enough. A 90% agreement rate is much better than we predicted and still unusable as an enforcement threshold. Being wrong about the mechanism did not change the conclusion, but we should have measured it before asserting it.

3. Correction: NC-3's bound was on the wrong quantity

The contract initially diverged between CPU and GPU at token 102 — which should be impossible for exact integer arithmetic. It was our bug, and it exposed a real defect in the contract.

MEASURED: MLX's Metal matmul is exact on integer operands up to 2^11 − 1 = 2047, and inexact from 2^12, regardless of accumulator magnitude:

max operand 2^b max |C| < 2^24 exact
2,047 2^11 797,436 yes yes
4,095 2^12 1,676,075 yes no
32,767 2^15 15,834,624 yes no

NC-3 bounded the accumulator at 2^24. On this backend that bound was never the binding one — the operand mantissa is, and it is 13 bits tighter. Every earlier experiment satisfied it by accident, because int8 operands (≤127) are comfortably under 2047. Attention probabilities at 2^15 were the first operand to exceed it, and the contract broke silently.

Fix: carry attention probabilities at 2^10 instead of 2^15, and add an explicit operand bound to the contract (NC-3b) that matmul_exact now asserts rather than assumes. After the fix, 10/10 runs are bit-identical.

The general lesson is the one this experiment exists to teach: a numeric contract is only as good as the quantity it bounds, and the binding constraint is backend-specific. We had measured the accumulator limit and assumed it was the whole story.

4. Quality cost

Perplexity on held-out text (16 sequences):

path perplexity vs float
float64 5.1708 —
standard quantisation (int8) 5.1639 −0.0069
NC-0.2 contract 5.2370 +0.0662 (+1.28%)

The extra cost of determinism over ordinary quantisation is +0.073 perplexity, about +1.4%.

That is materially more than E017 found on MNIST, where it was zero. Both results are honest and the difference is the point: a classifier's argmax absorbs small perturbations; a language model's perplexity does not. Attention softmax is now on the critical path, and generation compounds.

Generated text remains in the same quality class — see the samples in the records — but it is different text: the contract diverges from float at a median of token 6. That is expected (it is a quantised model, not the same function) and is only a problem if a job specification promises float outputs rather than contract outputs.

5. What this means for the programme

The contract survives a transformer. Attention, softmax, residuals, layernorm and GELU all run in exact integers, and the result is bit-identical across CPU and GPU over 3,000 generated tokens. That was the open question from E015 and it is now closed affirmatively.

Its price is now known on two workload classes: ~0 on MNIST classification, +1.4% perplexity on character-level generation. The honest summary is "small but no longer zero", and it should be quoted per workload rather than as a single number.

And it delivers the thing v1 actually needs. Nodus's premise is heterogeneous consumer hardware. Under float, honest providers disagree 10% of the time and no enforcement threshold exists. Under the contract they never disagree, so replication becomes an enforcement mechanism rather than a signal.

6. Limitations

  • Small model — 2 layers, d=128, vocab 65. Deeper models compound more, and larger vocabularies have tighter logit margins, so the 10% float divergence rate is probably an underestimate for real models.
  • 10 prompts, 300 tokens, one seed, greedy decoding only. Sampling at temperature > 0 would need the RNG to be part of the contract too — untested.
  • One machine, two backends that share a vendor. Cross-vendor would likely be worse for float and unchanged for the contract.
  • Perplexity is measured on 16 sequences; the ±0.07 difference deserves more sequences and more seeds before being quoted precisely.
  • The operand bound of 2^11 is measured for MLX on Metal only. Other backends will have different limits, and the contract must state the bound it requires rather than inheriting whatever a backend happens to provide.

1. Give v1 the finding it can act on [do first, it is small]

v1's REPLICATED_MISMATCH path treats an output-hash mismatch as a fraud signal. We now have the false-positive rate between honest heterogeneous backends: ~10% over 300 tokens on a small model, and probably worse on a large one.

Two concrete changes, neither of which is research:

  • document that exact-match replication cannot be used for enforcement on float inference, with this number attached;
  • make replication comparison contract-aware: under NC-0.2 an exact mismatch is real evidence, under float it is not.

2. Establish the operand bound properly, per backend

NC-3b currently carries one measured number (2^11 for MLX on Metal). A contract cannot inherit whatever a backend happens to provide — it must state the bound it requires, and providers must demonstrate conformance (NC-8).

Needed: the same measurement for NVIDIA cuBLAS with TF32 on and off, and for a non-Apple CPU BLAS. Blocked on hardware, and it is now the most consequential thing that hardware access would unblock.

3. Scale the model

Everything here is 2 layers and d=128. The two numbers most likely to move:

  • the float divergence rate should rise with depth and vocabulary size;
  • the perplexity cost of the contract may rise or fall — deeper models have more rounding but also more redundancy.

A 6–12 layer model on the same corpus is cheap and would show the trend.

4. Sampling, not just greedy

All of this is temperature 0. Sampling makes the RNG part of the computation, so the contract must specify the generator and its seed, or two honest providers will differ by construction. Untested and straightforward to add.

5. Firm up the perplexity number

16 sequences and one seed give ±0.07 with no error bar. More sequences and 3 seeds before the +1.4% figure is quoted anywhere load-bearing.