NodusLab
E020 · CP-018

What a provable transformer costs in quality

E019 showed our lookup tables block the sumcheck. If we change the model to match the protocol instead, what does that cost?

Checkpoint: CP-018 · Status: COMPLETE, run 2026-09-02 Question: E019 showed our lookup tables block the sumcheck. If we change the model to match the protocol instead, what does that cost?

Answer: nothing for activations, +4.2% for attention

variant FFN attention perplexity vs baseline sumcheck-ready
A GELU softmax 4.9939 — no
B x² softmax 4.9703 −0.5% no
C x² squared 5.2044 +4.2% yes

Replacing GELU with x² is free — variant B is marginally better than the baseline, inside seed noise. SafetyNets' activation restriction, the thing that looks like the painful compromise, costs nothing here.

Replacing softmax is what costs, and it is the change that actually matters.

E019 under-stated the problem, and this shows why

SafetyNets' networks had no attention — their only softmax was the final classification layer, which the paper explicitly excludes from the proof. A transformer cannot make that exclusion: softmax sits mid-chain in every attention block, and a sumcheck stops there exactly as it stops at a lookup table.

So variant B — the change everyone quotes — is not sumcheck-compatible. Only C is, and it required replacing attention itself with masked(s²) / rowsum.

The division is fine for a sumcheck: the prover supplies the quotient as a witness and the verifier checks quotient × rowsum == numerator, which is arithmetic. It is softmax's exponential that is unreachable.

The trade, measured

quality cost communication per layer
LUT contract (NC-0.3) +1.4% perplexity 49 MB
Polynomial model +4.2% perplexity ~800 bytes

Roughly three extra points of perplexity buys a ~60,000× bandwidth reduction — which on E016's cost model is the difference between a scheme that loses money and one that is free.

The caveat

A 2-layer character model on Shakespeare flatters squared attention. Softmax's advantage generally widens with scale, task difficulty and context length, so +4.2% is likely a floor rather than an estimate. Variant C's seed variance is 5× the baseline's and training needed gradient clipping — both signs the optimisation is less well-behaved, and both things SafetyNets also reported.

Files

file contents
implementation/polytrain.py all three variants, trained identically
analysis.md the result, the two architectures it leaves us with

Raw records: ../../benchmarks/results/020-polynomial-model.jsonl

Status: complete. Date: 2026-09-02. Apple M5; MLX 0.32.2. Model: 2-layer char transformer, d=128, 4 heads, ctx 64, tiny-shakespeare, 3,000 steps, 2 seeds per variant. Raw records: ../../benchmarks/results/020-polynomial-model.jsonl


1. The result

variant FFN attention perplexity vs baseline sumcheck-compatible
A GELU softmax 4.9939 ± 0.02 — no
B x² softmax 4.9703 ± 0.01 −0.5% no
C x² squared 5.2044 ± 0.10 +4.2% yes

Two things fall out, and they point in opposite directions.

Replacing GELU with x² is free. Variant B is marginally better than the baseline, comfortably inside seed noise. SafetyNets' famous activation restriction — the thing that looks like the painful compromise — costs nothing on this workload.

Replacing softmax is what costs. Variant C pays +4.2% perplexity, and it is the change that actually matters, because attention is on the critical path of every transformer and softmax is not polynomial.

2. Why B is not enough, and E019 under-stated the problem

E019 framed option (a) as "adopt SafetyNets' quadratic activations". That framing was incomplete, and this experiment shows why.

SafetyNets' networks had no attention. They were CNNs and fully-connected nets; their only softmax was the final classification layer, which the paper explicitly excludes from the proof as "not amenable to an interactive proof". For a classifier that exclusion is harmless — the argmax is taken after it.

A transformer cannot make that exclusion. Softmax sits inside every attention block, mid-chain, and a sumcheck stops there exactly as it stops at a lookup table. Variant B is not sumcheck-compatible, despite having the activation change everyone quotes.

Only variant C is, and getting there required replacing attention itself:

softmax(s)  ->  masked(s²) / rowsum(masked(s²))

The division is not a problem for a sumcheck even though it is not a polynomial: the prover supplies the quotient as a witness and the verifier checks the multiplication quotient × rowsum == numerator, which is arithmetic. It is the exponential in softmax that is unreachable, not the division.

3. What the trade actually is

Two architectures that both give a verifiable transformer:

quality cost communication per layer
LUT contract (NC-0.3, E015/E018) +1.4% perplexity 49 MB
Polynomial model (variant C) +4.2% perplexity (plus the contract's own ~1.4% if also quantised) ~800 bytes

Roughly three extra points of perplexity buys a ~60,000× reduction in bandwidth.

On E016's cost model that is decisive on the money: 49 MB at cloud egress is $4.4 × 10⁻³ per layer against a compute saving worth $2.7 × 10⁻⁵ — a loss of two orders of magnitude — while 800 bytes is free. The polynomial model turns a scheme that loses money into one that does not.

Whether ~3 points of perplexity is acceptable is a product decision, not a research one, and it is the operator's call.

4. The caveat that could reverse this

This is a 2-layer character model on Shakespeare, and squared attention is being flattered by that. Softmax's advantage over polynomial and linear attention variants is generally reported to widen with model scale, task difficulty and context length — the sharpness of softmax is what lets attention be selective over long contexts, and a squared kernel is a blunter instrument.

So +4.2% is very likely a floor, not an estimate. Variant C's seed variance is also 5× the baseline's (±0.10 against ±0.02), and training needed gradient clipping to stay finite — both consistent with SafetyNets reporting the same need, and both signs that the optimisation is less well-behaved.

Anyone acting on this number should re-measure it at the scale they intend to run, and should expect it to get worse.

MEASURED IN E020b, AND IT DOES. At context 128 the gap is +6.8% at 2 layers and +19.8% at 8, with softmax improving on depth while squared attention regresses. The floor was not close to the ceiling. See analysis-scaling.md — this closes the polynomial architecture.

5. What this changes

The programme now has two coherent architectures rather than one, with the trade-off measured rather than assumed:

  • Keep the model, pay the bandwidth. LUT contract + one-random-layer audit + a stake equal to the job's compute value (E016). Works today, on any transformer, at 39% of the cost of re-execution for a 32-layer model.
  • Change the model, pay the quality. Polynomial transformer + sumcheck. ~800-byte transcripts, no audit, no stake needed for correctness — but a measurably worse model, by an amount that is probably worse at scale.

The useful structural finding is where the cost sits. It is not in the activations, which are free to change. It is entirely in attention. That is where any future work on provable transformers has to aim.

6. Limitations

  • One small model, one corpus, two seeds, 3,000 steps. Perplexity differences of a few percent on a 2-layer model are suggestive, not settled.
  • Variants B and C were trained from scratch, not distilled from A. Distillation or longer training might close the gap; SafetyNets found quadratic networks converged faster, which we also see in B.
  • No quantised variant of C, so the +4.2% and the contract's +1.4% have not been measured composed — they are assumed roughly additive, which is untested.
  • Squared attention is one of several polynomial alternatives. Linear attention with feature maps, or a higher-degree polynomial kernel, may do better and were not tried.

Analysis — scaling

analysis-scaling.md ↗

Status: complete. Result: negative, and it settles the architecture. Date: 2026-09-02. Apple M5; MLX 0.32.2. Context 128, 1,500 steps, 2 seeds. Raw records: ../../benchmarks/results/020b-attention-scaling.jsonl


1. The result

layers softmax squared gap
2 5.6240 ± 0.00 6.0060 ± 0.10 +6.8%
4 5.3730 ± 0.04 5.7282 ± 0.24 +6.6%
8 5.2468 ± 0.01 6.2839 ± 0.04 +19.8%

E020 measured +4.2% at 2 layers with a 64-token context. Lengthening the context to 128 took that to +6.8% before depth entered at all. At 8 layers it is +19.8%.

The gap grows on both axes we can vary. The caveat E020 flagged as "likely a floor" was correct, and the floor is not close to the ceiling.

2. The mechanism is visible in the columns

Read the two columns separately, because they behave differently:

  softmax:   5.6240  ->  5.3730  ->  5.2468     improves with depth
  squared:   6.0060  ->  5.7282  ->  6.2839     stops improving, then regresses

Softmax attention gets better with depth. Squared attention stops benefiting from it and then goes backwards. The gap is not a fixed tax that a larger model amortises — it is a divergence.

That is consistent with the usual explanation: softmax's sharpness is what lets attention be selective, and selectivity is what deep stacks compound. A squared kernel spreads its mass, so each additional layer has less to work with.

3. What this settles

The polynomial architecture is dead for transformers.

Nobody trades ~20% model quality for a bandwidth saving, and the trend says 20% is what you pay at 8 layers with the gap still widening. E019's option (a) — change the model to fit the proof — is closed on evidence.

The programme's architecture question therefore has an answer:

Keep the model, pay the bandwidth. NC-0.3's lookup-table contract, plus one-random-layer audit, plus a stake equal to the job's compute value (E016).

That design works today on an unmodified transformer, costs +1.4% perplexity, and runs at 39% of the cost of re-execution for a 32-layer model.

4. And it says something about SafetyNets

SafetyNets' restriction to quadratic activations and sum pooling is often read as a dated compromise that better engineering would remove. These measurements suggest something stronger:

The restriction is not incidental to SafetyNets — it is why the approach works, and it is why the approach does not transfer to transformers.

Their networks were CNNs and fully-connected nets, which have no attention to lose. A transformer's entire inductive bias lives in softmax, and softmax is exactly what a sumcheck cannot pass through. The deeper the model, the more that costs.

This is a genuinely useful negative result for the interactive-proof line of work applied to modern architectures, and we have not found it stated anywhere.

5. The alternative explanation, which we cannot yet rule out

Under-training. All runs use a fixed 1,500 steps. Deeper models generally need more, and squared attention needed gradient clipping to stay finite at all, so it may need more still. Some of the +19.8% could be an optimisation artefact rather than a capacity limit.

Two things argue against that being the whole story: the softmax column improves monotonically with depth on the same budget, and the squared column regresses between 4 and 8 layers rather than merely lagging — a model that is simply under-trained usually still improves with capacity.

What would settle it: train to convergence, or match on steps-to-a-fixed-loss rather than a fixed step count. Until then the direction is solid and the magnitude at depth 8 carries a caveat.

6. Limitations

  • d=128, character-level, one corpus, 2 seeds, 1,500 steps.
  • Squared attention is the crudest polynomial kernel. Linear attention with feature maps or a higher-degree kernel might scale better; neither was tried, and either could reopen the question.
  • The depth-4 squared runs have ±0.24 variance, five times the softmax column's. The 4-layer point is the weakest in the table.

1. Re-measure at scale before anyone acts on +4.2%

The single most likely way this result misleads is by being measured on a 2-layer character model. Softmax's advantage over polynomial attention widens with depth, context length and task difficulty.

Cheapest useful version: the same three variants at 6–12 layers with a longer context, on the same corpus. If the gap grows from 4% to 15%, the polynomial architecture is dead and the audit-plus-stake design wins by default. This is the measurement that decides between the programme's two remaining architectures, and it is a few hours.

2. Try the other polynomial attentions

Squared attention is the crudest sumcheck-compatible kernel. Linear attention with feature maps, and higher-degree polynomial kernels, are both established and may cost far less than 4.2%. Cheap to add — one function each — and the result could move the whole trade.

3. Compose the two costs

The +4.2% (polynomial) and +1.4% (integer contract) have never been measured together. A verifiable production model needs both, and they are assumed roughly additive with no evidence. Train variant C under the NC-0.3 contract and measure the composed figure.

4. Measure prover overhead for the sumcheck

Still outstanding from E019. SafetyNets reports 5%; we have measured communication and correctness but not the prover's cost. The architecture comparison should not be finalised without it.

5. Then decide, and write the decision down

Two coherent architectures now exist with measured trade-offs:

  • keep the model, pay the bandwidth — LUT contract + layer audit + stake
  • change the model, pay the quality — polynomial transformer + sumcheck

After items 1–4 the choice is a product decision about whether a few points of perplexity are worth removing an audit-and-stake mechanism from the network. It should be recorded in docs/architecture.md as a decision with its evidence, not left implicit.