Paper readiness
What still has to be true before the preprint is submittable.
Status: honest assessment, 2026-09-02. Revised whenever a blocker clears.
The research charter says: "Never exaggerate the paper." This document exists so that the case for publishing is argued against the same standard as every result in the repository.
1. Is there a paper here?
Yes — a measurement and systems paper, not a cryptography paper. We have introduced no new primitive. Freivalds is from 1977, sumcheck from the early 1990s, and quantised inference is standard practice. A crypto venue would correctly reject this.
What we have is a set of measurements nobody appears to have made, and one named open problem answered.
2. What would be genuinely novel
Ordered by how well it would survive review.
C1 — Portable numeric determinism, demonstrated on a real model
Strongest single result in the programme. Qwen2.5-0.5B produces byte-identical logits on Apple ARM CPU, Apple Metal GPU, two Intel x86 CPUs on different OSes, and an NVIDIA L4 with TF32 enabled — five backends, two CPU vendors, two GPU vendors. EigenAI (arXiv:2602.00182, Jan 2026) achieves bit-exact inference by locking the network to identical H100s, and lists as future work:
"Cross-Architecture Reproducibility. Determinism currently holds only within fixed GPU families. Future work includes portable numeric normalization to enable heterogeneous verifier sets."
We answer that: exact integer semantics give bit-identical output across CPU and GPU, at +0.0001 accuracy on classification and +1.4% perplexity on generation. Answering a named open problem from a recent paper is the single strongest thing we have.
C2 — The cross-backend divergence rate for greedy decoding
9 of 10 honest CPU/GPU pairs produced identical 300-token outputs; 1 diverged at token 72. We have not found this measured anywhere, and it is the number that decides whether output-hash replication can be used for enforcement. It also contradicts our own prior, which strengthens rather than weakens it.
C3 — The determinism/provability tension, with the cost localised
Lookup-table non-linearities give cross-backend determinism and block the sumcheck. Making a transformer sumcheck-compatible costs +4.2% perplexity, and the cost is entirely in attention — replacing GELU with x² is free (−0.5%). This corrects the natural reading of SafetyNets, whose networks had no attention and whose only softmax was the excluded final layer.
C4 — The operand-magnitude trap, cross-vendor and predicted in advance
Apple's Metal matmul is exact on integer operands only to 2^11, regardless of accumulator width. From that we predicted NVIDIA's TF32 — an unrelated design with an 11-bit effective mantissa — would show the same bound, recorded it before the hardware was available, and measured exactly 2^11 (E022).
Two vendors, one number, with the prediction registered first. That is a stronger form of the finding than a single-platform gotcha, and it means the bound is a property of accelerator matmul rather than of one framework.
It also sharpens a threat-model point: TF32 is on by default in many frameworks because it is faster, so a provider "running float32" very often is not — threat T7 occurring with no intent to cheat.
C6 — Interactive proofs do not transfer to transformers, and why
SafetyNets' quadratic-activation restriction is usually read as a dated compromise. E020b suggests it is load-bearing: the polynomial attention a sumcheck requires costs +6.8% perplexity at 2 layers and +19.8% at 8, with softmax improving on depth while squared attention regresses. Their networks were CNNs and fully-connected nets with no attention to lose; a transformer's inductive bias lives in exactly the operation a sumcheck cannot pass through. A clean negative result for the interactive-proof line applied to modern architectures, which we have not found stated anywhere.
C5 — A cost decomposition with a re-execution baseline
Compute, bandwidth and accuracy measured on the same workload, always against "just run it again". The finding that bandwidth dominates by 3–5 orders of magnitude is stated qualitatively in one sentence of SafetyNets and, as far as we can tell, never priced. Our figure: 49 MB of intermediates against a 0.02 MB answer, break-even at $164/hour compute against a $1 GPU-hour.
3. What is not novel, and must be presented as such
- Freivalds' algorithm, the sumcheck protocol, quantised inference.
- That a per-layer interactive proof has lower bandwidth than transmitting intermediates — SafetyNets says this explicitly in 2017, and we rediscovered it by construction before reading them. That must be cited, not claimed.
- The overall architecture (determinism + audit + stake) is EigenAI's.
4. What blocks submission today
| # | blocker | severity | fix |
|---|---|---|---|
| — | |||
| close the residual accuracy gap | |||
| B3 | done | — | |
| — | |||
| — | |||
| B5 | Security claims are empirical, not proved. | acceptable, if framed | present as a measurement paper |
| — | done |
B1 is the one that matters. B1, B4 and B7 are all cleared. What
remains is B5, which is a framing decision rather than a measurement: present
this as a measurement study, which the paper already does.
Both B7 and B4 moved a number that looked settled, and in the same direction: the figure we had was a measurement of our own code rather than of the thing we meant to measure. That is now the default assumption for any number in this repository that has not been checked twice.
B7 is worth noting as the blocker we did not know we had: the accuracy numbers looked settled, and the check moved the headline by a full point and showed the original sample could not have supported it. The remaining unquantified figures should be assumed to have the same problem until measured.
5. Would it be "better" than SafetyNets or EigenAI?
No — and that is the wrong frame. Honestly:
- SafetyNets has better proof numbers than anything we have built. 5% prover overhead and under 8 KB of communication. We are not competitive on that axis and should not pretend to be.
- EigenAI has a deployed system with restaked capital behind it. We have a research repository.
What we would have is narrower and complementary: the portable-determinism result they named as open, plus a measurement study of what actually costs what. That is a real contribution to their line of work, not a replacement for it.
The honest positioning is "we measured the thing everyone assumes, and one assumption was wrong" — which is a good paper, and a much more defensible one than "we beat SafetyNets".
6. What would make it strong
In order:
Clear B1.Done, and TF32 did show the same trap — C4 is now a cross-vendor result with a pre-registered prediction behind it.- Clear B2/B3. One real model on one standard corpus.
Resolve B6.Done — and it did become the clean negative result (C6). Remaining: re-measure the depth-8 gap at convergence rather than at a fixed 1,500 steps, so under-training can be ruled out as a partial explanation.- Keep the corrections section. Seven documented retractions is unusual and, in a measurement paper, is a credibility asset rather than an embarrassment.