# Paper readiness

**Status:** honest assessment, 2026-09-02. Revised whenever a blocker clears.

The research charter says: *"Never exaggerate the paper."* This document exists
so that the case for publishing is argued against the same standard as every
result in the repository.

---

## 1. Is there a paper here?

**Yes — a measurement and systems paper, not a cryptography paper.** We have
introduced no new primitive. Freivalds is from 1977, sumcheck from the early
1990s, and quantised inference is standard practice. A crypto venue would
correctly reject this.

What we have is a set of **measurements nobody appears to have made**, and one
**named open problem answered**.

## 2. What would be genuinely novel

Ordered by how well it would survive review.

### C1 — Portable numeric determinism, demonstrated on a real model
**Strongest single result in the programme.** Qwen2.5-0.5B produces byte-identical
logits on Apple ARM CPU, Apple Metal GPU, two Intel x86 CPUs on different OSes,
and an **NVIDIA L4 with TF32 enabled** — five backends, two CPU vendors, two GPU
vendors.
EigenAI (arXiv:2602.00182, Jan 2026) achieves bit-exact inference by locking the
network to identical H100s, and lists as future work:

> *"Cross-Architecture Reproducibility. Determinism currently holds only within
> fixed GPU families. Future work includes portable numeric normalization to
> enable heterogeneous verifier sets."*

We answer that: exact integer semantics give bit-identical output across CPU and
GPU, at **+0.0001 accuracy on classification** and **+1.4% perplexity on
generation**. Answering a named open problem from a recent paper is the single
strongest thing we have.

### C2 — The cross-backend divergence rate for greedy decoding
**9 of 10 honest CPU/GPU pairs produced identical 300-token outputs; 1 diverged
at token 72.** We have not found this measured anywhere, and it is the number
that decides whether output-hash replication can be used for enforcement.
It also *contradicts our own prior*, which strengthens rather than weakens it.

### C3 — The determinism/provability tension, with the cost localised
Lookup-table non-linearities give cross-backend determinism *and* block the
sumcheck. Making a transformer sumcheck-compatible costs **+4.2% perplexity, and
the cost is entirely in attention** — replacing GELU with x² is free (−0.5%).
This corrects the natural reading of SafetyNets, whose networks had no attention
and whose only softmax was the excluded final layer.

### C4 — The operand-magnitude trap, cross-vendor and predicted in advance
Apple's Metal matmul is exact on integer operands only to **2^11**, regardless
of accumulator width. From that we *predicted* NVIDIA's TF32 — an unrelated
design with an 11-bit effective mantissa — would show the same bound, recorded
it before the hardware was available, and **measured exactly 2^11** (E022).

Two vendors, one number, with the prediction registered first. That is a
stronger form of the finding than a single-platform gotcha, and it means the
bound is a property of accelerator matmul rather than of one framework.

It also sharpens a threat-model point: **TF32 is on by default** in many
frameworks because it is faster, so a provider "running float32" very often is
not — threat T7 occurring with no intent to cheat.

### C6 — Interactive proofs do not transfer to transformers, and why
SafetyNets' quadratic-activation restriction is usually read as a dated
compromise. E020b suggests it is load-bearing: the polynomial attention a
sumcheck requires costs **+6.8% perplexity at 2 layers and +19.8% at 8**, with
softmax *improving* on depth while squared attention regresses. Their networks
were CNNs and fully-connected nets with no attention to lose; a transformer's
inductive bias lives in exactly the operation a sumcheck cannot pass through.
A clean negative result for the interactive-proof line applied to modern
architectures, which we have not found stated anywhere.

### C5 — A cost decomposition with a re-execution baseline
Compute, bandwidth and accuracy measured on the same workload, always against
"just run it again". The finding that **bandwidth dominates by 3–5 orders of
magnitude** is stated qualitatively in one sentence of SafetyNets and, as far as
we can tell, never priced. Our figure: 49 MB of intermediates against a 0.02 MB
answer, break-even at $164/hour compute against a $1 GPU-hour.

## 3. What is not novel, and must be presented as such

- Freivalds' algorithm, the sumcheck protocol, quantised inference.
- That a per-layer interactive proof has lower bandwidth than transmitting
  intermediates — **SafetyNets says this explicitly in 2017**, and we
  rediscovered it by construction before reading them. That must be cited, not
  claimed.
- The overall architecture (determinism + audit + stake) is EigenAI's.

## 4. What blocks submission today

| # | blocker | severity | fix |
|---|---|---|---|
| ~~**B1**~~ | ~~Every measurement comes from one Apple M5.~~ **CLEARED (E021, E022):** bit-identical across Apple ARM CPU, Intel x86 CPU, Apple Metal GPU, and NVIDIA L4 with TF32 on *and* off. | ~~fatal~~ **done** | — |
| ~~**B2**~~ | ~~Scale.~~ **CLEARED (E023):** Qwen2.5-0.5B, 24 layers, vocab 151,936, wikitext-2 — bit-identical across **five** backends from two CPU and two GPU vendors, including NVIDIA with TF32 enabled. Accuracy is +10.3%, not at parity. | ~~severe~~ minor | close the residual accuracy gap |
| **B3** | ~~No standard benchmark.~~ **Cleared (E023):** wikitext-2 test, the standard perplexity benchmark. | done | — |
| ~~**B4**~~ | ~~Sumcheck prover overhead unmeasured.~~ **CLEARED (E026):** 10.4% at n=4096, falling as n^-0.92, 5% by n~8,800. The sumcheck rounds are **0.5%** of it; the rest is extensions and field entry. Our first measurement said 36.1% and was measuring our own int64 fallback. | ~~moderate~~ **done** | — |
| ~~**B7**~~ | ~~Accuracy figures are point estimates with no uncertainty.~~ **CLEARED (E025):** 120 paired chunks, bootstrap intervals. Determinism costs **+2.43% [1.63, 3.23]** — and the old +3.5% was outside that interval, from a sample whose own CI contained zero. | ~~severe~~ **done** | — |
| **B5** | Security claims are empirical, not proved. | acceptable, if framed | present as a measurement paper |
| ~~B6~~ | ~~The 4.2% attention figure may not survive scale.~~ **RESOLVED (E020b): it does not.** +19.8% at 8 layers. Became contribution C6 rather than a blocker. | — | done |

~~**B1 is the one that matters.**~~ **B1, B4 and B7 are all cleared.** What
remains is B5, which is a framing decision rather than a measurement: present
this as a measurement study, which the paper already does.

Both B7 and B4 moved a number that looked settled, and in the same direction:
the figure we had was a measurement of our own code rather than of the thing we
meant to measure. That is now the default assumption for any number in this
repository that has not been checked twice.

B7 is worth noting as the blocker we did not know we had: the accuracy numbers
looked settled, and the check moved the headline by a full point and showed the
original sample could not have supported it. The remaining unquantified figures
should be assumed to have the same problem until measured.

## 5. Would it be "better" than SafetyNets or EigenAI?

**No — and that is the wrong frame.** Honestly:

- **SafetyNets has better proof numbers than anything we have built.** 5% prover
  overhead and under 8 KB of communication. We are not competitive on that axis
  and should not pretend to be.
- **EigenAI has a deployed system** with restaked capital behind it. We have a
  research repository.

What we would have is **narrower and complementary**: the portable-determinism
result they named as open, plus a measurement study of what actually costs what.
That is a real contribution to their line of work, not a replacement for it.

The honest positioning is *"we measured the thing everyone assumes, and one
assumption was wrong"* — which is a good paper, and a much more defensible one
than *"we beat SafetyNets"*.

## 6. What would make it strong

In order:

1. ~~**Clear B1.**~~ **Done, and TF32 did show the same trap** — C4 is now a
   cross-vendor result with a pre-registered prediction behind it.
2. **Clear B2/B3.** One real model on one standard corpus.
3. ~~Resolve B6.~~ **Done — and it did become the clean negative result (C6).**
   Remaining: re-measure the depth-8 gap at convergence rather than at a fixed
   1,500 steps, so under-training can be ruled out as a partial explanation.
4. Keep the corrections section. Seven documented retractions is unusual and, in
   a measurement paper, is a credibility asset rather than an embarrassment.
