The layer-audit protocol
Overview
README.md ↗Checkpoint: CP-022 · Status: COMPLETE, run 2026-09-02
E016 derived a design on paper: commit to every layer, challenge one at random, back it with a stake. This is that design built and run against Qwen2.5-0.5B.
provider coordinator
-------- -----------
run the model
hash each layer output
build a Merkle tree
submit (root, final output) --> record the commitment
<-- challenge: layer k, drawn uniformly
AFTER the commitment
open layer k:
input + two Merkle paths --> verify the input was committed
recompute the layer
check ITS OWN digest against the root
Results
| honest provider, all 24 layers challenged | accepted |
| cheating on 1 of 24 layers (200 trials) | detected 4.5% (1/L = 4.2%) |
| cheat everywhere, then answer the challenge honestly | REJECTED |
| audit traffic vs verifying every layer | 4.2% |
Two things building it fixed
The output never crosses the wire. E016 assumed the opening carries a layer's input and output. It doesn't — the verifier recomputes the output and checks its own digest against the commitment. Halves the opening (0.459 → 0.230 MB) and is strictly more secure, because the provider never gets to say what the output was.
Layer 0 is free. Its input is the embedding, which the verifier derives from token ids it already holds. The opening is two Merkle paths and nothing else: 160 bytes.
What it is not
Not a proof. A single-layer cheat escapes with probability 1 - 1/L = 0.958 at
24 layers — the defence is the stake, not the probability. And the whole
scheme rests on NC-0.5: without exact recomputability the verifier cannot tell a
cheat from its own rounding, and the commitment commits to nothing stable.
Files: implementation/audit.py (protocol), implementation/run.py (the
experiment), analysis.md (the result).
Analysis
analysis.md ↗Status: complete. The design E016 derived on paper now exists, runs against
a real model, and two of its details were wrong.
Date: 2026-09-02. Qwen2.5-0.5B under NC-0.5, 24 layers, 32 tokens.
Raw records: ../../benchmarks/results/024-layer-audit.jsonl
1. It works
| honest provider, every layer challenged in turn | accepted on all 24 |
| cheating on 1 of 24 layers, 200 trials | detected 9/200 = 0.045 (expected 1/L = 0.042) |
| cheat on all layers, then answer the challenge honestly | REJECTED — "opened input was not the one committed to" |
The third row is the one that matters, and it is why the commitment exists. Without it the audit is theatre: a provider learns which layer is checked and computes only that one. With it, every layer is fixed before the challenge is drawn, so answering one honestly requires having done them all honestly.
That is now demonstrated rather than asserted.
2. What crosses the wire
| bytes | |
|---|---|
| submission — Merkle root + the answer | 0.229 MB |
| opening one challenged layer | 0.230 MB |
| opening layer 0 | 0.160 kB |
| all intermediates (what E016 measured) | 5.505 MB |
The audit transmits 4.2% of what verifying every layer requires.
That 4.2% is 1/L, and so is the detection rate — the same ratio governs both, which is the honest way to state the trade: you send one layer of L, and you catch one cheat in L.
3. Two things the paper design got wrong
Building it surfaced two details that derivation missed. Both halve or eliminate transmission, and neither is visible until you write the verifier.
The layer's output does not need transmitting. E016 assumed the opening carries a layer's input and output. It does not: the verifier recomputes the output itself and checks the digest of its own result against the commitment. If the provider committed something else, the Merkle path fails.
This halves the opening (0.459 MB → 0.230 MB) and is strictly more secure — the provider never gets to state what the output was, so it cannot substitute one after the fact.
Layer 0 costs nothing at all. Its input is the embedding, which the verifier derives from token ids it already holds. The opening is then two Merkle paths and nothing else: 160 bytes, a 1,400× reduction against a typical layer.
The same argument extends: any layer whose input the verifier can derive independently is free to audit. For a transformer that is only layer 0, but for an architecture with more derivable state it would be more.
4. Where this leaves the design
E016's arithmetic said: bandwidth flat in depth, compute saved scaling with depth, and a stake equal to the job's compute value making cheating negative-expected-value at any depth. All of that survives contact with an implementation, and the constant is better than assumed.
What the protocol still does not do, and must not be described as doing:
- It is not a proof. A single-layer cheat escapes with probability
1 - 1/L= 0.958 at 24 layers. The defence is the stake, not the probability. E016 §4 gives the threshold, and themandLcancel: a stake equal to the job's compute value suffices at any depth and any level of greed. - It constrains the answer, not the work (E013). A provider that obtains the right intermediates cheaply still passes.
- It requires the numeric contract. The entire scheme rests on a layer being exactly recomputable by a different machine. Without NC-0.5 the verifier cannot distinguish a cheat from its own rounding, and the commitment commits to nothing stable.
5. Limitations
- 32 tokens, one prompt, one model. The bandwidth figures scale with
s·dand the 1/L ratio does not, so the shape holds, but the absolute megabytes are specific to this shape. - 200 trials gives the detection rate to about ±0.015. It agrees with 1/L, which is what the mechanism guarantees by construction; the empirical run is a check on the implementation, not a measurement of an uncertain quantity.
- No stake, slashing or settlement is implemented here. Those live in the coordinator's ledger, and the economic half of E016 §4 remains unbuilt.
- The challenge uses
secrets.randbelowlocally. In deployment the randomness must be unpredictable to the provider and auditable by the consumer, which is a commit-reveal or beacon problem this experiment does not address.
Addendum — the challenge is now a two-party coin
The protocol above drew its challenge with secrets.randbelow. That is a
private coin, and it quietly assumed the one thing this programme refuses to
assume everywhere else: that the coordinator is honest. A coordinator paid by
the provider, or merely lazy, can pick a layer it knows is clean, and the
consumer has no way to tell the difference from a fair draw.
Replaced with commit-reveal:
- the coordinator publishes
H(nonce)before the provider runs - the provider runs and submits
(root, answer) - the coordinator reveals
nonce challenge = H(nonce ‖ root ‖ answer) mod L
Each side is bound before it can see the other's contribution, so neither can steer the outcome, and step 4 is reproducible by anyone holding the commitment, the nonce and the submission. Cost on the wire: 32 bytes once per job.
Measured on Qwen2.5-0.5B, 24 layers:
| consumer re-derives the challenged layer | verified |
| coordinator claims a different layer | rejected |
| coordinator reveals a nonce it did not commit to | rejected |
| uniformity of the derived coin, 4000 draws | χ² = 14.6 on 23 df |
The ordering is the security property, not the hashing
A coordinator that commits after seeing the root can simply grind nonces until the coin lands where it wants. Measured: reaching a chosen layer took 17 guesses against an expected ~L = 24. That is free.
So the ordering cannot be left to documentation. Coordinator.challenge()
raises if commit() has not run, and raises again if a nonce is reused across
submissions — the two ways the guarantee is lost in practice. Both are tested.
What this still does not give
The commitment must be published somewhere the consumer can see it before the job runs, and nothing here does that. In deployment it needs a timestamped channel — a chain, a transparency log, or the consumer holding it directly. Without that, a colluding coordinator can backdate its commitment and the grinding attack returns at full strength.
This closes the protocol hole. The publication requirement is now the explicit assumption, which is the honest place for it to sit.