NVIDIA TF32, and a prediction confirmed
Overview
README.md ↗Checkpoint: CP-020 · Status: COMPLETE, run 2026-09-02 Hardware: NVIDIA L4 (Ada, compute 8.9), torch 2.6.0+cu124, GCP
The prediction, recorded in advance
E021 §3, written before any NVIDIA hardware was available:
"NVIDIA's TF32 carries a 10-bit explicit mantissa — 11 with the implicit bit — so a TF32 path should show an operand bound of 2^11, the same figure as Metal, arrived at from a different vendor's design."
The result
| operand bound | float32 rel. error | |
|---|---|---|
| TF32 off | none to 2^19 (test ceiling) | 2.8 × 10⁻⁷ |
| TF32 on | 2^11 | 2.9 × 10⁻⁴ |
Exact to 2^11, inexact from 2^12 — precisely the Metal figure. The accumulator peaked at 4.7 × 10⁶ against a 1.68 × 10⁷ limit, so every failure is an operand-precision limit with the accumulator clear.
Apple and NVIDIA share no design lineage. NC-3b's bound is a property of accelerator matmul, not an Apple quirk.
And the contract holds on every backend tested
| backend | exact int8 matmul | whole integer layer |
|---|---|---|
| Apple M5 CPU (Accelerate, ARM) | reference | reference |
| Intel Xeon 8581C (OpenBLAS, x86) | bit-identical | bit-identical |
| NVIDIA L4, TF32 off | bit-identical | bit-identical |
| NVIDIA L4, TF32 on | bit-identical | bit-identical |
Including with TF32 enabled, because int8 operands are ≤ 2^7 — four bits inside the bound. The contract doesn't need TF32 disabled; NC-3b guarantees the operands are small enough that the reduced-precision path can't lose anything.
Why it matters
This clears B1, the only blocker docs/paper-readiness.md rated fatal.
The central claim is now demonstrated across four backends from two silicon
vendors, two BLAS libraries, two operating systems, and three NumPy versions.
It also sharpens a threat-model point: TF32 is on by default in many frameworks because it's faster. A provider "running float32" very often isn't — threat T7 happening with no intent to cheat, now confirmed on a second vendor.
Reproduce
conformance/check_cuda.py — needs torch with CUDA.
python3 check_cuda.py --reference reference.json
Raw data: ../../conformance/cuda_results_l4.json
Analysis
analysis.md ↗Status: complete. A pre-registered prediction, tested on different
hardware from a different vendor, and confirmed exactly.
Date: 2026-09-02. NVIDIA L4 (Ada, compute 8.9), torch 2.6.0+cu124,
Intel Xeon @ 2.20 GHz, Ubuntu 22.04.
Raw data: ../../conformance/cuda_results_l4.json
1. The prediction, made before the hardware existed
E018 measured Apple's Metal matmul as exact on integer operands only to
2^11, regardless of accumulator size. E021 found neither CPU tested had any
operand limit, which made this look like a property of GPU matmul paths rather
than of hardware in general, and produced a prediction recorded in
experiments/021-cross-vendor/analysis.md §3:
"NVIDIA's TF32 carries a 10-bit explicit mantissa (11 with the implicit bit), so a TF32 path should show an operand bound of 2^11 — the same figure as Metal, arrived at from a different vendor's design."
2. The result
| operand | max |C| | TF32 off | TF32 on | error (TF32 on) |
|---|---|---|---|---|
| 2^10 | 9,436 | exact | exact | 0 |
| 2^11 | 19,039 | exact | exact | 0 |
| 2^12 | 36,469 | exact | INEXACT | 8 |
| 2^13 | 86,213 | exact | INEXACT | 17 |
| 2^16 | 587,731 | exact | INEXACT | 123 |
| 2^19 | 4,717,423 | exact | INEXACT | 1,058 |
TF32 on: exact to 2^11, inexact from 2^12. Exactly the Metal figure.
TF32 off is exact through 2^19, the sweep's ceiling — no operand limit at all.
The accumulator never exceeded 4.7 × 10^6 against float32's exact-integer limit of 1.68 × 10^7, so every failure above is an operand-precision limit with the accumulator entirely clear. The test was designed to make that separation unambiguous.
3. Two vendors, one number
| backend | operand bound | float32 relative error |
|---|---|---|
| Apple M5 CPU (Accelerate, NEON) | none to 2^22 | 5.0 × 10⁻⁷ |
| Intel Xeon 8581C (OpenBLAS, AVX-512) | none to 2^22 | 3.2 × 10⁻⁷ |
| Apple M5 GPU (Metal) | 2^11 | 7.7 × 10⁻⁴ |
| NVIDIA L4 (TF32 on) | 2^11 | 2.9 × 10⁻⁴ |
| NVIDIA L4 (TF32 off) | none to 2^19 | 2.8 × 10⁻⁷ |
Apple and NVIDIA share no design lineage, and their reduced-precision matmul paths land on the same 11-bit effective mantissa. NC-3b's bound is not an Apple quirk; it is a property of how accelerators do float32 matmul.
Two further readings:
The CPUs and the TF32-off GPU behave identically. Precision is not about vendor or device class per se — it is about whether the matmul path is a reduced-precision tensor-core kernel. A GPU asked to do real float32 does real float32.
And the reduced-precision path is the default. TF32 is on by default in many frameworks precisely because it is faster, so a provider "running float32" is very often not. That is threat T7 occurring without any intent to cheat, now confirmed on a second vendor.
4. The contract holds on all four backends
| backend | exact int8 matmul | whole integer layer |
|---|---|---|
| Apple M5 CPU | reference | reference |
| Intel Xeon (OpenBLAS) | bit-identical | bit-identical |
| NVIDIA L4, TF32 off | bit-identical | bit-identical |
| NVIDIA L4, TF32 on | bit-identical | bit-identical |
Including with TF32 enabled — because int8 operands are ≤ 127 = 2^7, four bits inside the 2^11 bound. The contract does not need TF32 turned off; it needs operands small enough that the reduced-precision path cannot lose anything, and NC-3b is what guarantees that.
That is the design working as intended rather than by luck, and it is the first time the contract has been checked against a GPU from a vendor we do not own.
5. What this closes
docs/paper-readiness.md lists B1 — every measurement from one machine — as
the only fatal blocker. It is now cleared:
- two CPU vendors: Apple ARM, Intel x86
- two GPU vendors: Apple Metal, NVIDIA Ada
- two BLAS libraries, two operating systems, three Python and three NumPy versions
The programme's central claim — exact integer semantics make heterogeneous hardware agree bit-for-bit — has moved from asserted, to demonstrated on one machine, to demonstrated across four backends from two silicon vendors.
And contribution C4 in the paper assessment is upgraded: the operand trap is no longer "a gotcha we found in MLX" but a cross-vendor property of accelerator matmul, confirmed by a prediction made in advance.
6. What a float-tolerance scheme would face
Separation between honest divergence and a bfloat16 cheat (2.84 × 10⁻³):
| network composition | separation |
|---|---|
| CPUs only, two vendors | ≈ 5,900× |
| CPU + NVIDIA L4 with TF32 on | ≈ 9.7× |
| CPU + Apple Metal GPU | ≈ 3.7× |
NVIDIA's TF32 is a little gentler than Metal, and 9.7× is still far too tight to set an enforcement threshold — it leaves under one order of magnitude between an honest GPU and a cheat, before any allowance for other hardware. The conclusion from E009-pre survives on NVIDIA.
7. Limitations
- One NVIDIA GPU (L4, Ada). A100/H100 (Ampere/Hopper) use the same TF32 format and should match, but were not tested.
- torch's TF32 toggle was used as the control. We verified it changes the result (float error moves 2.8 × 10⁻⁷ → 2.9 × 10⁻⁴) but did not inspect which kernel cuBLAS dispatched.
- bf16 and fp16 tensor-core paths were not swept; both should be tighter still.
- The sweep stops at 2^19; "no operand limit" for the CPUs and TF32-off means "none below the test ceiling", not "none at all".