NodusLab
E022 · CP-020

NVIDIA TF32, and a prediction confirmed

Completerun 2026-09-02experiments/022-nvidia-tf32/ ↗

Checkpoint: CP-020 · Status: COMPLETE, run 2026-09-02 Hardware: NVIDIA L4 (Ada, compute 8.9), torch 2.6.0+cu124, GCP

The prediction, recorded in advance

E021 §3, written before any NVIDIA hardware was available:

"NVIDIA's TF32 carries a 10-bit explicit mantissa — 11 with the implicit bit — so a TF32 path should show an operand bound of 2^11, the same figure as Metal, arrived at from a different vendor's design."

The result

operand bound float32 rel. error
TF32 off none to 2^19 (test ceiling) 2.8 × 10⁻⁷
TF32 on 2^11 2.9 × 10⁻⁴

Exact to 2^11, inexact from 2^12 — precisely the Metal figure. The accumulator peaked at 4.7 × 10⁶ against a 1.68 × 10⁷ limit, so every failure is an operand-precision limit with the accumulator clear.

Apple and NVIDIA share no design lineage. NC-3b's bound is a property of accelerator matmul, not an Apple quirk.

And the contract holds on every backend tested

backend exact int8 matmul whole integer layer
Apple M5 CPU (Accelerate, ARM) reference reference
Intel Xeon 8581C (OpenBLAS, x86) bit-identical bit-identical
NVIDIA L4, TF32 off bit-identical bit-identical
NVIDIA L4, TF32 on bit-identical bit-identical

Including with TF32 enabled, because int8 operands are ≤ 2^7 — four bits inside the bound. The contract doesn't need TF32 disabled; NC-3b guarantees the operands are small enough that the reduced-precision path can't lose anything.

Why it matters

This clears B1, the only blocker docs/paper-readiness.md rated fatal. The central claim is now demonstrated across four backends from two silicon vendors, two BLAS libraries, two operating systems, and three NumPy versions.

It also sharpens a threat-model point: TF32 is on by default in many frameworks because it's faster. A provider "running float32" very often isn't — threat T7 happening with no intent to cheat, now confirmed on a second vendor.

Reproduce

conformance/check_cuda.py — needs torch with CUDA.

python3 check_cuda.py --reference reference.json

Raw data: ../../conformance/cuda_results_l4.json

Status: complete. A pre-registered prediction, tested on different hardware from a different vendor, and confirmed exactly. Date: 2026-09-02. NVIDIA L4 (Ada, compute 8.9), torch 2.6.0+cu124, Intel Xeon @ 2.20 GHz, Ubuntu 22.04. Raw data: ../../conformance/cuda_results_l4.json


1. The prediction, made before the hardware existed

E018 measured Apple's Metal matmul as exact on integer operands only to 2^11, regardless of accumulator size. E021 found neither CPU tested had any operand limit, which made this look like a property of GPU matmul paths rather than of hardware in general, and produced a prediction recorded in experiments/021-cross-vendor/analysis.md §3:

"NVIDIA's TF32 carries a 10-bit explicit mantissa (11 with the implicit bit), so a TF32 path should show an operand bound of 2^11 — the same figure as Metal, arrived at from a different vendor's design."

2. The result

operand max |C| TF32 off TF32 on error (TF32 on)
2^10 9,436 exact exact 0
2^11 19,039 exact exact 0
2^12 36,469 exact INEXACT 8
2^13 86,213 exact INEXACT 17
2^16 587,731 exact INEXACT 123
2^19 4,717,423 exact INEXACT 1,058

TF32 on: exact to 2^11, inexact from 2^12. Exactly the Metal figure.

TF32 off is exact through 2^19, the sweep's ceiling — no operand limit at all.

The accumulator never exceeded 4.7 × 10^6 against float32's exact-integer limit of 1.68 × 10^7, so every failure above is an operand-precision limit with the accumulator entirely clear. The test was designed to make that separation unambiguous.

3. Two vendors, one number

backend operand bound float32 relative error
Apple M5 CPU (Accelerate, NEON) none to 2^22 5.0 × 10⁻⁷
Intel Xeon 8581C (OpenBLAS, AVX-512) none to 2^22 3.2 × 10⁻⁷
Apple M5 GPU (Metal) 2^11 7.7 × 10⁻⁴
NVIDIA L4 (TF32 on) 2^11 2.9 × 10⁻⁴
NVIDIA L4 (TF32 off) none to 2^19 2.8 × 10⁻⁷

Apple and NVIDIA share no design lineage, and their reduced-precision matmul paths land on the same 11-bit effective mantissa. NC-3b's bound is not an Apple quirk; it is a property of how accelerators do float32 matmul.

Two further readings:

The CPUs and the TF32-off GPU behave identically. Precision is not about vendor or device class per se — it is about whether the matmul path is a reduced-precision tensor-core kernel. A GPU asked to do real float32 does real float32.

And the reduced-precision path is the default. TF32 is on by default in many frameworks precisely because it is faster, so a provider "running float32" is very often not. That is threat T7 occurring without any intent to cheat, now confirmed on a second vendor.

4. The contract holds on all four backends

backend exact int8 matmul whole integer layer
Apple M5 CPU reference reference
Intel Xeon (OpenBLAS) bit-identical bit-identical
NVIDIA L4, TF32 off bit-identical bit-identical
NVIDIA L4, TF32 on bit-identical bit-identical

Including with TF32 enabled — because int8 operands are ≤ 127 = 2^7, four bits inside the 2^11 bound. The contract does not need TF32 turned off; it needs operands small enough that the reduced-precision path cannot lose anything, and NC-3b is what guarantees that.

That is the design working as intended rather than by luck, and it is the first time the contract has been checked against a GPU from a vendor we do not own.

5. What this closes

docs/paper-readiness.md lists B1 — every measurement from one machine — as the only fatal blocker. It is now cleared:

  • two CPU vendors: Apple ARM, Intel x86
  • two GPU vendors: Apple Metal, NVIDIA Ada
  • two BLAS libraries, two operating systems, three Python and three NumPy versions

The programme's central claim — exact integer semantics make heterogeneous hardware agree bit-for-bit — has moved from asserted, to demonstrated on one machine, to demonstrated across four backends from two silicon vendors.

And contribution C4 in the paper assessment is upgraded: the operand trap is no longer "a gotcha we found in MLX" but a cross-vendor property of accelerator matmul, confirmed by a prediction made in advance.

6. What a float-tolerance scheme would face

Separation between honest divergence and a bfloat16 cheat (2.84 × 10⁻³):

network composition separation
CPUs only, two vendors ≈ 5,900×
CPU + NVIDIA L4 with TF32 on ≈ 9.7×
CPU + Apple Metal GPU ≈ 3.7×

NVIDIA's TF32 is a little gentler than Metal, and 9.7× is still far too tight to set an enforcement threshold — it leaves under one order of magnitude between an honest GPU and a cheat, before any allowance for other hardware. The conclusion from E009-pre survives on NVIDIA.

7. Limitations

  • One NVIDIA GPU (L4, Ada). A100/H100 (Ampere/Hopper) use the same TF32 format and should match, but were not tested.
  • torch's TF32 toggle was used as the control. We verified it changes the result (float error moves 2.8 × 10⁻⁷ → 2.9 × 10⁻⁴) but did not inspect which kernel cuBLAS dispatched.
  • bf16 and fp16 tensor-core paths were not swept; both should be tighter still.
  • The sweep stops at 2^19; "no operand limit" for the CPUs and TF32-off means "none below the test ceiling", not "none at all".