Cross-vendor conformance
does the numeric contract hold on hardware that isn't a Mac?
Overview
README.md ↗Checkpoint: CP-019 · Status: COMPLETE, run 2026-09-02 Question: does the numeric contract hold on hardware that isn't a Mac?
Every determinism measurement before this came from one Apple M5 — an Apple CPU and an Apple GPU, sharing a vendor, a memory system and a compiler. This is the first test somewhere the claim could actually have failed.
| reference | under test | |
|---|---|---|
| CPU | Apple M5 (ARM) | Intel Xeon Platinum 8581C (x86-64) |
| BLAS | Accelerate | scipy-openblas 0.3.31 |
| SIMD | NEON | AVX-512 |
| OS | macOS 26.2 | Debian 12 |
| Python / NumPy | 3.13.9 / 2.5.2 | 3.11.2 / 2.4.6 |
Result: it holds
| test | outcome |
|---|---|
| exact int8 matmul | BIT-IDENTICAL |
| whole integer layer (matmul, requantise, LUT GELU, matmul, integer layernorm) | BIT-IDENTICAL |
| operand bound | 2^22 on both — no operand limit on either CPU |
| float32 matmul | differs, as expected |
And it corrects how we've been describing our own result
Measuring the two honest CPUs against each other:
| comparison | divergence |
|---|---|
| Apple vs Intel — both honest CPUs | 4.83 × 10⁻⁷ |
| Apple CPU vs Apple GPU — both honest | 7.70 × 10⁻⁴ |
| bfloat16 cheat | 2.84 × 10⁻³ |
We have been writing "heterogeneous hardware diverges". The accurate statement is "heterogeneous CPUs agree; GPUs diverge, because their matmul paths round operands." It's a device class, not a vendor.
Separation between honest divergence and a cheat:
| network | separation |
|---|---|
| CPUs only, multiple vendors | ≈ 5,900× |
| CPU + Metal GPU | 3.7× |
So a CPU-only heterogeneous network could safely use a float tolerance. The contract is what lets you put GPUs in the network — which is the case that matters commercially, and the one EigenAI solves by excluding all hardware but one GPU family.
A prediction this makes
Apple's Metal GPU is exact only to operands of 2^11. NVIDIA's TF32 carries a 10-bit explicit mantissa — 11 with the implicit bit — so a TF32 path should show the same 2^11 bound, reached from a different vendor's design. If it does, NC-3b generalises across accelerators.
How to run it
conformance/check.py — python3 and numpy, no GPU, no network, no data files.
python3 check.py --emit-reference > reference.json # on the reference machine
python3 check.py --reference reference.json # on the machine under test
Raw records: ../../benchmarks/results/021-cross-vendor.jsonl
Analysis
analysis.md ↗Status: complete. The programme's central claim is now tested somewhere it
could have failed, and it held.
Date: 2026-09-02.
Raw records: ../../benchmarks/results/021-cross-vendor.jsonl,
../../conformance/reference.json
| reference | machine under test | |
|---|---|---|
| CPU | Apple M5 (ARM) | Intel Xeon Platinum 8581C (x86-64, Emerald Rapids) |
| BLAS | Accelerate | scipy-openblas 0.3.31 |
| SIMD | NEON | AVX-512 |
| OS | macOS 26.2 | Debian 12, Linux 6.1 |
| Python / NumPy | 3.13.9 / 2.5.2 | 3.11.2 / 2.4.6 |
Four independent things vary at once. Nothing was shared between the machines except the script — all data is generated from fixed seeds.
1. Result
| test | outcome |
|---|---|
| T2 exact int8 matmul (512×2048 @ 2048×512) | BIT-IDENTICAL |
| T4 whole integer layer — matmul, requantise, LUT GELU, matmul, integer layernorm | BIT-IDENTICAL |
| T1 operand bound | 2^22 on both — no operand-precision limit on either CPU |
| T3 float32 matmul | differs, as expected |
The numeric contract holds across two different CPU vendors, instruction sets, BLAS libraries, operating systems, Python versions and NumPy versions.
Every prior determinism measurement in this programme came from a single Apple M5 — an Apple CPU and an Apple GPU, sharing a vendor, a memory system and a compiler. That was the largest weakness in the work. It is now substantially reduced.
2. The sharper finding: vendors are not the problem, GPUs are
Measuring the actual divergence between the two honest CPUs, rather than each machine's error against its own float64:
| comparison | mean relative divergence |
|---|---|
| Apple/Accelerate vs Intel/OpenBLAS — both honest CPUs | 4.83 × 10⁻⁷ |
| Apple CPU vs Apple Metal GPU — both honest (E009-pre) | 7.70 × 10⁻⁴ |
| a bfloat16 cheat (E009-pre) | 2.84 × 10⁻³ |
Crossing vendors costs 4.8 × 10⁻⁷. Crossing from CPU to GPU costs 1,600× more than that.
This corrects the way this programme has been describing its own result. We have been writing "heterogeneous hardware diverges". The accurate statement is:
Heterogeneous CPUs agree closely. GPUs diverge, because their matmul paths round operands. The problem is a device class, not a vendor.
The consequence is quantitative. Separation between honest divergence and a bfloat16 cheat:
| network composition | separation |
|---|---|
| CPUs only, multiple vendors | ≈ 5,900× |
| CPU + Metal GPU (E009-pre) | 3.7× |
A 5,900× separation is a threshold anyone can set. A CPU-only heterogeneous network could safely use a float tolerance. It is admitting GPUs that collapses the window — and admitting GPUs is exactly what a compute network exists to do.
So the numeric contract is not what makes heterogeneous verification possible in general. It is what makes heterogeneous verification possible with GPUs in the network — which is the case that matters commercially, and the case EigenAI solves by excluding all hardware but one GPU family.
3. T1: the operand limit is a GPU phenomenon
Both CPUs are exact to 2^22 operands (the test's ceiling, with accumulator headroom maintained throughout) — i.e. neither has an operand-precision limit, which is what true float32 should give.
Apple's Metal GPU was measured at 2^11 (E018). The limit is therefore not a property of "some hardware is sloppy" but specifically of GPU matmul paths that round their inputs.
Testable prediction, not yet run: NVIDIA's TF32 carries a 10-bit explicit mantissa (11 with the implicit bit), so a TF32 path should show an operand bound of 2^11 — the same figure as Metal, arrived at from a different vendor's design. If it does, NC-3b's bound generalises across accelerators and becomes a statement about GPU arithmetic rather than a quirk of one framework.
4. What this does to the paper's blockers
docs/paper-readiness.md lists B1 — every measurement from one machine — as the
only fatal blocker. It is now partially cleared:
- cross-vendor CPU: demonstrated. ARM vs x86, two BLAS libraries, two OSes.
- cross-vendor GPU: still open. No NVIDIA hardware available.
The central claim has moved from asserted to demonstrated on two architectures, with the GPU case outstanding and a specific prediction attached to it.
5. Limitations
- No GPU on the test machine. The interesting adversarial case — a reduced-precision accelerator path — is exactly what remains untested on non-Apple hardware.
- Two CPUs. Both are recent, well-behaved, IEEE-754-conformant server/desktop parts. An older or more exotic CPU (or a compiler with fast-math enabled) could still break the contract.
- The float divergence figure is from one matmul shape at one size.
- NumPy and Python versions also differ between the machines, so strictly this demonstrates portability across software stacks as well as hardware — which is a stronger result, but means the variables are not isolated.