NodusLab
E021 · CP-019

Cross-vendor conformance

does the numeric contract hold on hardware that isn't a Mac?

Completerun 2026-09-02experiments/021-cross-vendor/ ↗

Checkpoint: CP-019 · Status: COMPLETE, run 2026-09-02 Question: does the numeric contract hold on hardware that isn't a Mac?

Every determinism measurement before this came from one Apple M5 — an Apple CPU and an Apple GPU, sharing a vendor, a memory system and a compiler. This is the first test somewhere the claim could actually have failed.

reference under test
CPU Apple M5 (ARM) Intel Xeon Platinum 8581C (x86-64)
BLAS Accelerate scipy-openblas 0.3.31
SIMD NEON AVX-512
OS macOS 26.2 Debian 12
Python / NumPy 3.13.9 / 2.5.2 3.11.2 / 2.4.6

Result: it holds

test outcome
exact int8 matmul BIT-IDENTICAL
whole integer layer (matmul, requantise, LUT GELU, matmul, integer layernorm) BIT-IDENTICAL
operand bound 2^22 on both — no operand limit on either CPU
float32 matmul differs, as expected

And it corrects how we've been describing our own result

Measuring the two honest CPUs against each other:

comparison divergence
Apple vs Intel — both honest CPUs 4.83 × 10⁻⁷
Apple CPU vs Apple GPU — both honest 7.70 × 10⁻⁴
bfloat16 cheat 2.84 × 10⁻³

We have been writing "heterogeneous hardware diverges". The accurate statement is "heterogeneous CPUs agree; GPUs diverge, because their matmul paths round operands." It's a device class, not a vendor.

Separation between honest divergence and a cheat:

network separation
CPUs only, multiple vendors ≈ 5,900×
CPU + Metal GPU 3.7×

So a CPU-only heterogeneous network could safely use a float tolerance. The contract is what lets you put GPUs in the network — which is the case that matters commercially, and the one EigenAI solves by excluding all hardware but one GPU family.

A prediction this makes

Apple's Metal GPU is exact only to operands of 2^11. NVIDIA's TF32 carries a 10-bit explicit mantissa — 11 with the implicit bit — so a TF32 path should show the same 2^11 bound, reached from a different vendor's design. If it does, NC-3b generalises across accelerators.

How to run it

conformance/check.py — python3 and numpy, no GPU, no network, no data files.

python3 check.py --emit-reference > reference.json    # on the reference machine
python3 check.py --reference reference.json           # on the machine under test

Raw records: ../../benchmarks/results/021-cross-vendor.jsonl

Status: complete. The programme's central claim is now tested somewhere it could have failed, and it held. Date: 2026-09-02. Raw records: ../../benchmarks/results/021-cross-vendor.jsonl, ../../conformance/reference.json

reference machine under test
CPU Apple M5 (ARM) Intel Xeon Platinum 8581C (x86-64, Emerald Rapids)
BLAS Accelerate scipy-openblas 0.3.31
SIMD NEON AVX-512
OS macOS 26.2 Debian 12, Linux 6.1
Python / NumPy 3.13.9 / 2.5.2 3.11.2 / 2.4.6

Four independent things vary at once. Nothing was shared between the machines except the script — all data is generated from fixed seeds.


1. Result

test outcome
T2 exact int8 matmul (512×2048 @ 2048×512) BIT-IDENTICAL
T4 whole integer layer — matmul, requantise, LUT GELU, matmul, integer layernorm BIT-IDENTICAL
T1 operand bound 2^22 on both — no operand-precision limit on either CPU
T3 float32 matmul differs, as expected

The numeric contract holds across two different CPU vendors, instruction sets, BLAS libraries, operating systems, Python versions and NumPy versions.

Every prior determinism measurement in this programme came from a single Apple M5 — an Apple CPU and an Apple GPU, sharing a vendor, a memory system and a compiler. That was the largest weakness in the work. It is now substantially reduced.

2. The sharper finding: vendors are not the problem, GPUs are

Measuring the actual divergence between the two honest CPUs, rather than each machine's error against its own float64:

comparison mean relative divergence
Apple/Accelerate vs Intel/OpenBLAS — both honest CPUs 4.83 × 10⁻⁷
Apple CPU vs Apple Metal GPU — both honest (E009-pre) 7.70 × 10⁻⁴
a bfloat16 cheat (E009-pre) 2.84 × 10⁻³

Crossing vendors costs 4.8 × 10⁻⁷. Crossing from CPU to GPU costs 1,600× more than that.

This corrects the way this programme has been describing its own result. We have been writing "heterogeneous hardware diverges". The accurate statement is:

Heterogeneous CPUs agree closely. GPUs diverge, because their matmul paths round operands. The problem is a device class, not a vendor.

The consequence is quantitative. Separation between honest divergence and a bfloat16 cheat:

network composition separation
CPUs only, multiple vendors ≈ 5,900×
CPU + Metal GPU (E009-pre) 3.7×

A 5,900× separation is a threshold anyone can set. A CPU-only heterogeneous network could safely use a float tolerance. It is admitting GPUs that collapses the window — and admitting GPUs is exactly what a compute network exists to do.

So the numeric contract is not what makes heterogeneous verification possible in general. It is what makes heterogeneous verification possible with GPUs in the network — which is the case that matters commercially, and the case EigenAI solves by excluding all hardware but one GPU family.

3. T1: the operand limit is a GPU phenomenon

Both CPUs are exact to 2^22 operands (the test's ceiling, with accumulator headroom maintained throughout) — i.e. neither has an operand-precision limit, which is what true float32 should give.

Apple's Metal GPU was measured at 2^11 (E018). The limit is therefore not a property of "some hardware is sloppy" but specifically of GPU matmul paths that round their inputs.

Testable prediction, not yet run: NVIDIA's TF32 carries a 10-bit explicit mantissa (11 with the implicit bit), so a TF32 path should show an operand bound of 2^11 — the same figure as Metal, arrived at from a different vendor's design. If it does, NC-3b's bound generalises across accelerators and becomes a statement about GPU arithmetic rather than a quirk of one framework.

4. What this does to the paper's blockers

docs/paper-readiness.md lists B1 — every measurement from one machine — as the only fatal blocker. It is now partially cleared:

  • cross-vendor CPU: demonstrated. ARM vs x86, two BLAS libraries, two OSes.
  • cross-vendor GPU: still open. No NVIDIA hardware available.

The central claim has moved from asserted to demonstrated on two architectures, with the GPU case outstanding and a specific prediction attached to it.

5. Limitations

  • No GPU on the test machine. The interesting adversarial case — a reduced-precision accelerator path — is exactly what remains untested on non-Apple hardware.
  • Two CPUs. Both are recent, well-behaved, IEEE-754-conformant server/desktop parts. An older or more exotic CPU (or a compiler with fast-math enabled) could still break the contract.
  • The float divergence figure is from one matmul shape at one size.
  • NumPy and Python versions also differ between the machines, so strictly this demonstrates portability across software stacks as well as hardware — which is a stronger result, but means the variables are not isolated.