Matrix compute benchmarks
Correctness-gated Matrix throughput through the public OA operation path, with exact workload and hardware boundaries.
Matrix controlled workload
The checked tutorial calls oa::FnMatrix::matMulNt and accepts the runtime-selected route. Every row passes an independent CPU FP32 reference before timing is reported. The current reference is a Release run on Intel Iris Xe Graphics, Tiger Lake GT2, Vulkan 1.4, using the verified FP32 route.
- One warm-up process followed by the recorded pass.
- Median synchronized wall time is the primary result; p95 exposes variance.
- GFLOP/s uses
2 · M · N · K / p50 wall time. - Allocation, recording, submission, GPU work, and synchronization remain visible.
Public-operation throughput
| Shape | M | N | K | p50 ms | p95 ms | GFLOP/s |
|---|---|---|---|---|---|---|
| square-512 | 512 | 512 | 512 | 0.709 | 0.877 | 378.8 |
| square-1024 | 1024 | 1024 | 1024 | 4.901 | 5.967 | 438.2 |
| square-2048 | 2048 | 2048 | 2048 | 41.615 | 47.522 | 412.8 |
| tall-skinny | 4096 | 128 | 1024 | 2.696 | 4.354 | 398.3 |
| short-wide | 128 | 4096 | 1024 | 2.690 | 2.778 | 399.1 |
| gemv-decode | 1 | 4096 | 4096 | 5.051 | 5.879 | 6.6 |
Large square and projection shapes sustain roughly 379–438 GFLOP/s through the synchronized public path. Single-row decode falls into a GEMV-like regime and needs a dedicated route; it is not evidence for changing the square-GEMM tile.
Model-relevant projections
| Shape | M | N | K | p50 ms | p95 ms | GFLOP/s |
|---|---|---|---|---|---|---|
| nlp-qkv | 1024 | 32 | 32 | 0.076 | 0.253 | 27.4 |
| nlp-ffn-up | 1024 | 64 | 32 | 0.070 | 0.092 | 59.7 |
| nlp-ffn-down | 1024 | 32 | 64 | 0.078 | 0.088 | 53.8 |
| alm-qkv | 4096 | 384 | 384 | 2.911 | 3.789 | 414.9 |
| alm-ffn-up | 4096 | 1536 | 384 | 12.310 | 13.751 | 392.5 |
| alm-ffn-down | 4096 | 384 | 1536 | 13.346 | 18.112 | 362.0 |
Small NLP projections remain dominated by fixed dispatch costs. Larger Alm projections approach the large-GEMM regime, so fusion and submission work matter most for the former while kernel efficiency matters more for the latter.
Historical RTX 5090 training comparison
The recovered v0.6.35 snapshot measured a two-layer Fashion-MNIST MLP on an NVIDIA GeForce RTX 5090 Laptop GPU: 784 → 128 ReLU → 10, 101,770 parameters, batch 64, AdamW, and disabled presentation. OA used Vulkan 1.4.329 on NVIDIA driver 595.58.03. The comparison was recorded on 2026-06-03.
| Field | Historical contract |
|---|---|
| OA workload | 5 Fashion-MNIST epochs; 4,686 submitted training steps |
| OA forward path | FP16 W1 through NV cooperative-matrix2; BF16 W2 through KHR cooperative matrix; FP32 accumulation |
| Vendor comparator | cublasGemmEx with FP32 input/output storage and selectable FP32, TF32, BF16, or FP16 compute |
| Comparator run | 20 warmups followed by 2,000 measured steps; same model shape and batch size |
| Timing | GPU timestamp throughput and complete wall throughput reported separately |
| Runtime | Compute precision | Steps | GPU sample/s | Wall sample/s | Accuracy |
|---|---|---|---|---|---|
| OA API1 module | FP16 + BF16 | 4,686 | 1,257,071 | 906,838 | 86.40% |
| OA API2 function | FP16 + BF16 | 4,686 | 1,257,862 | 910,240 | 86.70% |
| cuBLAS 1:1 | FP32 | 2,020 | 1,065,819 | 664,532 | 84.44% |
| cuBLAS 1:1 | TF32 | 2,020 | 1,201,095 | 712,348 | 84.72% |
| cuBLAS 1:1 · matched | BF16 | 2,020 | 1,194,267 | 717,858 | 84.59% |
| cuBLAS 1:1 | FP16 | 2,020 | 1,206,289 | 726,731 | 84.40% |
| CUDA fused · carried forward | Historical | 2,000 | 412,411 | 309,957 | 86.21% |
| PyTorch CUDA · carried forward | Historical | 2,000 | — | 104,454 | 85.81% |
The precision-matched headline is OA API1 at 1.257M GPU sample/s versus cuBLAS BF16 at 1.194M: +5.3% for OA in device time. Complete wall throughput was 906.8K versus 717.9K sample/s, a +26.3% result in this harness. The nearby TF32 and FP16 cuBLAS rows show that the vendor path also occupied the 1.2M GPU sample/s tier; the older FP32-only comparison is not the appropriate headline.
This is historical evidence, not a current-package result. OA and cuBLAS used the same shape and batch, but not the same number of optimizer steps, and OA's mixed FP16/BF16 route is not numerically identical to all-BF16 compute. Accuracy therefore remains a convergence gate, not a direct quality ranking. The CUDA-fused and PyTorch rows were inherited from the preceding v0.6.34/v0.6.33 harnesses rather than rerun for v0.6.35.
Evidence boundary
These numbers are a dated hardware and driver snapshot, not a universal performance claim. The recovered RTX 5090 training result remains separate from the current Matrix tables: it measures an end-to-end MLP with older kernels and a different timing method. Older RTX Matrix tables remain excluded because their batch GFLOP/s normalization was incorrect. New hardware replaces the current Matrix baseline only after the same correctness and fresh-process protocol is rerun.
Read the matrix multiplication tutorial for the operation contract, validation formula, replay modes, and refresh procedure.
C++ terminal
cmake --build build/release --target TutorialCoreMatMulIntro -j./bin/release/sdk/tutorials/core/tuCoreMatMulIntro# Optional exhaustive CSV gridOA_AUTOTUNE_BENCH=1 ./bin/release/sdk/tutorials/core/tuCoreMatMulIntro