Performance
NLP architecture and tokenizer matrix
A controlled 15-model sweep on the Intel Iris Xe reference system. Every row trains the same dense next-token workload and passes evaluation, generation, and checkpoint checks.
45 measured processesVulkan FP32July 2026 reference
Controlled workload
- 300 optimizer steps, batch 64, sequence 16, and 1,024 predicted positions per step.
- Model width 32 and hidden/FFN width 64; MoE uses four DFF=16 experts with top-2 routing.
- One excluded warm-up followed by three measured processes per executable.
- Wall time includes forward, cross entropy, backward, AdamW, submission, synchronization, metrics, and callbacks.
End-to-end wall performance
| Architecture | Tokenizer | Params | ms/step | token/s | byte/s |
|---|---|---|---|---|---|
| RNN | Byte | 31,104 | 7.01 ± 0.24 | 146.28 ± 4.89K | 146.28 ± 4.89K |
| RNN | BPE | 37,312 | 8.46 ± 0.73 | 121.99 ± 9.91K | 338.59 ± 27.51K |
| RNN | Char | 8,891 | 5.75 ± 0.39 | 178.72 ± 11.50K | — |
| GRU | Byte | 43,648 | 14.37 ± 0.88 | 71.55 ± 4.50K | 71.55 ± 4.50K |
| GRU | BPE | 49,856 | 14.91 ± 0.25 | 68.70 ± 1.16K | 190.67 ± 3.20K |
| GRU | Char | 21,435 | 11.89 ± 0.39 | 86.25 ± 2.85K | — |
| Transformer | Byte | 25,760 | 9.24 ± 0.28 | 110.95 ± 3.42K | 110.95 ± 3.42K |
| Transformer | BPE | 29,920 | 9.59 ± 0.07 | 106.73 ± 0.73K | 296.24 ± 2.01K |
| Transformer | Char | 10,875 | 8.02 ± 0.06 | 127.65 ± 0.96K | — |
| MoE Transformer | Byte | 28,068 | 13.55 ± 0.02 | 75.60 ± 0.08K | 75.60 ± 0.08K |
| MoE Transformer | BPE | 32,228 | 14.01 ± 0.10 | 73.10 ± 0.52K | 202.89 ± 1.46K |
| MoE Transformer | Char | 13,183 | 12.62 ± 0.13 | 81.15 ± 0.85K | — |
| Mamba-3 | Byte | 25,800 | 35.56 ± 1.34 | 28.83 ± 1.10K | 28.83 ± 1.10K |
| Mamba-3 | BPE | 29,960 | 38.60 ± 3.15 | 26.71 ± 2.12K | 74.12 ± 5.90K |
| Mamba-3 | Char | 10,915 | 36.55 ± 1.40 | 28.06 ± 1.06K | — |
Interpretation
Raw token accuracy is not compared across tokenizer families. BPE uses a 320-token vocabulary and compresses the 576-byte corpus to 210 model tokens, so byte throughput is the useful Byte-versus-BPE metric. At these shapes BPE increases represented text throughput by 2.32× to 2.68× despite similar model-token throughput.
These are whole-application tutorial measurements, not normalized peak-silicon claims. Hardware, driver, precision, and workload are part of the result.