Performance

NLP architecture and tokenizer matrix

A controlled 15-model sweep on the Intel Iris Xe reference system. Every row trains the same dense next-token workload and passes evaluation, generation, and checkpoint checks.

45 measured processesVulkan FP32July 2026 reference

Controlled workload

  • 300 optimizer steps, batch 64, sequence 16, and 1,024 predicted positions per step.
  • Model width 32 and hidden/FFN width 64; MoE uses four DFF=16 experts with top-2 routing.
  • One excluded warm-up followed by three measured processes per executable.
  • Wall time includes forward, cross entropy, backward, AdamW, submission, synchronization, metrics, and callbacks.

End-to-end wall performance

ArchitectureTokenizerParamsms/steptoken/sbyte/s
RNNByte31,1047.01 ± 0.24146.28 ± 4.89K146.28 ± 4.89K
RNNBPE37,3128.46 ± 0.73121.99 ± 9.91K338.59 ± 27.51K
RNNChar8,8915.75 ± 0.39178.72 ± 11.50K
GRUByte43,64814.37 ± 0.8871.55 ± 4.50K71.55 ± 4.50K
GRUBPE49,85614.91 ± 0.2568.70 ± 1.16K190.67 ± 3.20K
GRUChar21,43511.89 ± 0.3986.25 ± 2.85K
TransformerByte25,7609.24 ± 0.28110.95 ± 3.42K110.95 ± 3.42K
TransformerBPE29,9209.59 ± 0.07106.73 ± 0.73K296.24 ± 2.01K
TransformerChar10,8758.02 ± 0.06127.65 ± 0.96K
MoE TransformerByte28,06813.55 ± 0.0275.60 ± 0.08K75.60 ± 0.08K
MoE TransformerBPE32,22814.01 ± 0.1073.10 ± 0.52K202.89 ± 1.46K
MoE TransformerChar13,18312.62 ± 0.1381.15 ± 0.85K
Mamba-3Byte25,80035.56 ± 1.3428.83 ± 1.10K28.83 ± 1.10K
Mamba-3BPE29,96038.60 ± 3.1526.71 ± 2.12K74.12 ± 5.90K
Mamba-3Char10,91536.55 ± 1.4028.06 ± 1.06K

Interpretation

Raw token accuracy is not compared across tokenizer families. BPE uses a 320-token vocabulary and compresses the 576-byte corpus to 210 model tokens, so byte throughput is the useful Byte-versus-BPE metric. At these shapes BPE increases represented text throughput by 2.32× to 2.68× despite similar model-token throughput.

These are whole-application tutorial measurements, not normalized peak-silicon claims. Hardware, driver, precision, and workload are part of the result.