Engineering
From Python + CUDA to Pure Rust + Metal 4: How zml-flow-rs Stops Us Writing a Kernel Per Model
Every new diffusion model ships as PyTorch with CUDA assumptions baked in, and every port to Apple Silicon has meant hand-writing its kernels again. zml-flow-rs reads only the weights, discovers the semantics from one reference trace, synthesizes its fused kernels from data, and compiles the result into a standalone Rust binary. Two different architectures, zero model-specific code, and a Krea-2 PNG that is bit-identical to the reference.
September 25, 2026
23 min read
Read Story →