“FP32 is less common in modern ML workloads and often less optimized on recent hardware compared to FP16 or BF16, which may partly explain why it’s easier to achieve performance gains over PyTorch with FP32 kernels.”
The headline numbers beat PyTorch on FP32, the precision PyTorch optimizes least. The “FP32” Conv2D kernel computes in half precision, and the tolerance allows it. Each kernel is tuned to one fixed problem size. On FP16 Flash Attention the generated kernel reaches 9% of PyTorch.