How to Optimize a CUDA Matmul Kernel: a Worklog

The blog is inspired by https://siboehm.com/articles/22/CUDA-MMM.

Benchmark

We choose FLOPS as the only thing ot evaluate the kernel.

We record the time cuda needs, then we calculate the FLOPS.

Autotuning

< Back to all posts