How to Optimize a CUDA Matmul Kernel: a Worklog
The blog is inspired by https://siboehm.com/articles/22/CUDA-MMM.
Benchmark
We choose FLOPS as the only thing ot evaluate the kernel.
We record the time cuda needs, then we calculate the FLOPS.
The blog is inspired by https://siboehm.com/articles/22/CUDA-MMM.
We choose FLOPS as the only thing ot evaluate the kernel.
We record the time cuda needs, then we calculate the FLOPS.