Kernel Fusion in NVIDIA CUDA: Optimizing Memory Traffic and Launch Overhead

Originally published at: Kernel Fusion in NVIDIA CUDA: Optimizing Memory Traffic and Launch Overhead | NVIDIA Technical Blog

There are many ways to optimize code for GPUs. In this post, you’ll learn how kernel fusion can improve memory bandwidth and reduce kernel launch overhead, along with multiple ways to apply it in NVIDIA CUDA code. A common bottleneck when writing GPU code is that GPU compute is so fast that even high-bandwidth device…

Great post. The 3× result in the `sum(abs(x))` example shows how eliminating intermediate global-memory traffic can accelerate a single operation. In a full CFD solver, those savings can compound across many stages.

Brae, a GPU-native CUDA port of OpenFOAM, applies kernel fusion alongside other GPU-native optimizations throughout its solver loops. For `rhoSimpleFoam`, a 10M-cell case running 100 SIMPLE iterations completed in 109 seconds on one GH200, versus 921 seconds on all 64 Grace CPU cores, about 8.4× faster; while matching OpenFOAM to `7e-08`.

Across six cases, Brae warm was more than 20× faster than SPUMA by geometric mean, reaching 136× on the 1M-cell case. It’s a useful real-world example of kernel fusion contributing to order-of-magnitude application-level speedups: