My expectation is that RTX4090 (and other variants of the Ada Lovelace family) will be compute capability 8.9 devices, and as such are not capable of cluster programming that requires CUDA 9.0. See here.
Related topics
| Topic | Replies | Views | Activity | |
|---|---|---|---|---|
| Confusing H100 SXM thread block cluster | 12 | 372 | October 29, 2025 | |
| Thread block clustering in Blackwell GPUs | 21 | 4841 | January 29, 2025 | |
| Does only H100 support the kernel launch configuration attribute for clustering? | 10 | 272 | March 18, 2025 | |
| What is the inter-SM linkage of DSM(cluster)? | 9 | 934 | May 8, 2026 | |
| More blocks than SMs may not make sense | 13 | 2989 | November 11, 2010 | |
| understand the mapping of the block threads to SMs in GPU | 3 | 2876 | August 2, 2018 | |
| NVIDIA Hopper Architecture In-Depth | 3 | 1307 | August 22, 2025 | |
| the 1024 threads can work concurrently? | 4 | 962 | July 24, 2017 | |
| Does 4080 support Distributed Shared Memory? | 5 | 736 | December 26, 2022 | |
| Why is the amount of thread blocks per cluster and the dynamic shared memory that I can allocate much lower than expected? | 7 | 488 | December 15, 2024 |