You might want to file a bug and refer to this thread.
rs277
6
Related topics
| Topic | Replies | Views | Activity | |
|---|---|---|---|---|
| Incorrect PTX complier warning? Potential performance loss for WGMMA pipeline crossing func boundary | 0 | 126 | October 17, 2025 | |
| PTXAS says registers not enough for m64n256k32 integer wgmma after setting regcount from 128 to 200 by setmaxnreg.inc | 7 | 158 | April 21, 2026 | |
| Sm_90a: only one WGMMA of a 2-mma group overlaps the accumulator XOR-reduction — why? | 5 | 98 | June 24, 2026 | |
| Example with wgmma.mma_async | 4 | 2660 | November 23, 2024 | |
| Question about wgmma instruction in Hopper | 3 | 1583 | October 25, 2024 | |
| How many tensor cores to execute the wmma.mma.sync.aligned.{alayout}.{blayout}.m16n16k16 instruction? | 23 | 478 | December 12, 2025 | |
| Throughput and latency of mma.sync instruction | 4 | 670 | August 19, 2024 | |
| Wmma vs Wgmma On H100 GPU | 4 | 421 | December 15, 2025 | |
| Fastest Tiled WMMA for Matrices of Any Size? | 3 | 529 | October 26, 2024 | |
| PTX instruction `mma` not lowered to tensor core related SASS instruction | 2 | 1452 | March 22, 2022 |