I’m not sure what you mean by “down to the mma instruction itself”. You may wish to read the relevant sections of the PTX doc. Or study an example. You can find examples on these forums. Here is one.
Related topics
| Topic | Replies | Views | Activity | |
|---|---|---|---|---|
| Can we directly use register value for tensor core calculation? | 3 | 824 | October 18, 2023 | |
| Register use by wmma | 2 | 145 | September 25, 2024 | |
| [Question] How does the threads in a warp work collectively? | 4 | 307 | July 8, 2024 | |
| Is the documentation on MMA from the NVIDIA GTC 2020 talk incorrect? | 1 | 402 | November 23, 2023 | |
| Get wrong result using tensor core example | 7 | 599 | June 17, 2024 | |
| When using tensor core with "wmma" problem | 0 | 421 | January 13, 2023 | |
| How many tensor cores to execute the wmma.mma.sync.aligned.{alayout}.{blayout}.m16n16k16 instruction? | 23 | 527 | December 12, 2025 | |
| Why does WMMA and MMA support different matrix tile size? | 1 | 2365 | October 28, 2023 | |
| Order of registers in MMA calls | 2 | 1332 | June 13, 2022 | |
| WMMA - What does "warp matrix operations" mean? | 6 | 8202 | October 4, 2022 |