I benched my Jetson Orin nano devkit on cutlass under f16 and the highest performance mma was around 9124 GFlop/s. This seems really high, the A100 does 256 flop/cycle per tensor core, so 1024 flop/cycle/SM. If my jetson is running at 625 MHz, total performance would be around (625e6 Hz) * (8 SMs) * (1024 flop/cycle) = 5.12e12 flop/s. These types of calculations seem to work perfectly for the A100, 3090, and 4090 from what I’ve seen personally. Is there something I’m missing?
Related topics
| Topic | Replies | Views | Activity | |
|---|---|---|---|---|
| Jetson orin nano fp16/int8 performance | 7 | 2189 | March 18, 2025 | |
| The performance of the Jetson Orin Nano module does not match the data provided on the official website | 14 | 3084 | September 28, 2023 | |
| Verifying TOPS with Jetson Orin Nano | 1 | 610 | December 30, 2024 | |
| Performance discrepancy: TensorRT achieves ~10 TFLOPS vs. 17 TFLOPS spec on Orin Nano (Super mode) | 6 | 222 | February 25, 2026 | |
| Jetson Orin AI Performance | 6 | 1757 | February 1, 2023 | |
| FFT runs slower than expected on Jetson Orin Nano Super | 1 | 252 | July 21, 2025 | |
| The tensor core performance detail of Jetson AGX Orin 32GB | 13 | 1612 | June 13, 2023 | |
| Jetson AGX Orin TOPs / CUDA Cores Explained | 7 | 7961 | May 11, 2023 | |
| Could you share me with some intuitive examples as below which tell the use cases for different TOPS value?(in the range of 0.5-10 TOPS) | 3 | 220 | July 10, 2024 | |
| Jetson Nano Performance | 1 | 662 | December 19, 2021 |