How to collect GPU memory bandwidth usage data with nsys profiler on Pegasus?

yudong.wu · October 8, 2021, 11:35pm

Hello,

We are now looking for ways to profile GPU memory bandwidth data when running CUDA kernels on GPU. However, when we try to use nsys profiler, it reports that the GPUs on Pegasus do not support --gpu-metrics sampling. Is there any other ways to collect information similar to what --gpu-metrics sampling will report?

Please provide the following info (check/uncheck the boxes after creating this topic):
Software Version
DRIVE OS Linux 5.2.6
DRIVE OS Linux 5.2.0
DRIVE OS Linux 5.2.0 and DriveWorks 3.5
NVIDIA DRIVE™ Software 10.0 (Linux)
NVIDIA DRIVE™ Software 9.0 (Linux)
other DRIVE OS version
other

Target Operating System
Linux
QNX
other

Hardware Platform
NVIDIA DRIVE™ AGX Xavier DevKit (E3550)
NVIDIA DRIVE™ AGX Pegasus DevKit (E3550)
other

SDK Manager Version
1.6.1.8175
1.6.0.8170
other

Host Machine Version
native Ubuntu 18.04
other

Thanks.

SivaRamaKrishnaNV · October 10, 2021, 4:33am

Dear @yudong.wu,
Could you share the used nsys version? If you used CLI, please share the used command.

SivaRamaKrishnaNV · October 11, 2021, 4:27pm

Dear @yudong.wu ,
Pegasus do not support --gpu-metrics sampling

It is supported for GPU architecture greater than Turing. It seems you have volta based dGPU?

yudong.wu · October 11, 2021, 6:24pm

Hi @SivaRamaKrishnaNV,

Thanks for your comments. It seems that we have TU104 dGPU on our platform, however when we use command nsys profile --gpu-metrics-device=help, it responds that neither GPU devices on Pegasus support the GPU sampling. I am wondering whether --gpu-metrics sampling doesn’t support the Turing GPU on Pegasus? If it is the truth, may I ask whether there are any other profiling tools that we can use to collect information similar to --gpu-metrics, like VRAM read/write bandwidth, SM occupancy, etc?

Thank you so much!

SivaRamaKrishnaNV · October 11, 2021, 6:28pm

Dear @yudong.wu,
Could you share the CUDA device query sample output

yudong.wu · October 11, 2021, 6:33pm

Hi @SivaRamaKrishnaNV,

Here it is:

./deviceQuery
./deviceQuery Starting…

CUDA Device Query (Runtime API) version (CUDART static linking)

Detected 2 CUDA Capable device(s)

Device 0: “Graphics Device”
CUDA Driver Version / Runtime Version 10.2 / 10.2
CUDA Capability Major/Minor version number: 7.5
Total amount of global memory: 7680 MBytes (8052998144 bytes)
(44) Multiprocessors, ( 64) CUDA Cores/MP: 2816 CUDA Cores
GPU Max Clock rate: 1440 MHz (1.44 GHz)
Memory Clock rate: 1440 Mhz
Memory Bus Width: 256-bit
L2 Cache Size: 4194304 bytes
Maximum Texture Dimension Size (x,y,z) 1D=(131072), 2D=(131072, 65536), 3D=(16384, 16384, 16384)
Maximum Layered 1D Texture Size, (num) layers 1D=(32768), 2048 layers
Maximum Layered 2D Texture Size, (num) layers 2D=(32768, 32768), 2048 layers
Total amount of constant memory: 65536 bytes
Total amount of shared memory per block: 49152 bytes
Total number of registers available per block: 65536
Warp size: 32
Maximum number of threads per multiprocessor: 1024
Maximum number of threads per block: 1024
Max dimension size of a thread block (x,y,z): (1024, 1024, 64)
Max dimension size of a grid size (x,y,z): (2147483647, 65535, 65535)
Maximum memory pitch: 2147483647 bytes
Texture alignment: 512 bytes
Concurrent copy and kernel execution: Yes with 3 copy engine(s)
Run time limit on kernels: No
Integrated GPU sharing Host Memory: No
Support host page-locked memory mapping: Yes
Alignment requirement for Surfaces: Yes
Device has ECC support: Disabled
Device supports Unified Addressing (UVA): Yes
Device supports Compute Preemption: Yes
Supports Cooperative Kernel Launch: Yes
Supports MultiDevice Co-op Kernel Launch: Yes
Device PCI Domain ID / Bus ID / location ID: 1 / 1 / 0
Compute Mode:
< Default (multiple host threads can use ::cudaSetDevice() with device simultaneously) >

Device 1: “Xavier”
CUDA Driver Version / Runtime Version 10.2 / 10.2
CUDA Capability Major/Minor version number: 7.2
Total amount of global memory: 28305 MBytes (29680066560 bytes)
( 8) Multiprocessors, ( 64) CUDA Cores/MP: 512 CUDA Cores
GPU Max Clock rate: 1109 MHz (1.11 GHz)
Memory Clock rate: 1109 Mhz
Memory Bus Width: 256-bit
L2 Cache Size: 524288 bytes
Maximum Texture Dimension Size (x,y,z) 1D=(131072), 2D=(131072, 65536), 3D=(16384, 16384, 16384)
Maximum Layered 1D Texture Size, (num) layers 1D=(32768), 2048 layers
Maximum Layered 2D Texture Size, (num) layers 2D=(32768, 32768), 2048 layers
Total amount of constant memory: 65536 bytes
Total amount of shared memory per block: 49152 bytes
Total number of registers available per block: 65536
Warp size: 32
Maximum number of threads per multiprocessor: 2048
Maximum number of threads per block: 1024
Max dimension size of a thread block (x,y,z): (1024, 1024, 64)
Max dimension size of a grid size (x,y,z): (2147483647, 65535, 65535)
Maximum memory pitch: 2147483647 bytes
Texture alignment: 512 bytes
Concurrent copy and kernel execution: Yes with 1 copy engine(s)
Run time limit on kernels: No
Integrated GPU sharing Host Memory: Yes
Support host page-locked memory mapping: Yes
Alignment requirement for Surfaces: Yes
Device has ECC support: Disabled
Device supports Unified Addressing (UVA): Yes
Device supports Compute Preemption: Yes
Supports Cooperative Kernel Launch: Yes
Supports MultiDevice Co-op Kernel Launch: Yes
Device PCI Domain ID / Bus ID / location ID: 0 / 0 / 0
Compute Mode:
< Default (multiple host threads can use ::cudaSetDevice() with device simultaneously) >

Peer access from Graphics Device (GPU0) → Xavier (GPU1) : No
Peer access from Xavier (GPU1) → Graphics Device (GPU0) : No

deviceQuery, CUDA Driver = CUDART, CUDA Driver Version = 10.2, CUDA Runtime Version = 10.2, NumDevs = 2
Result = PASS

Topic		Replies	Views
Tesla V100 doesn't support GPU-Metrics Collection Profiling Linux Targets	1	212	August 8, 2024
[nsys profile] gpu-metrics-devices fails with "Already under profiling" Profiling Linux Targets profiling	11	747	June 2, 2025
Unable to get gpu metrics on Quadro GV100 Profiling Linux Targets	3	567	January 5, 2024
Monitoring GPU utilization of dGPU on DRIVE AGX Pegasus DRIVE AGX Xavier General drive-devtools	12	2565	April 30, 2021
Nsys command line on agx pegasus Profiling DRIVE Targets drive-devtools	13	2156	November 16, 2021
Nsys profile doesn't collect tensor core utilization and the metrics about tensor active/SM instructions are not shown in the GUI Profiling Linux Targets cuda	5	858	February 10, 2024
Using nsys DRIVE AGX Orin General drive-devtools	5	149	October 30, 2025
Can not use GPU metrics on nsight system Profiling Linux Targets wsl	1	307	December 9, 2024
Gpu-metrics-set not found for GH200 Profiling Linux Targets	5	658	August 15, 2024
Running nsys profiling for GPU memory data on python Profiling Linux Targets	4	1230	June 25, 2024

How to collect GPU memory bandwidth usage data with nsys profiler on Pegasus?

Related topics