`nvidia-smi` Performance degredation

I recently updated my cuda drivers to 560.28.03.

I am now seeing an issue where nvidia-smi takes 10 seconds to run. strace shows that the overwhelming majority of this time is spent on this doozy of a call:

mmap(NULL, 51539607552, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS, -1, 0) = 0x7fa544400000

Is this intended behavior? Why does nvidia-smi need 52 Gb of memory? Is there any way to avoid this behavior? For expanded context, the syscalls before this seem to suggest it has something to do with nvidia-smi’s interaction with nvidia-persistenced

connect(19, {sa_family=AF_UNIX, sun_path="/var/run/nvidia-persistenced/socket"}, 37) = 0
rt_sigprocmask(SIG_SETMASK, ~[RTMIN RT_1], [], 8) = 0
prlimit64(0, RLIMIT_NOFILE, NULL, {rlim_cur=1024, rlim_max=1073741816}) = 0
mmap(NULL, 4294967296, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS, -1, 0) = 0x7f43d7600000
mmap(NULL, 51539607552, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS, -1, 0) = 0x7f37d7600000

Some subcommands of nvidia-smi do not exhibit this behavior. nvidia-smi dmon, for example, runs as fast as it always has, and without a large memory footprint.

I’m running Debian Trixie, with the latest versions of nvidia software as provided in https://developer.download.nvidia.com/compute/cuda/repos/debian12/x86_64/.
My machine has 4 A10Gs on it.

It doesn’t need 52GB of memory. It is mapping a virtual address (VA) space. Carving out a virtual reservation. this is a typical activity involved in “starting up a GPU”. When a GPU is truly/completely idle (no process running on it, persistence mode not enabled) it has no bearing on the machine VA space - it is invisible (mostly, leaving aside things like the BARs and things visible in I/O space.)

When you start up a GPU to do “real work”, then the CUDA runtime (the operating system for the GPU ) needs to make it possible for the various CUDA allocators to request allocations within a VA space that has already been reserved for such things. This is doing that VA space reservation.

Some requests that you might make of nvidia-smi require it to “start up the GPU”. Some don’t. This you may see variation in behavior, depending on what exactly you request on the command line.

And I’m not suggesting this is intended to be a detailed description for all possible questions that could be asked in this vein, or even all of the question you have asked, nor would I be able to provide such a description. For example I’m not suggesting that having persistence mode enabled makes every issue go away, although it demonstrably helps in some cases. If you don’t have persistence mode enabled (which might happen on a driver upgrade) you should probably re-enable it.

And if you observe that in otherwise identical scenarios that a particular nvidia-smi from a particular driver version involves very low run time whereas from another driver version involves very high run time, then that would probably be a candidate for a bug report. I’m running R565 drivers on a particular machine of mine and nvidia-smi runs in less than a second on a single GPU with 8GB of system ram.

Is that a cloud machine?

Thank you for such a fast reply!

The confusing wrinkle is that I do have persistence mode enabled, and this behavior remains even if I’m currently running processes on the GPU. Also, initializing cuda contexts and making device calls all happens with subsecond latency.

If I disable persistence mode, and verify that there are no active processes running, I see an additional ~300ms start up penalty, spent opening the underlying character devices and calling into the kernel. I’m very much guessing here, but it seems like that is the device initialization step you’re referring to, and a few hundred milliseconds is in line with my memories for the normal initialization latency. After this step is completed, I still see the 10 second 52 GiB allocation.

Also, I could be wrong about this, but my understand of device VM mapping is that the necessary mmap couldn’t be MAP_ANONYMOUS – it would need to map a fd corresponding to the device memory, right? Or, alternatively, a MAP_SHARED into a region that another process is brokering into a non-anonymous device map. My understanding of MAP_ANONYMOUS | MAP_PRIVATE is that this is allocating new memory (reserving VA space for, no hardware pages allocated) that does not correspond to any real backing (a.k.a ANONYMOUS), and is not shared with any other process (a.k.a PRIVATE).

I believe that is the case here. I never timed nvidia-smi with the old drivers, but it was certainly significantly faster. I will spin up a snapshot and benchmark directly before filing a bug report, but that’s my plan for right now.

Yes – a g5.12xlarge from aws

That’s exactly what I was suggesting. A generic VA reservation. I understand you find my viewpoints expressed here to be questionable, which is fine. I stated this:

Perhaps I shouldn’t have said anything at all. I’ll just restate what I said before: if you find evidence of a significant regression, my suggestion is to file a bug.

I would include that info in the bug report, or provide a link back to this forum post in the bug report.

Apologies if I came off as negative: I was only trying to understand what was happening, and didn’t mean to express any ill will. I appreciate you helping me try to understand my issue.

I’ve re-provisioned several test machines and re-installed all components of Nvidia’s stack directly, and I cannot reproduce the behavior I’m seeing. On identical fresh hosts, nvidia-smi completes with sub-second latency, never mapping more than a few hundred MBs of memory. It’s clear I’ve done something in the past to directly break my installation. If I had to guess, I preformed an unsupported dist-upgrade several weeks back, and must’ve made a mistake re-propagating i386 flags for apt sources?

As the issue will not reproduce on other hosts, and the fresh environments works well for me, I will neglect to submit a bug report, and chalk it up to PEBCAK.