I recently updated my cuda drivers to 560.28.03.
I am now seeing an issue where nvidia-smi takes 10 seconds to run. strace shows that the overwhelming majority of this time is spent on this doozy of a call:
mmap(NULL, 51539607552, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS, -1, 0) = 0x7fa544400000
Is this intended behavior? Why does nvidia-smi need 52 Gb of memory? Is there any way to avoid this behavior? For expanded context, the syscalls before this seem to suggest it has something to do with nvidia-smi’s interaction with nvidia-persistenced
connect(19, {sa_family=AF_UNIX, sun_path="/var/run/nvidia-persistenced/socket"}, 37) = 0
rt_sigprocmask(SIG_SETMASK, ~[RTMIN RT_1], [], 8) = 0
prlimit64(0, RLIMIT_NOFILE, NULL, {rlim_cur=1024, rlim_max=1073741816}) = 0
mmap(NULL, 4294967296, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS, -1, 0) = 0x7f43d7600000
mmap(NULL, 51539607552, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS, -1, 0) = 0x7f37d7600000
Some subcommands of nvidia-smi do not exhibit this behavior. nvidia-smi dmon, for example, runs as fast as it always has, and without a large memory footprint.
I’m running Debian Trixie, with the latest versions of nvidia software as provided in https://developer.download.nvidia.com/compute/cuda/repos/debian12/x86_64/.
My machine has 4 A10Gs on it.