This appears to be a Jetson R39.2 APT package issue, not an upstream source issue.
Upstream NVIDIA Container Toolkit v1.19.1 already has the corrected nvidia-cdi-refresh.service, but the Jetson R39.2 nvidia-container-toolkit-base package appears to ship an older/intermediate copy of that unit.
Observed package
- Jetson Linux: R39.2
- Ubuntu: 24.04.4 LTS
- Kernel:
6.8.12-1021-tegra
- Package:
nvidia-container-toolkit-base
- Version:
1.19.1-1
- Unit:
/etc/systemd/system/nvidia-cdi-refresh.service
- Unit SHA256:
798ece5e5812f525a60048ca2852bbfe0914790d294bae7de13b24e008f6c74c
Cause
The packaged unit appears to match intermediate upstream commit 2a80b6f / fe5a920, not the final upstream v1.19.1 unit.
The packaged unit is missing the final v1.19.1 restart/readiness behavior:
StartLimitBurst=5
StartLimitIntervalSec=10s
ExecStart readiness gate using nvidia-smi -L
Restart=on-failure
RestartSec=1s
Behavior
On boot, nvidia-cdi-refresh.service can run before the NVIDIA driver stack is ready.
When that happens, CDI generation can fail with failed to initialize nvml: Driver Not Loaded.
GPU CDI devices may then be missing until the service is restarted manually.
Fix tested
A local systemd drop-in restoring the final upstream v1.19.1 unit behavior resolves the CDI boot race.
After applying that drop-in, CDI generation succeeded and GPU CDI devices remained present across reboot testing.
Request
Please rebuild or update the Jetson R39.2 nvidia-container-toolkit-base package so nvidia-cdi-refresh.service matches the final upstream NVIDIA Container Toolkit v1.19.1 unit.
As an addendum, I isolated a second Jetson Linux R39.2 boot race involving the same nvidia-cdi-refresh.service. Its early nvidia-smi -L readiness probe can run before nvpmodel.service has applied the configured GPU power state. On affected boots, that probe appears to initialize the GPU and create the golden image context while the FBP/TPC masks are still 0/0 and the GPU maximum frequency is still 1020 MHz. When nvpmodel.service subsequently attempts to apply the configured 2/240 masks and 918 MHz limit, the power-gating state is already locked and nvpmodel exits with status 234. This appears to be the trigger for the intermittent nvpmodel.service failure already listed as a known Jetson Linux R39.2 issue.
Suggested possible fix
Make nvidia-cdi-refresh.service wait until nvpmodel.service has completed successfully once during the current boot before allowing the CDI readiness probe or CDI generation to touch the GPU.
Because nvpmodel.service is a Type=oneshot unit with RemainAfterExit=no, a small boot-scoped gate avoids rerunning nvpmodel if nvidia-cdi-refresh.path triggers another CDI refresh later in the same boot.
Create the gate:
# /etc/systemd/system/nvpmodel-ready.service
[Unit]
Description=Wait for nvpmodel initialization
Requires=nvpmodel.service
After=nvpmodel.service
Before=nvidia-cdi-refresh.service
[Service]
Type=oneshot
ExecStart=/usr/bin/true
RemainAfterExit=yes
Add a CDI service drop-in:
# /etc/systemd/system/nvidia-cdi-refresh.service.d/20-after-nvpmodel.conf
[Unit]
Requires=nvpmodel-ready.service
After=nvpmodel-ready.service
Apply the change:
sudo install -d -m 0755 \
/etc/systemd/system/nvidia-cdi-refresh.service.d
sudo systemctl daemon-reload
sudo reboot
This enforces the following boot order:
nvpmodel.service
↓
nvpmodel-ready.service
↓
nvidia-cdi-refresh.service
↓
nvidia-smi -L
↓
nvidia-ctk cdi generate
The dependency should be attached to the CDI refresh side, because CDI is the consumer that must not initialize the GPU before the Jetson power model has been established.
With this ordering applied, all consecutive boot tests (I ran 50 consecutive) completed without the failure of either the nvmodel.service or the nvidia-cdi-refresh.service.
This is only one possible solution, and the appropriate upstream implementation may differ. However, I believe the underlying root causes of both boot-time service races described above have been correctly isolated..
Hi,
Our environment (setup with JetPack 7.2) shows 1.19.1-1.
Which version do you observe in your environment?
$ apt show nvidia-container-toolkit-base
Package: nvidia-container-toolkit-base
Version: 1.19.1-1
Priority: optional
Section: utils
Source: nvidia-container-toolkit
Maintainer: NVIDIA CORPORATION <cudatools@nvidia.com>
Installed-Size: 24.9 MB
Breaks: nvidia-container-runtime (<= 3.5.0-1), nvidia-container-runtime-hook, nvidia-container-toolkit (<= 1.10.0-1)
Replaces: nvidia-container-runtime (<= 3.5.0-1), nvidia-container-runtime-hook
Homepage: https://github.com/NVIDIA/nvidia-container-toolkit
Download-Size: 4842 kB
APT-Manual-Installed: no
APT-Sources: https://repo.download.nvidia.com/jetson/common r39.2/main arm64 Packages
Description: NVIDIA Container Toolkit Base
Provides tools such as the NVIDIA Container Runtime and NVIDIA Container Toolkit CLI to enable GPU support in containers.
Thanks.
It shows the same:
jetadmin@elroy-jetson:~$ apt show nvidia-container-toolkit-basePackage: nvidia-container-toolkit-baseVersion: 1.19.1-1Priority: optionalSection: utilsSource: nvidia-container-toolkitMaintainer: NVIDIA CORPORATION cudatools@nvidia.comInstalled-Size: 24.9 MBBreaks: nvidia-container-runtime (<= 3.5.0-1), nvidia-container-runtime-hook, nvidia-container-toolkit (<= 1.10.0-1)Replaces: nvidia-container-runtime (<= 3.5.0-1), nvidia-container-runtime-hookHomepage: Download-Size: 4,842 kBAPT-Manual-Installed: noAPT-Sources: https://repo.download.nvidia.com/jetson/common r39.2/main arm64 PackagesDescription: NVIDIA Container Toolkit BaseProvides tools such as the NVIDIA Container Runtime and NVIDIA Container Toolkit CLI to enable GPU support in containers.
The issue is that package published by the apt-get repo doesn’t match up with the source for the github release (https://github.com/NVIDIA/nvidia-container-toolkit/tree/v1.19.1) , and appears to be from a previous commit, prior to the service restart features being added back to the nvidia-cdi-refresh.service, as noted above. This causes a race condition where the nvidia-cdi-refresh.service can fail on boot and doesn’t recover. Simply adapting the upstream code for the systemd service definition from the github repo fixes that race condition, so the fix there is to simply adopt the service definition from the upstream source, and integrate it into the presented Jetson package.
That service is also connected directly with a mentioned known issue for R39.2, where the nvpmodel.service can intermittently crash due do an entirely different race condition. I have also provided an isolated analysis of that issue and provided a suggested fix for it as well. I did look, but I have not found any other solutions in the forum for that specific race issue, and it was directly mentioned in the release notes as known.
Hi,
Sorry for the late update.
Let’s focus on the CDI issue first.
Suppose this issue has a failure rate, right?
We test several times (reboot) on Orin Nano, but it doesn’t hit the race condition.
How many time you tried to hit the issue?
$ sha256sum /etc/systemd/system/nvidia-cdi-refresh.service
798ece5e5812f525a60048ca2852bbfe0914790d294bae7de13b24e008f6c74c /etc/systemd/system/nvidia-cdi-refresh.service
$ systemctl status nvidia-cdi-refresh.service
○ nvidia-cdi-refresh.service - Refresh NVIDIA CDI specification file
Loaded: loaded (/etc/systemd/system/nvidia-cdi-refresh.service; enabled; preset: enabled)
Active: inactive (dead) since Fri 2026-07-24 04:22:16 UTC; 1min 44s ago
TriggeredBy: ● nvidia-cdi-refresh.path
Main PID: 902 (code=exited, status=0/SUCCESS)
CPU: 344ms
Jul 24 04:22:15 tegra-ubuntu nvidia-ctk[902]: time="2026-07-24T04:22:15Z" level=info msg="Selecting /opt/nvidia/l4t-gpu-libs/openrm/libcuda_instrumentation.so as /opt/nvi>
Jul 24 04:22:15 tegra-ubuntu nvidia-ctk[902]: time="2026-07-24T04:22:15Z" level=info msg="Selecting /opt/nvidia/l4t-gpu-libs/openrm/libcuda.so.1.1 as /opt/nvidia/l4t-gpu->
Jul 24 04:22:15 tegra-ubuntu nvidia-ctk[902]: time="2026-07-24T04:22:15Z" level=info msg="Selecting /etc/nv_tegra_release as /etc/nv_tegra_release"
Jul 24 04:22:15 tegra-ubuntu nvidia-ctk[902]: time="2026-07-24T04:22:15Z" level=warning msg="Failed to get soname symlinks for {HostPath:/usr/lib/aarch64-linux-gnu/gstrea>
Jul 24 04:22:15 tegra-ubuntu nvidia-ctk[902]: time="2026-07-24T04:22:15Z" level=warning msg="Failed to get soname symlinks for {HostPath:/usr/lib/aarch64-linux-gnu/gstrea>
Jul 24 04:22:15 tegra-ubuntu nvidia-ctk[902]: time="2026-07-24T04:22:15Z" level=warning msg="Failed to get soname symlinks for {HostPath:/usr/lib/aarch64-linux-gnu/gstrea>
Jul 24 04:22:15 tegra-ubuntu nvidia-ctk[902]: time="2026-07-24T04:22:15Z" level=warning msg="Failed to locate symlink /usr/lib/aarch64-linux-gnu/tegra"
Jul 24 04:22:15 tegra-ubuntu nvidia-ctk[902]: time="2026-07-24T04:22:15Z" level=info msg="Generated CDI spec with version 0.7.0"
Jul 24 04:22:16 tegra-ubuntu systemd[1]: nvidia-cdi-refresh.service: Deactivated successfully.
Jul 24 04:22:16 tegra-ubuntu systemd[1]: Finished nvidia-cdi-refresh.service - Refresh NVIDIA CDI specification file.
Thanks.
Yes. The CDI failure was intermittent, but it reproduced very frequently in my testing.
For hardware context, these are Jetson Orin Nano Super modules. I have four of them. The quantified failure-rate test below was performed on one module, with the stock R39.2 package/service state and without adding retries or modifying the units during the sample.
Across 20 consecutive reboots:Results:
nvidia-cdi-refresh.service failed on 11/20 boots (55%)
nvpmodel.service failed on 13/20 boots (65%)
- both failed on 5/20 boots (25%)
- only 1/20 boots completed without either failure
So for the CDI race you asked about, my measured failure rate was 55% (11 of 20 boots).
These appear to be two separate races. After restoring the readiness/restart behavior from the final upstream v1.19.1 nvidia-cdi-refresh.service, I stopped reproducing the CDI failure in the subsequent validation. I then addressed the separate nvpmodel.service ordering race, after which I ran 50 consecutive reboots without a failure of either service.
The SHA256 you posted is also the same packaged service file I tested:
798ece5e5812f525a60048ca2852bbfe0914790d294bae7de13b24e008f6c74c
Although I primarily tested on one module, I was able to validate the same issue occurred on all four of my devices at (seemingly) similar rates.
I have not had a single issue with either service since implementing the fixes I suggested above, one of which addresses a ‘known issue’ in the release notes.