RTX 5090 not working as eGPU on ubuntu 22.04

So I’ve got an eGPU setup for my 5090 via Razer Core X V2. I tested this setup on a Windows machine (on a Dell alienware) and the GPU was duly recognized and was working (although training speed was awfully slow but that’s another topic).

However, I’m having issues getting it to work on my ubuntu 22.04 running on a Dell XPS 13 9340 (which is equipped with thunderbolt 4). I’ve tried various cuda driver & toolkit versions, varying from 570-580 and 12.8 - 13.0 for toolkit. But I don’t think the issue is here as I’ll explain below.

Just as a note: I have an eGPU solution for a RTX 3090 on this laptop which is working perfectly fine.

My default grub setup was this:

GRUB_CMDLINE_LINUX_DEFAULT=“quiet splash i915.enable_psr=0”

I encounter the following error:

NVRM: This PCI I/O region assigned to your NVIDIA device is invalid
nvidia: probe of 0000:03:00.0 failed with error -1
NVRM: None of the NVIDIA devices were initialized

And nvidia-smi fails with: NVIDIA-SMI has failed because it couldn’t communicate with the NVIDIA driver. Make sure that the latest NVIDIA driver is installed and running.

Similar to this issue: NVIDIA RTX 5090 Not Detected by nvidia-smi on Ubuntu Server 24.04 - #12 by ys12345678

One of the top suggestions given in the above thread is disabling and re-enabling resizable BAR. Unfortunately my XPS 13 9340 doesn’t offer this option in its BIOS settings.

This other thread NVRM: This PCI I/O region assigned to your NVIDIA device is invalid suggests trying pci=realloc. This doesn’t result in any changes for me (meaning same error messages). What (somewhat) helps is the following:

GRUB_CMDLINE_LINUX_DEFAULT=“quiet splash i915.enable_psr=0 pci=realloc=off pci=nocrs”

Like this I no longer have the BAR issue and nvidia-smi now gives a No devices were found. Logs, however, show a different error, namely: Error nvidia-drm failed to allocate NvKmsKapiDevice.

The issue appears to be that the GPU is allocated only 256MB BAR space, while the RTX 5090
apparently requires much more (?).

I’ve also tried enabling BAR reallocation with hotplug memory reservation:

GRUB_CMDLINE_LINUX_DEFAULT=“quiet splash i915.enable_psr=0 pci=realloc pci=assign-busses pci=hpmmiosize=128M,hpmemsize=4096M”

Doesn’t help, unfortunately.

Here’s where I’m currently stuck, and the 5090 is still not usable.

Hi @georg9alem,

on eGPU topics I am out. I have no idea how PCI channel assignment works through TB or how/if ReBAR even works through TB.

[addition] I reached out internally if we have some experts on the topic, since we do test eGPUs at least in our general test farms. A quick search in our issue database at least did not show anything standing out either that might help here.

Hey @MarkusHoHo ,

thanks a lot for taking the time to also reach out internally.

I see, so I suppose there’s not much you guys could suggest me to try out?

Hi @georg9alem ,

Can you please capture a NVIDIA bug report after you hit the initialization errors. Please run $sudo nvidia-bug-report.sh and attach the generated nvidia-bug-report.log.gz to this thread. Thanks.

@abchauhan Thanks for the follow-up. I’ve attached the bug report.

nvidia-bug-report.log (4.7 MB)

Thank you. The logs show GPU initialization failures. Filed Bug #5618878 internally to request Engineering review of the logs.

I will share Engineering feedback when available.

@abchauhan I have an update.

I tried unbinding & re-scaning the TB bridge, and it works now. So I did the following:

echo 1 > /sys/bus/pci/devices/$TB_BRIDGE/remove
echo 1 > /sys/bus/pci/rescan

Now nvidia-smi works!



Thu Oct 30 15:47:12 2025
±----------------------------------------------------------------------------------------+
| NVIDIA-SMI 580.95.05              Driver Version: 580.95.05      CUDA Version: 13.0     |
±----------------------------------------±-----------------------±---------------------+
| GPU  Name                 Persistence-M | Bus-Id          Disp.A | Volatile Uncorr. ECC |
| Fan  Temp   Perf          Pwr:Usage/Cap |           Memory-Usage | GPU-Util  Compute M. |
|                                         |                        |               MIG M. |
|=========================================+========================+======================|
|   0  NVIDIA GeForce RTX 5090        Off |   00000000:03:00.0 Off |                  N/A |
| 34%   38C    P0             56W /  600W |       0MiB /  32607MiB |      2%      Default |
|                                         |                        |                  N/A |
±----------------------------------------±-----------------------±---------------------+

±----------------------------------------------------------------------------------------+
| Processes:                                                                              |
|  GPU   GI   CI              PID   Type   Process name                        GPU Memory |
|        ID   ID                                                               Usage      |
|=========================================================================================|
|  No running processes found                                                             |
±----------------------------------------------------------------------------------------+

It’s very unstable though. As you can see its consuming 56W without even being used, and I also notice that my CPU becomes busy.

Also, initially torch.cuda.is_available()returns True but after a bit it fails and returns False with the typical cuda error:

>>> torch.cuda.is_available()
/home/rejnald/.pyenv/versions/neoface-training/lib/python3.10/site-packages/torch/cuda/init.py:138: UserWarning: CUDA initialization: CUDA unknown error - this may be due to an incorrectly set up environment, e.g. changing env variable CUDA_VISIBLE_DEVICES after program start. Setting the available devices to be zero. (Triggered internally at ../c10/cuda/CUDAFunctions.cpp:108.)
return torch._C._cuda_getDeviceCount() > 0
False

Following errors can be observed in dmesg:

[ 1589.292316] nvidia-modeset: Unloading
[ 1589.363223] nvidia-nvlink: Unregistered Nvlink Core, major device number 509
[ 1589.527467] nvidia-nvlink: Nvlink Core is being initialized, major device number 509
[ 1589.534466] nvidia 0000:03:00.0: vgaarb: changed VGA decodes: olddecodes=none,decodes=none:owns=none
[ 1589.540609] NVRM: loading NVIDIA UNIX Open Kernel Module for x86_64 580.95.05 Release Build (dvs-builder@U22-I3-B17-02-5) Tue Sep 23 09:55:41 UTC 2025
[ 1589.563626] nvidia-modeset: Loading NVIDIA UNIX Open Kernel Mode Setting Driver for x86_64 580.95.05 Release Build (dvs-builder@U22-I3-B17-02-5) Tue Sep 23 09:42:01 UTC 2025
[ 1589.571092] [drm] [nvidia-drm] [GPU ID 0x00000300] Loading driver
[ 1590.952743] [drm] Initialized nvidia-drm 0.0.0 20160202 for 0000:03:00.0 on minor 1
[ 1590.954310] nvidia 0000:03:00.0: [drm] Cannot find any crtc or sizes
[ 1591.015814] [drm] [nvidia-drm] [GPU ID 0x00000300] Unloading driver
[ 1591.324128] nvidia-modeset: Unloading
[ 1591.394616] nvidia-nvlink: Unregistered Nvlink Core, major device number 509

Also, in general my system starts lagging every few seconds (by about ~300ms or so), and sometimes even unexpectedly logs off. So very unstable.

This settles down after a few minutes though, the lag is gone, and watt usage goes down significantly:



Thu Oct 30 15:57:26 2025
±----------------------------------------------------------------------------------------+
| NVIDIA-SMI 580.95.05              Driver Version: 580.95.05      CUDA Version: 13.0     |
±----------------------------------------±-----------------------±---------------------+
| GPU  Name                 Persistence-M | Bus-Id          Disp.A | Volatile Uncorr. ECC |
| Fan  Temp   Perf          Pwr:Usage/Cap |           Memory-Usage | GPU-Util  Compute M. |
|                                         |                        |               MIG M. |
|=========================================+========================+======================|
|   0  NVIDIA GeForce RTX 5090        Off |   00000000:03:00.0 Off |                  N/A |
|  0%   39C    P8              6W /  600W |       7MiB /  32607MiB |      0%      Default |
|                                         |                        |                  N/A |
±----------------------------------------±-----------------------±---------------------+

±----------------------------------------------------------------------------------------+
| Processes:                                                                              |
|  GPU   GI   CI              PID   Type   Process name                        GPU Memory |
|        ID   ID                                                               Usage      |
|=========================================================================================|
|  No running processes found                                                             |
±----------------------------------------------------------------------------------------+

I’m attaching a new nvidia bug report

nvidia-bug-report.log (38.6 MB)

RTX 5090 AORUS AI BOX — Thunderbolt 4 Hard Lock on Any GPU Write (Bug #5618878)

Hi, adding another data point to this thread. I have the same fundamental issue as @georg9alem but with different hardware and more detailed testing.

My Setup

  • Host: Intel NUC 15 Pro+ (Arrow Lake, Core Ultra 9 288V), 96GB RAM
  • eGPU: GIGABYTE AORUS GeForce RTX 5090 AI BOX (TB4)
  • OS: Proxmox VE 9.1.5 (Debian Bookworm, bare metal, no desktop)
  • Kernel: 6.17.9-1-pve
  • Connection: Thunderbolt 4 (Intel Gen14 controller)

Drivers Tested

  1. 580.142 closed modules — crash
  2. 590.48.01 open modules (stock) — crash
  3. 590.48.01 open modules with PR #984 patch (RmForceExternalGpu=1) — detection fixed, crash persists

Kernel Parameters

pci=hpbussize=0x10,hpmmiosize=64M,hpmmioprefsize=384M,realloc
pcie_port_pm=off pcie_aspm.policy=performance thunderbolt.clx=0

What Works

  • GPU detected on PCI bus: 04:00.0 VGA compatible controller: NVIDIA Corporation GB202 [GeForce RTX 5090]
  • Driver loads, nvidia-smi shows GPU at idle (77W, 56°C, 32GB VRAM)
  • Read-only nvidia-smi queries work indefinitely

What Crashes

Any command that writes to the GPU causes an instant hard lock (no panic, no Xid, no SysRq, requires power cycle):

  • nvidia-smi -pm 1 → hard lock
  • nvidia-smi -pl 460 → hard lock
  • nvidia-smi -lgc 2407,2407 → hard lock
  • Any CUDA operation → hard lock

This is 100% reproducible. The boundary is clear: read = OK, write = instant death.

PR #984 Feedback

I built 590.48.01 open kernel modules from GitHub source with the RmForceExternalGpu registry key patch. The patch works perfectly for detection — the driver loads cleanly and recognizes the eGPU without any manual
PCI remove/rescan. I strongly recommend merging it.

However, the hard lock on write operations is a separate, deeper issue — likely in the PCIe power state management layer when communicating through the Thunderbolt bridge.

Comparison with Other Reports

  • @georg9alem: RTX 5090 + Razer Core X V2 + Dell XPS 13 (TB4) — same BAR/detection issues
  • GitHub #979: RTX 5080 + Sonnet Breakaway (TB5) — @roger-pmta got it working with clock locking. On my TB4 setup, even clock locking commands crash the
    system.
    This suggests TB4 may handle GPU power state transitions differently than TB5.

Request

Could the engineering team investigating bug #5618878 look at why write operations (persistence mode, power limit, clock locking) cause hard locks while read operations work? The read/write boundary seems like a
useful clue — it points to GPU state transitions triggering a fatal PCIe link event through the Thunderbolt bridge.

Happy to run any diagnostic commands or attach nvidia-bug-report.sh output (as long as they don’t write to the GPU 😅).

btw, I eventually got it work by plugging and unplugging an HMI cable directly to the GPU after unbinding & re-scaning the TB bridge (which makes the eGPU available to begin with).

This stabilizes the egpu from 56W → 6W usage and then I’m able to use it w/o any issues.

It still (seemingly) randomly drops off sometimes but it’s not that frequent and I can live with it for now.