RTX 5060 Ti eGPU (AORUS AI BOX) — CUDA hard-lock on Linux via Thunderbolt 4

nvidia-bug-report.log.gz (668.9 KB)

Hi NVIDIA team,

I’m experiencing a complete system hard-lock when attempting any CUDA compute operation on my RTX 5060 Ti connected via Thunderbolt 4. This is the same issue reported in open-gpu-kernel-modules GitHub issues #974 and #979.

Hardware:

  • Laptop: ASUS ROG Strix G18 G814JIR (Intel i9-14900HX, TB4)

  • eGPU: Gigabyte AORUS RTX 5060 Ti AI BOX (Blackwell GB206, 16 GB)

  • Internal GPU: RTX 4070 Laptop (works perfectly)

Software:

  • Fedora 43, kernels 6.18.13 / 6.19.7 / 6.19.10

  • Drivers tested: 580.126.18 and 595.58.03 (open kernel modules)

What works: nvidia-smi, cuInit, cuDeviceGet, cuDeviceGetName, cuDeviceTotalMem — all return correct results for the eGPU.

What crashes: cuCtxCreate_v2() causes immediate system hard-lock (no kernel panic, no SysRq, power cycle required). The GSP firmware fails to initialize:

NVRM: GPU1 _kgspRpcRecvPoll: LibOS heartbeat timed out
NVRM: GPU1 kgspInitRm_IMPL: SET_GUEST_SYSTEM_INFO failed: 0xf
NVRM: GPU1 RmInitAdapter: Cannot initialize GSP firmware RM

Exhaustive workarounds tested (all failed):

  • iommu=off / iommu=pt / default

  • pci=realloc with various hpmmio sizes / no pci= args

  • NVreg_EnableGpuFirmware=0, NVreg_EnableHMM=0, NVreg_DynamicPowerManagement=0

  • NVreg_EnableResizableBar=0

  • pcie_ports=native, pcie_aspm=off, pcie_port_pm=off, thunderbolt.clx=0

  • NVreg_RegistryDwordsPerDevice with RmForceExternalGpu=1

  • udev d3cold_allowed=0 + power/control=on

  • blacklist nvidia_drm + nvidia_modeset (compute-only mode)

  • Power cycling the eGPU enclosure

  • Three different kernel versions

  • Two driver versions (580 and 595)

  • Proprietary modules → refused by Blackwell (“requires open kernel modules”)

Pattern from community reports:

  • TB5 host + Blackwell eGPU → works (roger-pmta, GitHub #979)

  • TB4 host + Blackwell eGPU → crashes (mihau81, rvn2p, myself, GitHub #974)

This appears to be a fundamental issue with GSP firmware communication latency through TB4 PCIe tunneling. The AORUS AI BOX is marketed for AI workloads but is unusable for CUDA compute on Linux.

Questions:

  1. Is there an internal timeline for a fix?

  2. Is the TB4 vs TB5 difference acknowledged as a factor?

  3. Are there any driver-internal debug flags we can test?

See also my detailed report on GitHub: RTX 5080 via Thunderbolt 5 eGPU: Hard lock on CUDA operations (nvidia-smi works at idle) · Issue #979 · NVIDIA/open-gpu-kernel-modules · GitHub

Thanks :)

@fanfanmgz not sure why I can’t see your post, but I’ve got it in an email notification.
Anyway, I’m like 99% sure that it’s a Linux TB module problem, rather than an NV driver problem. It is possible that there are 2 independent problems though, but definitely the TB Linux module needs to be fixed first to properly handle TB5 devices.

@morgwai666 Sorry about the deleted post — the forum flagged it as
too similar when I tried to re-edit.

You may be right that it’s a combination of both TB module and NV driver.
On my setup, BAR0 shows as 0M@0x0 without pci=realloc (TB resource
allocation issue), but even with BARs properly assigned, the GSP firmware
fails to initialize (heartbeat timeout) and any GPU write causes a
hard-lock. Read operations through TB4 work fine though.

Interestingly, pre-Blackwell GPUs (Ada/Ampere) work over TB4 on Linux,
and Blackwell works over TB4 on Windows — so both layers seem involved.

I’m happy to test any patches or debug builds if NVIDIA or the community
comes up with something. My full setup details and nvidia-bug-report.log.gz
are on GitHub #979:

echo 1 | sudo tee /sys/bus/pci/devices/0000:00:07.0/remove
sleep 2
echo 1 | sudo tee /sys/bus/pci/rescan
sleep 3
boltctl
lspci -nnk -s 04:00.0
nvidia-smi -L
sudo dmesg -T | grep -iE ‘MSI-X|RmInitAdapter|BAR|fallen|probe|AER|DPC’ | tail -60

it will force to rescan but any command sent to egpu fails by frezzing cpu

Hi everyone,

I am experiencing the exact same hard locks and kernel panics with the AORUS RTX 5060 Ti AI Box under Linux.

My Setup:

  • Host: Minisforum MS-S1 Max (AMD Ryzen AI Max+ 395, 128 GB RAM)

  • OS: Ubuntu 26.04 Desktop (Kernel 7.0.0-22-generic)

  • eGPU: GIGABYTE AORUS RTX 5060 Ti AI Box (Connected via USB4 / Thunderbolt)

  • Driver: NVIDIA 595.71.05 (Open Kernel Module)

Where it crashes: The system boots up fine, and nvidia-smi recognizes the RTX 5060 Ti correctly on the desktop. However, the moment I trigger a heavy compute payload—specifically running the Wan2.1 video generation model (14B parameter model)—the system hard locks and crashes completely.

Looking at my journalctl logs right before the crash, the PCIe link speed drops and the GPU completely vanishes from the bus due to register read failures:

Plaintext

kesä 24 20:18:37 otokka kernel: NVRM: GPU0 _intrServiceStallCommonCheckBegin: Failed GPU reg read : 0xffffffff. Check whether GPU is present on the bus
kesä 24 20:18:37 otokka kernel: NVRM: GPU0 nvAssertOkFailedNoLog: Assertion failed: GPU lost from the bus [NV_ERR_GPU_IS_LOST] (0x0000000F)

The link completely dies under heavy data transfer over the USB4 tunnel. Disabling ASPM (pcie_aspm=off) or using pci=noaer has not resolved the underlying hard lock.

I am not a developer or a Linux nerd, but I have tried several combinations with the help of AI to debug this. It is clearly a major compatibility issue between the new Blackwell (GB206) architecture eGPUs and the Linux USB4/Thunderbolt subsystem.

When will NVIDIA release stable production drivers or fixes for Linux that address these Blackwell eGPU bus timeouts?

±----------------------------------------------------------------------------------------+
| NVIDIA-SMI 595.71.05 Driver Version: 595.71.05 CUDA Version: 13.2 |
±----------------------------------------±-----------------------±---------------------+
| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|=========================================+========================+======================|
| 0 NVIDIA GeForce RTX 5060 Ti Off | 00000000:95:00.0 Off | N/A |
| 0% 35C P8 4W / 180W | 13MiB / 16311MiB | 0% Default |
| | | N/A |
±----------------------------------------±-----------------------±---------------------+

±----------------------------------------------------------------------------------------+
| Processes: |
| GPU GI CI PID Type Process name GPU Memory |
| ID ID Usage |
|=========================================================================================|
| 0 N/A N/A 7193 G /usr/bin/gnome-shell 2MiB |
±----------------------------------------------------------------------------------------+
OpenGL renderer string: AMD Radeon Graphics (radeonsi, strix_halo, ACO, DRM 3.64, 7.0.0-22-generic)

I got the RTX 5060 TI AI Box to run on a Dell Latitude 5450 (TB4/USB4.1, no internal GPU) with the following configuration:

BIOS:
Kernel DMA protection on
Secure Boot off

Linux:
Kernel 6.12 and 6.18 both work
nvidia-open driver 610

GRUB_CMDLINE_LINUX_DEFAULT='... pci=assign-busses,hpmmiosize=64M,hpmmioprefsize=384M,realloc,hpbussize=0x33 pcie_port_pm=off pcie_aspm=off intel_iommu=on iommu=pt thunderbolt.clx=0'

To prevent the hard-locking, the driver needs to stay in P0 mode.

/etc/modprobe.d/nvidia.conf

blacklist nvidia_drm
blacklist nvidia_modeset

softdep nvidia post: nvidia-uvm

options nvidia NVreg_DynamicPowerManagement=0x00

options nvidia NVreg_RegistryDwords="RMSetDynamicPowerManagement=0;PowerMizerEnable=0x1;PerfLevelSrc=0x2222;PowerMizerDefaultAC=0x1;"

I blacklisted drm and modeset because I only need the eGPU for compute. Other settings might work as well.

nvidia-persistenced service needs to keep running with

sudo systemctl enable nvidia-persistenced
sudo systemctl start nvidia-persistenced

/etc/systemd/system/nvidia-persistenced.service.d/override.conf

[Service]
ExecStart=
ExecStart=/usr/bin/nvidia-persistenced --user root --verbose
ExecStartPost=/bin/sh -c 'sleep 3; \
  echo on > /sys/bus/pci/devices/0000:36:00.0/power/control; \
  CAP_OFF=$(setpci -s 36:00.0 34.b); \
  LNKCTL=$(printf "%x" $((0x$CAP_OFF + 0x10))); \
  LNKCTL2=$(printf "%x" $((0x$CAP_OFF + 0x30))); \
  setpci -s 36:00.0 $LNKCTL2.b=0x04; \
  CURRENT=$(setpci -s 36:00.0 $LNKCTL.w); \
  NEW=$(printf "%04x" $((0x$CURRENT | 0x0020))); \
  setpci -s 36:00.0 $LNKCTL.w=$NEW'

36:00.0 is the parent address of the connection port in my case, 37:00.0 the pci address of the eGPU. Replace it with your address.

I did not perform an ablation yet, so some of the settings might not even be necessary. With these setting my AI Box stays at stable 16GT/s and does not crash the OS anymore.

This pattern is common with Blackwell eGPUs on Linux: enumeration and light queries work, but any real DMA/compute triggers a hard-lock. The fact that nvidia-smi and cuInit work but cuMemAlloc/cuLaunchKernel fails points to Thunderbolt PCIe tunneling or power-state transitions, not the GPU itself.

Things to test to isolate the layer:

1. Force PCIe gen down. Add pci=realloc,pcie_port_pm=off,nomsi or limit the downstream port to Gen3 in BIOS/ACPI if possible. Gen4/Gen5 over TB4 is where most of these locks occur.

2. Disable runtime PM for the eGPU root port. echo on > /sys/bus/pci/devices/…/power/control for the bridge and the GPU to rule out ASPM L1.2 exits.

3. Boot with nouveau blacklisted and no GSP fallback. With open modules, try nvidia.NVreg_OpenRmEnableUnsupportedGpus=1 only if the GPU is otherwise unsupported; for 5060 Ti it should not be needed.

4. Check for DMA translation issues. On some Intel TB hosts, iommu=pt or intel_iommu=off changes behavior. Not a fix, but a diagnostic.

5. Capture dmesg over netconsole or serial. Hard locks usually leave a trace in the host kernel, not the GPU driver.

If it stabilizes with Gen3 forced and ASPM off, you have a TB/PCIe signal-margin problem. If it still locks, it is more likely a driver/firmware issue and the nvidia-bug-report.log you attached is the right path. Cross-reference the open-gpu-kernel-modules issues you mentioned; they are tracking similar reports.

Sorry AF for giving you AI response, but after two weeks i have Aorus 5060 ti ai box working under linux as i want with firebat mn56 mini-pc (amd) and really clean CachyOS instalation.

So… this is what ai says about what we do:

Technical Overview: Hot-Plug eGPU Architecture

System Specifications:

  • OS: CachyOS (x86_64)

  • Kernel: Stock linux-cachyos (No custom patches, custom builds, or kernel boot parameters required)

Segment 1: One-Time System Configuration

This segment covers the permanent configuration setup on CachyOS. It ensures device authorization, installs official repository drivers, and sets up the udev gate to prevent boot-time instability.

1. Device Authorization (boltctl)

Authorizes the external USB4/Thunderbolt enclosure so it can communicate with the PCIe bus automatically:

Bash

sudo boltctl enroll c4148780-00b3-5ee2-ffff-ffffffffffff

2. Driver Installation

Uses the official stock NVIDIA drivers provided by the CachyOS repositories:

Bash

sudo pacman -Syu nvidia-dkms nvidia-utils lib32-nvidia-utils

3. The udev Boot-Gate Rule

Creates a safety mechanism in /etc/udev/rules.d/99-disable-egpu-bridge.rules to prevent system hangs during boot or reboot:

Bash

sudo bash -c 'cat << EOF > /etc/udev/rules.d/99-disable-egpu-bridge.rules
ACTION=="add", SUBSYSTEM=="pci", KERNEL=="0000:08:00.0", TEST!="/tmp/egpu_allow", RUN+="/bin/sh -c \"echo 1 > /sys/bus/pci/devices/0000:08:00.0/remove\""
EOF'

  • How it works: Because /tmp is wiped on every reboot, the authorization flag /tmp/egpu_allow does not exist during system startup. When udev detects the eGPU PCIe bridge (08:00.0), it immediately detaches it before KWin Wayland or the kernel can attempt initialization. The system boots cleanly on the integrated AMD Radeon 780M GPU.

Segment 2: Desktop Activation Script (start-egpu.sh)

This script is executed manually from the desktop when you want to activate the eGPU. It dynamically handles the hot-plug sequence, enforces PCIe Gen3 bus stability, and transitions KDE Plasma (KWin Wayland) to the RTX 5060 Ti.

The Complete Script (start-egpu.sh)

Bash

#!/usr/bin/env bash
set -e

echo "=== 1. Creating Authorization Flag ==="
sudo touch /tmp/egpu_allow

echo "=== 2. Rescanning PCIe Bus ==="
echo 1 | sudo tee /sys/bus/pci/rescan

echo "=== 3. Disabling ASPM & Forcing PCIe Gen3 (8 GT/s) ==="
# Disable ASPM power management to prevent signal latency
sudo setpci -s 08:00.0 CAP_EXP+10.w=0000
sudo setpci -s 09:00.0 CAP_EXP+10.w=0000

# Lock link speed to PCIe Gen3 (0003) on both parent bridge and GPU
sudo setpci -s 08:00.0 CAP_EXP+30.w=0003
sudo setpci -s 08:00.0 CAP_EXP+10.w=0020

sudo setpci -s 09:00.0 CAP_EXP+30.w=0003
sudo setpci -s 09:00.0 CAP_EXP+10.w=0020
sleep 1

echo "=== 4. Loading Official NVIDIA Drivers ==="
sudo modprobe nvidia nvidia_uvm nvidia_modeset nvidia_drm

echo "=== 5. Setting Persistence Mode & Power Limits ==="
sudo nvidia-smi -pm 1
sudo nvidia-smi -pl 180

echo "=== 6. Unbinding iGPU to Route Wayland 100% to eGPU ==="
echo 1 | sudo tee /sys/bus/pci/devices/0000:c7:00.0/remove

echo "=== Success! RTX 5060 Ti is Active ==="

What Makes This Architecture Optimal?

  1. Zero Kernel Modifications: No GRUB/Limine flags (pci=noaer, pci=nomsi), custom kernel builds, or initramfs (mkinitcpio) hooks are required.

  2. Dynamic PCIe Handshake Control: Enforcing Gen3 via setpci before calling modprobe nvidia bypasses the unstable Gen4/Gen5 auto-negotiation over USB4 cables.

  3. Clean Session Switching: Unbinding the iGPU (c7:00.0) signals KWin Wayland and Vulkan applications to run natively on the eGPU with full 180W power availability.

Here is the fix:

@fanfanmgz could you please mark this thread as solved so that AI bots and aura farmers stop claiming #979’s solution as their own? ;-)
Thx!

Stable gen4 is also posible. And works.

Script:

echo "=== 1. Ustawienie zgody udev i Skanowanie magistrali PCIe ==="
# Tworzymy zielone światło dla udev
sudo touch /tmp/egpu_allow

# Teren czysty - udev przepuści urządzenie bez usuwania!
echo 1 | sudo tee /sys/bus/pci/rescan > /dev/null
sleep 1


sudo bash -c 'echo performance > /sys/module/pcie_aspm/parameters/policy' 2>/dev/null || true

echo "# 2. Wyłączenie ASPM L0s/L1 na mostku i karcie"
sudo setpci -s 08:00.0 CAP_EXP+10.w=0000
sudo setpci -s 09:00.0 CAP_EXP+10.w=0000

echo "# 3. Wyłączenie L1 Substates (L1.1 / L1.2)"
sudo setpci -s 08:00.0 ECAP_1E+04.l=00000000 2>/dev/null || true
sudo setpci -s 09:00.0 ECAP_1E+04.l=00000000 2>/dev/null || true

echo "# 4. Wymuszenie PCIe Gen4 (0004) + Retrain (0020) na MOSTKU (08:00.0)"
sudo setpci -s 08:00.0 CAP_EXP+30.w=0004
sudo setpci -s 08:00.0 CAP_EXP+10.w=0020

echo "# 5. Wymuszenie PCIe Gen4 (0004) + Retrain (0020) na KARCIE (09:00.0)"
sudo setpci -s 09:00.0 CAP_EXP+30.w=0004
sudo setpci -s 09:00.0 CAP_EXP+10.w=0020
sleep 1

echo "=== 4. LnkCtl2 Registers & Active Link Status / Rejestry LnkCtl2 i stan magistrali: ==="
echo "Bridge / Mostek 08:00.0 (Target Speed): $(sudo setpci -s 08:00.0 CAP_EXP+30.w)"
echo "Card / Karta   09:00.0 (Target Speed): $(sudo setpci -s 09:00.0 CAP_EXP+30.w)"

echo "--- Weryfikacja aktywnego połączenia (LnkSta) ---"
echo "Mostek 08:00.0: $(sudo lspci -vv -s 08:00.0 | grep LnkSta | xargs)"
echo "Karta  09:00.0: $(sudo lspci -vv -s 09:00.0 | grep LnkSta | xargs)"

echo "=== 5. Loading NVIDIA Drivers / Ładowanie sterownika NVIDIA ==="
sudo modprobe nvidia nvidia_uvm nvidia_modeset nvidia_drm

echo "=== 6. Setting Power Limits & Disabling Auto-PM / Ustawianie limitów mocy i wyłączenie Auto-PM ==="
sudo nvidia-smi -pl 180

sudo nvidia-smi -lmc 28002,28002
sudo nvidia-smi -lgc 3090,3090

echo "=== 7. Verifying PCIe Link Speed / Sprawdzić prędkość linku: ==="
sudo lspci -vv -s 09:00.0 | grep -iE "LnkSta:|LnkCtl2:"

powerprofilesctl set performance

# Usuwamy flagę zezwolenia, aby ewentualny kolejny rescan/unplug znów był chroniony
sudo rm -f /tmp/egpu_allow

echo "=== 8. Detaching iGPU from PCIe bus / Odcinanie iGPU z szyny PCIe ==="
echo 1 | sudo tee /sys/bus/pci/devices/0000:c7:00.0/remove

Seems stable, as my firebat mn56 is one hour active without reboot and ran gothic gemake, then crimson desert, then days gone and now cp77 and i can still see 16gt/s link.

I’ve just noticed here that you have AMD 8745hs CPU, correct? It supports PCIe up to gen4 only, so I wonder if now you can just remove setting CAP_EXP+30.w=0004 or is it still necessary?

it is up to you to test. ;) my script works in gen4 and after more than two weeks of testing egpu i need to chill. xD i am not developer xD

It is proff that gen4 x4 works under linux. What you do with that is your. ;)

it’s a proof that it works specifically on your hardware mix, not in general. As you might have seen in the Framework thread, people are getting dramatically different results between different AMD CPU generations.

In my case (Strix) so far it didn’t even work with gen3 ;-] ;-(

Also, which kernel are you running now?

this is what is running

these kernel flags are additional for reBAR and no need to stable gen4

I couldn’t find any info regarding capability 1E: could you please point me to some documentation regarding this?

Also FYI:

This does much more than the comment says: to disable ASPM L0 and L1 you need to clear bits 0 and 1 only, while the quoted commands clear all 16 bits of the word, which may have unexpected consequences (I didn’t check what all the other bits do). It may be safer to add a mask to indeed clear only bits 0 and 1:

sudo setpci -s 08:00.0 CAP_EXP+10.w=0000:0003
sudo setpci -s 09:00.0 CAP_EXP+10.w=0000:0003

I cant. I really really asked AI to do that, first AI gave me unworking but very beatufull commands instead.

AI also want this as one line but then it dont work.

I was curious if i can get stable gen 3, so why i can not get stable gen4. Simple things dont work and some days ago i abandoned that gen4 ■■■■ because it dont work. With extracly THAT script, it works.

USE TRANSLATOR: polish → your language

Przepraszam że piszę po po Polsku ale język angielski nie jest moją mocną stroną. Udało mi się zrobić skrypt do uruchomienia egpu na gen3 i działał on sprawnie i przede wszystkim: stabilnie. Miał po dwie linijki kodu dla każdej modyfikacji parametrów komunikacji. I to działa. Poprosiłem sztuczną inteligencję o przerobienie tego tak aby zablokować połączenie na gen4 zamiast gen3 - skoro głównym problemem jest skakanie z gen3 na gen4 i gen4 na gen3. Każdy wynik sztucznej inteligencji z dwóch linii kodu robił jedną i to mi nigdy nie zadziałało. Wymusiłem na sztucznej inteligencji zrobienie tego dokładnie tak samo jak dla gen3, prosiłem o podwójne linie tak samo jak dla gen3. To 1E to bardzo mocno przerobione polecenie które domyślnie nie działało, tzn system nie przyjmował prostej ładnej komendy do wyłączenia stanów uśpienia L1 w egpu - dostawałem błąd w stylu brak takiej opcji.

Jestem pewien że ten skrypt da się bardzo łatwo dostosować do każdego sprzętu z pomocą modeli sztucznej inteligencji, jednak trzeba być bardzo stanowczym aby ta sztuczna inteligencja zrobiła dokładnie to samo. Sztuczna inteligencja sama wszystko upraszcza do tego stopnia że zamiast dwóch działających linijek kodu robi jedną (to nie działa) albo najprościej odsyła do flag jądra (to też nie działa).

Cały skrypt wygląda tak:

echo "=== 1. Ustawienie zgody udev i Skanowanie magistrali PCIe ==="
# Tworzymy zielone światło dla udev
sudo touch /tmp/egpu_allow

# Teren czysty - udev przepuści urządzenie bez usuwania!
echo 1 | sudo tee /sys/bus/pci/rescan > /dev/null
sleep 1


sudo bash -c 'echo performance > /sys/module/pcie_aspm/parameters/policy' 2>/dev/null || true

echo "# 2. Wyłączenie ASPM L0s/L1 na mostku i karcie"
sudo setpci -s 08:00.0 CAP_EXP+10.w=0000
sudo setpci -s 09:00.0 CAP_EXP+10.w=0000

echo "# 3. Wyłączenie L1 Substates (L1.1 / L1.2)"
sudo setpci -s 08:00.0 ECAP_1E+04.l=00000000 2>/dev/null || true
sudo setpci -s 09:00.0 ECAP_1E+04.l=00000000 2>/dev/null || true

echo "# 4. Wymuszenie PCIe Gen4 (0004) + Retrain (0020) na MOSTKU (08:00.0)"
sudo setpci -s 08:00.0 CAP_EXP+30.w=0004
sudo setpci -s 08:00.0 CAP_EXP+10.w=0020

echo "# 5. Wymuszenie PCIe Gen4 (0004) + Retrain (0020) na KARCIE (09:00.0)"
sudo setpci -s 09:00.0 CAP_EXP+30.w=0004
sudo setpci -s 09:00.0 CAP_EXP+10.w=0020
sleep 1

echo "=== 4. LnkCtl2 Registers & Active Link Status / Rejestry LnkCtl2 i stan magistrali: ==="
echo "Bridge / Mostek 08:00.0 (Target Speed): $(sudo setpci -s 08:00.0 CAP_EXP+30.w)"
echo "Card / Karta   09:00.0 (Target Speed): $(sudo setpci -s 09:00.0 CAP_EXP+30.w)"

echo "--- Weryfikacja aktywnego połączenia (LnkSta) ---"
echo "Mostek 08:00.0: $(sudo lspci -vv -s 08:00.0 | grep LnkSta | xargs)"
echo "Karta  09:00.0: $(sudo lspci -vv -s 09:00.0 | grep LnkSta | xargs)"

echo "=== 5. Loading NVIDIA Drivers / Ładowanie sterownika NVIDIA ==="
sudo modprobe nvidia nvidia_uvm nvidia_modeset nvidia_drm

echo "=== 6. Setting Power Limits & Disabling Auto-PM / Ustawianie limitów mocy i wyłączenie Auto-PM ==="
sudo nvidia-smi -pl 180
sudo nvidia-smi -lmc 14002,28002
sudo nvidia-smi -lgc 2000,3090

echo "=== 7. Verifying PCIe Link Speed / Sprawdzić prędkość linku: ==="
sudo lspci -vv -s 09:00.0 | grep -iE "LnkSta:|LnkCtl2:"

powerprofilesctl set performance

# Usuwamy flagę zezwolenia, aby ewentualny kolejny rescan/unplug znów był chroniony
sudo rm -f /tmp/egpu_allow

echo "=== 8. Detaching iGPU from PCIe bus / Odcinanie iGPU z szyny PCIe ==="
echo 1 | sudo tee /sys/bus/pci/devices/0000:c7:00.0/remove


Ale żeby to działało, zanim wogóle odpalę skrypt na pulpicie, potrzebne były modyfikacje:

komenda do zaakceptowania mojej karty graficznej
sudo boltctl enroll c4148780-00b3-5ee2-ffff-ffffffffffff

instalacja sterownika nvidia
sudo pacman -Syu nvidia-dkms nvidia-utils nvidia-settings

regula udev - ważna aby sterownik nvidia nie ładował się przed wykonaniem skryptu.

sudo bash -c 'cat << EOF > /etc/udev/rules.d/99-disable-egpu-bridge.rules

ACTION==“add”, SUBSYSTEM==“pci”, KERNEL==“0000:08:00.0”, TEST!=“/tmp/egpu_allow”, RUN+=“/bin/sh -c “echo 1 > /sys/bus/pci/devices/0000:08:00.0/remove””
EOF’

przeładowanie udev
sudo udevadm control --reload-rules