Hi! I have Asus Tuf 15 (FA507UI) with installed PopOS, ubuntu based linux kernel (6.8.0-76060800daily20240311-generic).
I’m running the system in the hybrid gpu mode, which means gpu should be in standby most of the time.
I experience random microfreezes with 100% CPU spikes. Usually 1 core is responsible, but the actual core may change from time to time. I visually saw the spikes in the performance monitor GUI, and recorded the timestamps with CPU number, for reasons I explain below.
First of all, I’m attaching the result of the nvidia-bug-report.sh. Interestingly, freezes were also occuring during the run of the utility, so this may help.
I used atop to detect which process causes the spikes, it turned out to be kworker/*:*-kac..., that’s why later I filter the processes causing the spikes by kworker.
I used while true; do sleep 0.3; kill -USR1 9101; done to trigger more frequent snapshots from atop → this resulted into a 400mb file, which I can attach if needed.
I also used /proc/sysrq-trigger to record the backtraces when the problem happens. Exact one-liner is this:
while true; do sleep 0.1; if [[ $(top -bn1 -o '%CPU' | tail -n+8 | head -n1 | awk '$9 ~ /100/ && $12 ~ /kworker/ {print $9,$12}' | wc -l) = 1 ]]; then echo l > /proc/sysrq-trigger; echo "trigger `date`"; fi; done
Since I knew the timestamps and cpu numbers, I could extract the stacktraces from dmesg corresponding to the particular kernel. See the file attached.
I would really appreciate any advice on how to fix those micro-freezes, and still keep the nvidia gpu available.
Thank you!
nvidia-bug-report.log.gz (741.1 KB)
backtrace.txt (15.5 KB)