Cpuinfo_cur_freq 总是不对

thor 执行 jetson_clocks 后,各cpu 的 cpuinfo_cur_freq 总是不对,好像被锁频了,帮忙看看是怎么回事?多谢。

过程如下:

nvidia@thor:~$ sudo jetson_clocks
Persistence mode is already Enabled for GPU 00000000:01:00.0.
All done.

nvidia@thor:~$ sudo cat /sys/devices/system/cpu/cpu0/cpufreq/cpuinfo_cur_freq
[ 2524.115767] cpufreq: cpu0,cur:650000,set:2601000,delta:1951000,set ndiv:289
650000

nvidia@thor:~$ sudo cat /sys/devices/system/cpu/cpu5/cpufreq/cpuinfo_cur_freq
[ 2218.314605] cpufreq: cpu4,cur:649000,set:2601000,delta:1952000,set ndiv:289
649000

nvidia@thor:~$ sudo jetson_clocks --show
SOC family:tegra264 Machine:NVIDIA Jetson AGX Thor Developer Kit
Online CPUs: 0-13, Offline CPUs:
cpu0: Governor=schedutil MinFreq=2601000 MaxFreq=2601000 CurrentFreq=2601000 IdleStates: WFI=0 cc7=
0
cpu1: Governor=schedutil MinFreq=2601000 MaxFreq=2601000 CurrentFreq=2601000 IdleStates: WFI=0 cc7=
0
cpu2: Governor=schedutil MinFreq=2601000 MaxFreq=2601000 CurrentFreq=2601000 IdleStates: WFI=0 cc7=
0
cpu3: Governor=schedutil MinFreq=2601000 MaxFreq=2601000 CurrentFreq=2601000 IdleStates: WFI=0 cc7=
0
cpu4: Governor=schedutil MinFreq=2601000 MaxFreq=2601000 CurrentFreq=2601000 IdleStates: WFI=0 cc7=
0
cpu5: Governor=schedutil MinFreq=2601000 MaxFreq=2601000 CurrentFreq=2601000 IdleStates: WFI=0 cc7=
0
cpu6: Governor=schedutil MinFreq=2601000 MaxFreq=2601000 CurrentFreq=2601000 IdleStates: WFI=0 cc7=
0
cpu7: Governor=schedutil MinFreq=2601000 MaxFreq=2601000 CurrentFreq=2601000 IdleStates: WFI=0 cc7=
0
cpu8: Governor=schedutil MinFreq=2601000 MaxFreq=2601000 CurrentFreq=2601000 IdleStates: WFI=0 cc7=
0
cpu9: Governor=schedutil MinFreq=2601000 MaxFreq=2601000 CurrentFreq=2601000 IdleStates: WFI=0 cc7=
0
cpu10: Governor=schedutil MinFreq=2601000 MaxFreq=2601000 CurrentFreq=2601000 IdleStates: WFI=0 cc7=
0
cpu11: Governor=schedutil MinFreq=2601000 MaxFreq=2601000 CurrentFreq=2601000 IdleStates: WFI=0 cc7=
0
cpu12: Governor=schedutil MinFreq=2601000 MaxFreq=2601000 CurrentFreq=2601000 IdleStates: WFI=0 cc7=
0
cpu13: Governor=schedutil MinFreq=2601000 MaxFreq=2601000 CurrentFreq=2601000 IdleStates: WFI=0 cc7=
0
gpu-gpc-0 MinFreq=1575000000 MaxFreq=1575000000 CurrentFreq=1575000000
gpu-nvd-0 MinFreq=1692000000 MaxFreq=1692000000 CurrentFreq=1692000000
EMC MinFreq=4266000000 MaxFreq=4266000000 CurrentFreq=4266000000
PVA0_VPS0: Online=0 MinFreq=0 MaxFreq=1215000000 CurrentFreq=1215000000
PVA0_AXI: Online=0 MinFreq=0 MaxFreq=909000000 CurrentFreq=909000000
FAN Dynamic Speed Control=nvfancontrol hwmon1_pwm1=255
FAN Dynamic Speed Control=nvfancontrol hwmon1_pwm1_enable=1
NV Power Mode: MAXN

并且执行 7zip 压力测试主频也没上去,如下所示:

nvidia@thor:~$ taskset -c 5 7z b -mmt1

7-Zip 23.01 (arm64) : Copyright (c) 1999-2023 Igor Pavlov : 2023-06-20
64-bit arm_v:8 locale=en_US.UTF-8 Threads:14 OPEN_MAX:16384

mt1
Compiler: 13.2.0 GCC 13.2.0
Linux : 6.8.12-tegra : #2 SMP PREEMPT Thu Apr 16 13:55:03 CST 2026 : aarch64
PageSize:4KB THP:always hwcap:FFFFFFFF:CRC32:SHA1:SHA2:AES:ASIMD hwcap2:801AF3FF
LE

1T CPU Freq (MHz): 648 649 648 649 649 649 649

RAM size: 125772 MB, # CPU hardware threads: 1 / 14 : 0020
RAM usage: 437 MB, # Benchmark threads: 1

                   Compressing  |                  Decompressing

Dict Speed Usage R/U Rating | Speed Usage R/U Rating
KiB/s % MIPS MIPS | KiB/s % MIPS MIPS

22: 2224 100 2172 2164 | 14589 100 1247 1246
23: 2096 100 2140 2136 | 14463 100 1254 1252
24: 1994 100 2149 2144 | 14323 100 1260 1257

25: 1932 100 2212 2207 14151 100 1261 1260
Avr: 2062 100 2168 2163 14382 100 1255 1254
Tot: 100 1712 1708

跑 7zip 压力同时监测 cpu5 的状态 ,tegrastats 的输出和 cpuinfo_cur_freq 总是不同,如下:

nvidia@thor:~$ sudo cat /sys/devices/system/cpu/cpu5/cpufreq/cpuinfo_cur_freq
[ 3181.268394] cpufreq: cpu4,cur:650000,set:2601000,delta:1951000,set ndiv:289
650000

nvidia@thor:~$ tegrastats
09-02-2026 17:21:17 RAM 3033/125773MB (lfb 8x4MB) SWAP 0/2048MB (cached 0MB) CPU [1%@2601,0%@2601,1%
@2601,0%@2601,0%@2601,100%@2601,0%@2601,0%@2601,0%@2601,0%@2601,0%@2601,0%@2601,0%@2601,0%@2601] cpu
@52.75C tj@54.156C soc012@53.25C gpu@54.156C soc345@52.531C VDD_GPU 2837mW/2837mW VDD_CPU_SOC_MSS 70
89mW/7089mW VIN_SYS_5V0 4492mW/4492mW

please check the oc1/2/3 event count on your board after issue happened and see if any of them got non zero value.

root@thor:/home/nvidia# sudo su -c ‘grep “” /sys/class/hwmon/hwmon*/oc*’
/sys/class/hwmon/hwmon2/oc1_event_cnt:0
/sys/class/hwmon/hwmon2/oc1_throt_en:0
/sys/class/hwmon/hwmon2/oc2_event_cnt:0
/sys/class/hwmon/hwmon2/oc2_throt_en:1
/sys/class/hwmon/hwmon2/oc3_event_cnt:1
/sys/class/hwmon/hwmon2/oc3_throt_en:1

但是另一台设备也一样的输出,cpuinfo_cur_freq 是正常的。

Is tegrastats result a normal one?

nvidia@thor:~$ tegrastats
09-03-2026 15:12:29 RAM 2356/125773MB (lfb 6x4MB) SWAP 0/2048MB (cached 0MB) CPU [1%@972,1%@972,1%@
972,1%@972,0%@972,10%@972,0%@972,0%@972,0%@1458,0%@1458,0%@972,0%@972,0%@972,0%@972] cpu@54.687C tj
@55.187C soc012@55.187C gpu@56.156C soc345@54.406C VDD_GPU 710mW/710mW VDD_CPU_SOC_MSS 5683mW/5683m
W VIN_SYS_5V0 5491mW/5491mW
09-03-2026 15:12:30 RAM 2356/125773MB (lfb 6x4MB) SWAP 0/2048MB (cached 0MB) CPU [0%@972,2%@972,0%@
972,0%@972,0%@972,0%@972,0%@972,0%@972,0%@972,0%@972,0%@972,0%@972,0%@972,0%@972] cpu@54.656C tj@55
.187C soc012@55.187C gpu@56.156C soc345@54.406C VDD_GPU 0mW/355mW VDD_CPU_SOC_MSS 5683mW/5683mW VIN
_SYS_5V0 5391mW/5441mW
09-03-2026 15:12:31 RAM 2356/125773MB (lfb 6x4MB) SWAP 0/2048MB (cached 0MB) CPU [0%@972,1%@972,0%@
972,0%@972,0%@972,0%@972,0%@972,0%@972,0%@972,0%@972,0%@972,0%@972,0%@972,0%@972] cpu@54.625C tj@55
.218C soc012@55.218C gpu@56.156C soc345@54.437C VDD_GPU 0mW/237mW VDD_CPU_SOC_MSS 5446mW/5604mW VIN
_SYS_5V0 5391mW/5424mW
09-03-2026 15:12:32 RAM 2356/125773MB (lfb 6x4MB) SWAP 0/2048MB (cached 0MB) CPU [1%@972,1%@972,0%@
972,0%@972,0%@972,0%@972,0%@972,0%@972,0%@972,0%@972,0%@972,0%@972,0%@972,0%@972] cpu@54.562C tj@55
.156C soc012@55.156C gpu@56.156C soc345@54.343C VDD_GPU 0mW/178mW VDD_CPU_SOC_MSS 5683mW/5624mW VIN
_SYS_5V0 5391mW/5416mW
09-03-2026 15:12:33 RAM 2356/125773MB (lfb 6x4MB) SWAP 0/2048MB (cached 0MB) CPU [1%@972,0%@972,0%@
972,0%@972,0%@972,0%@972,0%@972,0%@972,0%@972,1%@972,0%@972,0%@972,0%@972,0%@972] cpu@54.593C tj@55
.187C soc012@55.187C gpu@56.156C soc345@54.375C VDD_GPU 0mW/142mW VDD_CPU_SOC_MSS 5446mW/5588mW VIN
_SYS_5V0 5391mW/5411mW
09-03-2026 15:12:34 RAM 2356/125773MB (lfb 6x4MB) SWAP 0/2048MB (cached 0MB) CPU [1%@972,0%@972,0%@
972,0%@972,0%@972,0%@972,0%@972,0%@972,0%@972,2%@972,0%@972,3%@972,0%@972,0%@972] cpu@54.625C tj@55
.218C soc012@55.218C gpu@56.156C soc345@54.375C VDD_GPU 0mW/118mW VDD_CPU_SOC_MSS 5683mW/5604mW VIN
_SYS_5V0 5391mW/5408mW

但是:

nvidia@thor:~$ sudo cat /sys/devices/system/cpu/cpu4/cpufreq/cpuinfo_cur_freq
[ 497.196459] cpufreq: cpu4,cur:243000,set:972000,delta:729000,set ndiv:108
243000
nvidia@thor:~$ sudo cat /sys/devices/system/cpu/cpu5/cpufreq/cpuinfo_cur_freq
[ 503.242902] cpufreq: cpu4,cur:242000,set:972000,delta:730000,set ndiv:108
242000

Could you share me the full dmesg?

Is this test on NV devkit or custom board?

custom board.

full message from debug serial port:
minicom.cap.txt (184.6 KB)

Could you check if your OC3 event count is always 1 right after boot up without running any stress?

/sys/class/hwmon/hwmon2/oc1_event_cnt:0
/sys/class/hwmon/hwmon2/oc1_throt_en:0
/sys/class/hwmon/hwmon2/oc2_event_cnt:0
/sys/class/hwmon/hwmon2/oc2_throt_en:1
/sys/class/hwmon/hwmon2/oc3_event_cnt:1
/sys/class/hwmon/hwmon2/oc3_throt_en:1

And do you have a NV devkit there to put this SOM on it and see if the issue is gone?

  1. OC3 event count is always 1,even under strees-ng

$ stress-ng --cpu $(nproc)

  1. SOM 和 NVME 安装到官方 thor devkit 上,没有出现频率不对的问题,数据正常。但 OC2 和 OC3 都也出现了1次报警,如下:

/sys/class/hwmon/hwmon2/oc1_event_cnt:0
/sys/class/hwmon/hwmon2/oc1_throt_en:0
/sys/class/hwmon/hwmon2/oc2_event_cnt:1
/sys/class/hwmon/hwmon2/oc2_throt_en:1
/sys/class/hwmon/hwmon2/oc3_event_cnt:1
/sys/class/hwmon/hwmon2/oc3_throt_en:1

Could you review your board design for OC3/OC2 pin? And if the pinmux is correct?

还有其它报警会影响或扼制主频吗?OC2和OC3 在做压力测试时并没有增长。

要澄清的事情是在你沒有跑壓力測試的時候, 你還有沒有看到降頻這件事情.

如果有的話代表說OC pin可能不小心被asserted. 這個不一定會顯示在OC event上. 但還是有機會被降頻

而且在NV devkit上無法複製出來. 代表說custom board上的design或pinmux設定機會比較大