Thor reboot failed

硬件环境:thor 核心板+定制底板
软件环境:jetpack 7.1
问题:执行 sudo reboot 后卡死,之前是可以正常重启的,突然就重启卡死了,每次都会出现。报错情况如下,请帮忙分析,多谢。

nvidia@thor:~$ sudo reboot
[sudo] password for nvidia:
Sorry, try again.
[sudo] password for nvidia:
nvidia@thor:~$
▒▒INFO: END TASK:MB▒▒
INFO: enter idle task.
INFO: END TASK:MB▒▒
INFO: enter idle task.
▒▒[ 261.007457] kauditd_printk_skb: 129 callbacks suppressed
[ 261.007463] audit: type=1305 audit(72433.540:675): op=set audit_pid=0 old=782 auid=4294967295 ses=4294967295 subj=unconfined res=1
[ 261.011368] audit: type=1131 audit(72433.544:676): pid=1 uid=0 auid=4294967295 ses=4294967295 subj=unconfined msg=‘unit=auditd comm=“systemd” exe=“/usr/lib/systemd/systemd” hostname=? addr=? terminal=? res=success’
[ 261.029860] audit: type=1131 audit(72433.544:677): pid=1 uid=0 auid=4294967295 ses=4294967295 subj=unconfined msg=‘unit=systemd-tmpfiles-setup comm=“systemd” exe=“/usr/lib/systemd/systemd” hostname=? addr=? terminal=? res=success’
[ 261.089371] audit: type=1131 audit(72433.624:678): pid=1 uid=0 auid=4294967295 ses=4294967295 subj=unconfined msg='unit=systemd-fsck@dev-nvme0n1p4 [ 261.114910] audit: type=1131 audit(72433.648:679): pid=1 uid=0 auid=4294967295 ses=4294967295 subj=unconfined msg='unit=systemd-tmpfiles-setup-dev [ 261.121678] audit: type=1131 audit(72433.648:680): pid=1 uid=0 auid=4294967295 ses=4294967295 subj=unconfined msg=‘unit=systemd-tmpfiles-setup-dev-[ 261.149208] audit: type=1334 audit(72433.684:681): prog-id=35 op=UNLOADinal=? res=success’
[ 261.207460] audit: type=1131 audit(72433.740:682): pid=1 uid=0 auid=4294967295 ses=4294967295 subj=unconfined msg=‘unit=systemd-remount-fs comm=“sy[ 261.213449] audit: type=1130 audit(72433.740:683): pid=1 uid=0 auid=4294967295 ses=4294967295 subj=unconfined msg='unit=systemd-reboot comm=“systemd” exe=”/usr/lib/systemd/systemd" hostname=? addr=? terminal=? res=success’
[ 261.233341] audit: type=1131 audit(72433.740:684): pid=1 uid=0 auid=4294967295 ses=4294967295 subj=unconfined msg=‘unit=systemd-reboot comm=“systemd” exe=“/usr/lib/systemd/systemd” hostname=? addr=? terminal=? res=success’
[ 261.305681] watchdog: watchdog0: watchdog did not stop!
[ 261.633303] nvsciipc nvsciipc: nvipc: Shutting down
[ 261.633622] nvsciipc: Unloaded module
▒▒[TS:282193579306]RmDeInit completed successfully
▒▒[ 262.803286] tegra-mc 8108020000.memory-controller: pcie1w: non-secure write @0x0000000000000000: EMEM address decode error (EMEM decode error)
[ 262.814935] tegra-mc 8108020000.memory-controller: pcie1w: non-secure write @0x0000000000000200: EMEM address decode error (EMEM decode error)
[ 263.875895] arm-smmu-v3 8105000000.iommu: CMD_SYNC timeout at 0x0001c7cf [hwprod 0x0001c7d0, hwcons 0x0001c7cf]
▒▒Rest▒▒759▒▒Rese▒▒96] ▒▒t re▒▒rebo▒▒ques▒▒ot: ▒▒ted
Reb▒▒arti▒▒ooti▒▒ng s▒▒ng s▒▒yste▒▒yste▒▒m
▒▒m …
tegra264_pcie_rp_deinit PCIE C1 L2 timeout, LTSSM_STATE=0x4011b
wait_flush_done: flush timed out, mask=4, reg=a8024b08
wait_flush_done: flush timed out, mask=4, reg=a8024b08

FATAL ERROR [FILE=platform/drivers/pg/soc/pg-soc.c, ERR_UID=203]: MC flush failed for partition pcie_c1
firmware tag: 13cc6a25acfdf2fe601c-0c7e5cbd77e
r0_usr 0x500f6cda
r1_usr 0x000000cb
r2_usr 0x500b0690
r3_usr 0x50217e84
r4_usr 0x500221d4
r5_usr 0x500b0690
r6_usr 0x000000cb
r7_usr 0x500f6cda
r8_usr 0x50217e84 r8_fiq 0x00000000
r9_usr 0x00000003 r9_fiq 0x00000000
r10_usr 0x10101010 r10_fiq 0x00000000
r11_usr 0x11111111 r11_fiq 0x00000000
r12_usr 0x50217af0 r12_fiq 0x00000000
sp_usr 0x50217e68 sp_fiq 0x50050880 sp_irq 0x50050c00
sp_svc 0x50051400 sp_abt 0x50050800 sp_und 0x50050770
lr_usr 0x50023543 lr_fiq 0x00000000 lr_irq 0x50020922
lr_svc 0x50037469 lr_abt 0x00000000 lr_und 0x500221d8
pc 0x500221d4
spsr 0x20000010
fpscr 0x80000010
00.0: base: 50000000 size: 00200000 XN P_RW U_RW Non-shareable outer: WB, no WA inner: WB, no WA
00.1: base: 50200000 size: 00200000 XN P_RW U_RW Non-shareable outer: WB, no WA inner: WB, no WA
00.2: base: 50400000 size: 00200000 XN P_RW U_RW Non-shareable outer: WB, no WA inner: WB, no WA
00.3: base: 50600000 size: 00200000 XN P_RW U_RW Non-shareable outer: WB, no WA inner: WB, no WA
00.4: base: 50800000 size: 00200000 XN P_RW U_RW Non-shareable outer: WB, no WA inner: WB, no WA
01: base: 80000000 size: 80000000 XN P_RW U_RW Shareable strongly-ordered
02: base: ffff0000 size: 00000040 X P_RO U_NA Non-shareable outer: WB, no WA inner: WB, no WA
03.2: base: 50020000 size: 00010000 X P_RO U_RO Non-shareable outer: WB, no WA inner: WB, no WA
03.3: base: 50030000 size: 00010000 X P_RO U_RO Non-shareable outer: WB, no WA inner: WB, no WA
03.4: base: 50040000 size: 00010000 X P_RO U_RO Non-shareable outer: WB, no WA inner: WB, no WA
04.2: base: 50080000 size: 00040000 X P_RO U_RO Non-shareable outer: WB, no WA inner: WB, no WA
04.3: base: 500c0000 size: 00040000 X P_RO U_RO Non-shareable outer: WB, no WA inner: WB, no WA
04.4: base: 50100000 size: 00040000 X P_RO U_RO Non-shareable outer: WB, no WA inner: WB, no WA
05: base: 50140000 size: 00040000 XN P_RW U_RW Non-shareable outer: WB, no WA inner: WB, no WA
06: base: 50200000 size: 00080000 XN P_RO U_NA Non-shareable outer: WB, no WA inner: WB, no WA
07.3: base: 50600000 size: 00200000 XN P_RO U_RO Non-shareable outer: WB, no WA inner: WB, no WA
07.4: base: 50800000 size: 00200000 XN P_RO U_RO Non-shareable outer: WB, no WA inner: WB, no WA
09: base: 40000000 size: 00040000 XN P_RW U_RW Non-shareable outer: WB, no WA inner: WB, no WA
15: base: 50216000 size: 00002000 XN P_RW U_RW Non-shareable outer: WB, no WA inner: WB, no WA
mail worker 3 sp 0x50217e68 stack: 50216000 - 50217ffc
Per task regions for mail worker 3
04: base: 50216000 size: 00002000 XN P_RW U_RW Non-shareable outer: WB, no WA inner: WB, no WA
Enable MMIO: 1, Enable data R/W: 1
call stack:
sp 0x50217e68 pc 0x500221d4
sp 0x50217e68 pc 0x50023542
sp 0x50217e80 pc 0x500b067a
sp 0x50217eb8 pc 0x500ae6c8
sp 0x50217ef0 pc 0x500ae67a
sp 0x50217f00 pc 0x500c709a
sp 0x50217f48 pc 0x500ad2b2
sp 0x50217f58 pc 0x500c796a
sp 0x50217f78 pc 0x500a6314
sp 0x50217fa8 pc 0x5002cd92
sp 0x50217fd0 pc 0x50030130
sp 0x50217fd8 pc 0x500c8578
sp 0x50217ff8 pc 0xaaaaaaaa
eht_idx_find: 0xaaaaaaaa not a valid code address
no eidx for 0xaaaaaaaa

Sorry for the late response.
Is this still an issue to support? Any result can be shared?

yes, still be an issue and need support.
Not any new result can be shared.

There is no update from you for a period, assuming this is not an issue anymore.
Hence, we are closing this topic. If need further support, please open a new one.
Thanks
~0422

I saw an error on PCIe C1. What device is in use on this controller?