Mmc0: Timeout waiting for hardware interrupt

Hi Nvidia,

We need to address the issue where interrupts are only aggregated and processed by CPU0. Therefore, we have incorporated the following patch: “re-enable GICv2m for PCIe MSI interrupts”.

R36.3 Patch to re-enable GICv2m for PCIe MSI interrupts and restore I/O performance - Jetson Systems / Jetson AGX Orin - NVIDIA Developer Forums

Currently, there is a problem with the test. Occasionally, there are situations where the MMC disk writes fail and the system freezes. From the log, it seems that there is an abnormality in the MMC interrupt response. Could you please advise on how to solve this?

Basic info: kernel is 5.15.148-rt-tegra, jetpack 6.2, Jetson Agx orin 64G

# cat /etc/nv_tegra_release 
# R36 (release), REVISION: 4.3, GCID: 38968081, BOARD: generic, EABI: aarch64, DATE: Wed Jan  8 01:49:37 UTC 2025
# KERNEL_VARIANT: oot
TARGET_USERSPACE_LIB_DIR=nvidia
TARGET_USERSPACE_LIB_DIR_PATH=usr/lib/aarch64-linux-gnu/nvidia

20260525-172127.log (1.4 MB)

mmc0: Timeout waiting for hardware interrupt.
[   28.872612] mmc0: running CQE recovery
[   28.874490] mmc0: cache flush error -110
[   40.159220] mmc0: running CQE recovery
[   40.161197] mmc0: cache flush error -110
[   18.763753] mmc0: running CQE recovery
[   19.265601] mmc0: cqhci: Failed to halt
[   30.177597] mmc0: Timeout waiting for hardware interrupt.
[   30.177603] mmc0: sdhci: ============ SDHCI REGISTER DUMP ===========
[   30.177607] mmc0: sdhci: Sys addr:  0x00000000 | Version:  0x00000505
[   30.177610] mmc0: sdhci: Blk size:  0x00007200 | Blk cnt:  0x00000080
[   30.177613] mmc0: sdhci: Argument:  0x401d0008 | Trn mode: 0x00000033
[   30.177615] mmc0: sdhci: Present:   0x11fb00f0 | Host ctl: 0x00000039
[   30.177618] mmc0: sdhci: Power:     0x0000000f | Blk gap:  0x00000000
[   30.177620] mmc0: sdhci: Wake-up:   0x00000000 | Clock:    0x0000000f
[   30.177623] mmc0: sdhci: Timeout:   0x0000000e | Int stat: 0x00000000
[   30.177625] mmc0: sdhci: Int enab:  0x00ff0003 | Sig enab: 0x00fc0003
[   30.177628] mmc0: sdhci: ACmd stat: 0x00000000 | Slot int: 0x00000000
[   30.177630] mmc0: sdhci: Caps:      0x3f6cd08c | Caps_1:   0x18002f73
[   30.177632] mmc0: sdhci: Cmd:       0x00002c1e | Max curr: 0x00000000
[   30.177635] mmc0: sdhci: Resp[0]:   0x00000900 | Resp[1]:  0x200021ae
[   30.177637] mmc0: sdhci: Resp[2]:   0x26468000 | Resp[3]:  0x00000000
[   30.177639] mmc0: sdhci: Host ctl2: 0x0000300d
[   30.177642] mmc0: sdhci: ADMA Err:  0x00000000 | ADMA Ptr: 0x0000007ffffe7090

Hi,

Could you help dump this information?

cat /sys/class/mmc_host/mmc0//mmc0:0001/manfid

# cat /sys/class/mmc_host/mmc0//mmc0:0001/manfid
0x000045

let me check internally for this.

Has there been any progress on the issue?

we will update you the result soon. Thanks for patience.

Hi @liteblue

We notice a known issue that Orin AGX with Sandisk eMMC needs to have a firmware update.

Please use this package to update the firmware.
nv_sandisk_emmc_fw.tar.gz (1.0 MB)

The MMC timeout issue has reoccurred. The MMC firmware has been upgraded. Could you please take a look at the problem and tell me how to solve it?

$ sudo cat /sys/block/mmcblk0/device/fwrev
0x3832313037393335

image

mmc-timeout.txt (28.2 KB)

[ 598.500128] mmc0: Timeout waiting for hardware interrupt.

Please put this SOM to NV devkit with sdkm image and see if you could still reproduce issue.

Also, upgrade your BSP to rel-36.5. JP6.2 might be too old with missing upstream kernel fix.

Hi WayneWWW,
Please look at the first comment. We encountered the MMC interrupt timeout issue after applying the ‘re-enable GICv2m for PCIe MSI…’ patch.
Testing revealed that only enabling this patch would cause the MMC interrupt timeout problem.
The official development board is not affected, which proves nothing.

Sorry that I didn’t remember the detail with a month ago.

Is it possible to do the test without RT kernel and only with the GICv2 patch?

RT kernel is actually a very large set. Need to clarify if that also leads to the reproducible.

Non-real-time kernel testing will be conducted in the near future.

Thanks for the detailed logs and system info. To help isolate the root cause, could you test whether the MMC timeout issue reproduces when you apply the GICv2m patch but use the non-RT kernel (standard 5.15.148-tegra without RT patches)?

This will help us determine whether the issue is specific to the RT kernel or caused by the GICv2m patch itself. Once you have results from that test, please share:

  1. Whether the timeout occurs with GICv2m + non-RT kernel
  2. If it does NOT occur, which non-RT kernel version you tested with
  3. Any relevant dmesg or SDHCI register dumps from the non-RT kernel test

This diagnostic will narrow down the scope significantly.

We’re hitting what looks like the same issue, and it might help with the RT-vs-GICv2m question. We get this on a completely stock kernel. No RT, no GICv2m patch, nothing custom.

Setup is an AGX Orin 64GB on a Syslogic carrier, L4T 36.4.7 via apt (nvidia-l4t-kernel 5.15.148-tegra-36.4.7-20250918154033, uname shows 5.15.148-tegra #1 SMP PREEMPT). eMMC is a SanDisk DG4064, manfid 0x000045, date 07/2025, fwrev 0x3733313033353137 (“73103517”). Rootfs on the eMMC, HS400-ES, CQE enabled.

We’ve had two incidents where the rootfs just dies underneath a running system, everything in RAM keeps going, network included, but nothing persists to disk anymore, and only a power cycle recovers it. We caught one in ramoops on July 19. The buildup looks like this:

mmc0: running CQE recovery
mmc0: cache flush error -110

repeating with shrinking intervals (13.5h apart, then 2.4h, then 1.7h, then seconds), then

mmc0: cqhci: Failed to halt
mmc0: cqhci: Failed to clear tasks

then a cbb-fabric SLAVE_ERR storm (MASTER_ID: CCPLEX, address 0x3460008, slave AXI2APB_8) and 21x “mmc0: Timeout waiting for hardware interrupt” with Int stat: 0x00000000 in the SDHCI dumps. We still see the occasional cache flush error -110 + CQE recovery on quiet days too (one on 7/21, five on 7/24, four of those within 4 minutes). eMMC health registers look fine, life_time 0x01/0x01, pre_eol 0x01, no filesystem errors ever.

So at least on our unit, neither the RT kernel nor the GICv2m patch is needed to trigger this.

Happy to attach the full ramoops log if useful.

May I know if your SOM is also with this one?

# cat /sys/class/mmc_host/mmc0//mmc0:0001/manfid
0x000045

And what is the method to reproduce this error?