Kernel panic when sharing NvSciBuf between processes

Please provide the following info (tick the boxes after creating this topic):
Software Version
DRIVE OS 6.0.8.1
DRIVE OS 6.0.6
DRIVE OS 6.0.5
DRIVE OS 6.0.4 (rev. 1)
DRIVE OS 6.0.4 SDK
other

Target Operating System
Linux
QNX
other

Hardware Platform
DRIVE AGX Orin Developer Kit (940-63710-0010-300)
DRIVE AGX Orin Developer Kit (940-63710-0010-200)
DRIVE AGX Orin Developer Kit (940-63710-0010-100)
DRIVE AGX Orin Developer Kit (940-63710-0010-D00)
DRIVE AGX Orin Developer Kit (940-63710-0010-C00)
DRIVE AGX Orin Developer Kit (not sure its number)
other

SDK Manager Version
1.9.3.10904
other

Host Machine Version
native Ubuntu Linux 20.04 Host installed with SDK Manager
native Ubuntu Linux 20.04 Host installed with DRIVE OS Docker Containers
native Ubuntu Linux 18.04 Host installed with DRIVE OS Docker Containers
other

It seems the APIs to allow sharing NvSciBuf handles via NvSciIpc can cause a kernel panic. Say process A has an NvSciBuf that it needs to send to process B. The last stage in that sequence of sending the buffer is A reads the NvSciBufObjIpcExportDescriptor from the IPC handle and imports the buffer. If process B closes its handles to the NvSciBuf after A has imported the buffer, then everything operates as expected and A can still use the buffer. However, if process B closes its handle either before or while A is importing the NvSciBuf, then this causes a kernel panic.

Checking the journal from the failed boot shows this:

Feb 01 13:52:05 tegra-ubuntu kernel: Unable to handle kernel NULL pointer dereference at virtual address 0000000000000008
Feb 01 13:52:05 tegra-ubuntu kernel: Mem abort info:
Feb 01 13:52:05 tegra-ubuntu kernel: ESR = 0x0000000096000006
Feb 01 13:52:05 tegra-ubuntu kernel: EC = 0x25: DABT (current EL), IL = 32 bits
Feb 01 13:52:05 tegra-ubuntu kernel: SET = 0, FnV = 0
Feb 01 13:52:05 tegra-ubuntu kernel: printk: console [ttyS2]: printing thread stopped
Feb 01 13:52:05 tegra-ubuntu kernel: EA = 0, S1PTW = 0
Feb 01 13:52:05 tegra-ubuntu kernel: FSC = 0x06: level 2 translation fault
Feb 01 13:52:05 tegra-ubuntu kernel: Data abort info:
Feb 01 13:52:05 tegra-ubuntu kernel: ISV = 0, ISS = 0x00000006
Feb 01 13:52:05 tegra-ubuntu kernel: CM = 0, WnR = 0
Feb 01 13:52:05 tegra-ubuntu kernel: user pgtable: 4k pages, 48-bit VAs, pgdp=0000000212091000
Feb 01 13:52:05 tegra-ubuntu kernel: [0000000000000008] pgd=0800000210fd8003, p4d=0800000210fd8003, pud=08000001040ea003, pmd=0000000000000000
Feb 01 13:52:05 tegra-ubuntu kernel: Internal error: Oops: 96000006 [#1] PREEMPT_RT SMP

I would expect the import in process A to simply fail at the API level if process B closes its handle too early and just return an NvSciError. Is this a known issue? Is there is a patch to address this?

In the described scenario, did the NvSciBufObjIpcImport() function still return NvSciError_Success even though the kernel panic occured? Do you expect the function to return an error?

The function NvSciBufObjIpcImport is what is causing the panic. So no it did not return at all. This is reproducible if process B closes its handle to the buffer before process A calls NvSciBufObjIpcImport. If that happens, then the panic occurs during the call to NvSciBufObjIpcImport which means it does not return any error code at all. This is unexpected. And yes for a failure case like this, NvScibufObjIpcImport should return an error code and not simply cause a kernel panic.

The journal didn’t have the trace, but the pstore did. This should help.

<1>[ 182.244107] Unable to handle kernel NULL pointer dereference at virtual address 0000000000000008
<1>[ 182.245767] Mem abort info:
<1>[ 182.246271] ESR = 0x0000000096000006
<1>[ 182.246958] EC = 0x25: DABT (current EL), IL = 32 bits
<1>[ 182.247975] SET = 0, FnV = 0
<6>[ 182.247976] printk: console [ttyS2]: printing thread stopped
<1>[ 182.248529] EA = 0, S1PTW = 0
<1>[ 182.250151] FSC = 0x06: level 2 translation fault
<1>[ 182.251025] Data abort info:
<1>[ 182.251549] ISV = 0, ISS = 0x00000006
<1>[ 182.252256] CM = 0, WnR = 0
<1>[ 182.252807] user pgtable: 4k pages, 48-bit VAs, pgdp=0000000226722000
<1>[ 182.253965] [0000000000000008] pgd=080000014b7f9003, p4d=080000014b7f9003, pud=080000022616b003, pmd=0000000000000000
<0>[ 182.255897] Internal error: Oops: 96000006 [#1] PREEMPT_RT SMP
<4>[ 182.256938] Modules linked in: xt_conntrack xt_MASQUERADE nf_conntrack_netlink nfnetlink xt_addrtype iptable_filter iptable_nat nf_nat nf_conntrack libcrc32c nf_defrag_ipv6 nf_defrag_ipv4 br_netfilter fuse 8021q garp mrp nvidia_modeset(O) lan743x(O) tegra_pcie_dma_test(O) tegra_pcie_edma(O) tegra210_adma spidev cdi_mgr(O) snd_soc_tegra_virt_t210ref_pcm(O) isc_mgr(O) cdi_pwm(O) isc_pwm(O) snd_soc_tegra210_virt_alt_admaif(O) cdi_dev(O) isc_dev(O) tegra_hv_pm_ctl(O) tegra_hv_vcpu_yield(O) cdi_gpio(O) nvidia(O) isc_gpio(O) tegra_fsicom(O) mttcan(O) phy_tegra194_p2u lm90 can_dev cam_fsync(O) tegra_aconnect tegra_uss_io_proxy(O) tegra_bpmp_thermal tegra_xudc spi_tegra114 tegra_dce(O) tsecriscv(O) watchdog_tegra_t18x(O) pcie_tegra194 safety_i2s(O) nvhost_isp5(O) nvhost_vi5(O) nvhost_nvcsi_t194(O) tegra_camera(O) v4l2_dv_timings v4l2_fwnode v4l2_async videobuf2_dma_contig nvhost_nvcsi(O) tegra_camera_platform(O) mc_utils(O) capture_ivc(O) tegra_drm_next(O) videobuf2_v4l2 videobuf2_memops
<4>[ 182.256995] videobuf2_common cec videodev mc nvhost_pva(O) drm_kms_helper nvhost_capture(O) drm bridge stp llc cpuidle_tegra_auto(O) nvhost_nvdla(O) nvhwpm(O) host1x_nvhost(O) camchar(O) camera_diagnostics(O) debug(O) tegra_camera_rtcpu(O) device_group(O) firmwares_class(O) reset_group(O) ivc_bus(O) hsp_mailbox_client(O) clk_group(O) nvgpu(O) tegra_gr_comm(O) nvmap(O) hvc_sysfs(O) tegra_nvvse_cryptodev(O) tegra_hv_vse_safety(O) host1x_fence(O) host1x_next(O) nvsciipc(O) userspace_ivc_mempool(O) ivc_cdev(O) ip_tables x_tables ipv6 nvme(E) nvme_core(E) oak_pci(OE) nvethernet(OE) pinctrl_tegra234(OE) nvpps(OE) tegra194_gte(OE) tegra_bpmp(OE) tegra_vblk(OE) tegra_hv_vblk_oops(OE) tegra_hv(OE) ivc_ext(OE)
<4>[ 182.283247] CPU: 6 PID: 2617 Comm: buffer_sharing_ Tainted: G OE 5.15.98-rt-tegra #1
<4>[ 182.284832] Hardware name: p3710-0010 (DT)
<4>[ 182.285553] pstate: 60400005 (nZCv daif +PAN -UAO -TCO -DIT -SSBS BTYPE=–)
<4>[ 182.286811] pc : nvmap_duplicate_handle+0x2f8/0x380 [nvmap]
<4>[ 182.287814] lr : nvmap_duplicate_handle+0x218/0x380 [nvmap]
<4>[ 182.288766] sp : ffff80001754bbf0
<4>[ 182.289354] x29: ffff80001754bc10 x28: ffff00009523a0b8 x27: 0000000000000000
<4>[ 182.290633] x26: ffff80001754bd80 x25: 0000000000010005 x24: ffff8000015a5000
<4>[ 182.291909] x23: 0000000000000001 x22: 0000000000000000 x21: ffff00010720b000
<4>[ 182.293171] x20: ffff000192917d00 x19: ffff00009523a000 x18: 0000000000000001
<4>[ 182.294437] x17: 0000000000000000 x16: ffff8000015815c0 x15: ffffffffffffffff
<4>[ 182.295701] x14: ffffff0000000000 x13: ffffffffffffffff x12: 0000000000000028
<4>[ 182.296993] x11: 0101010101010101 x10: 0000000000000000 x9 : 0000000000000000
<4>[ 182.298258] x8 : ffff000192917d80 x7 : 0000000000000000 x6 : 000000000000003f
<4>[ 182.299532] x5 : 0000000000000040 x4 : ffff000196798f00 x3 : 0000000000000004
<4>[ 182.300808] x2 : 0000000000000000 x1 : ffff000196798f00 x0 : 0000000000000000
<4>[ 182.302084] Call trace:
<4>[ 182.302521] nvmap_duplicate_handle+0x2f8/0x380 [nvmap]
<4>[ 182.303459] nvmap_get_handle_from_sci_ipc_id+0x268/0x5c0 [nvmap]
<4>[ 182.304554] nvmap_ioctl_handle_from_sci_ipc_id+0xb4/0x100 [nvmap]
<4>[ 182.305663] __traceiter_refcount_free_handle+0x2e24/0x6ad0 [nvmap]
<4>[ 182.306803] __arm64_sys_ioctl+0xbc/0x100
<4>[ 182.307547] invoke_syscall+0x5c/0x150
<4>[ 182.308223] el0_svc_common.constprop.0+0x64/0x120
<4>[ 182.309089] do_el0_svc+0x3c/0xb0
<4>[ 182.309694] el0_svc+0x20/0x70
<4>[ 182.310249] el0t_64_sync_handler+0xc0/0xd0
<4>[ 182.311005] el0t_64_sync+0x1a4/0x1a8
<0>[ 182.311674] Code: 95f454e2 17ffffaf 3900929f f9402660 (f9400400)
<4>[ 182.312782] —[ end trace 0000000000000002 ]—

Thank you for providing more details.

Could you please try reproducing the issue with the “rawstream” application? If the issue persists, please share any modifications you made and the commands you used. This will help us understand the problem better.

First patch the rawstream_consumer.c file with the provided patch file.
rawstream_consumer.c.patch.txt (186 Bytes)

patch rawstream_consumer.c rawstream_consumer.c.patch.txt

Run the producer

./rawstream -p

Run the Consumer

./rawstream -c

The consumer will stop and output:

Kill the rawstream producer task and then press enter.

Kill the producer task and then press enter in the consumer terminal to cause the panic.

I have informed our team about your suggestion and will update you if there are any changes.

The resolution for this issue will be incorporated in the upcoming release. Just wanted to keep you informed.

When is the next expected release? Will patches be made available prior to the release?

Please note that patches for issues are not provided in forum support. For information regarding the schedule of the next release, kindly contact your Nvidia representative.