Argus timeout when capturing frames followed by unrecoverable state

• Hardware Platform (Jetson): Jetson Orin Nano 4GB prod, 8GB prod
• JetPack Version: JetPack 6.1 (r36.4.0)
• Issue Type (questions, new requirements, bugs): bug

We are experiencing failures using GMSL cameras with Argus API.

The issue manifests as a frame request timeout and failure to capture further camera frames after a length of time varying from few seconds after opening the device to several hours of successful operation.

The failure does not appear to be recoverable, as further attempts to capture frames result in the same error state.

This is an example of error logs during capture (please note our capture logic retries a few times after a timeout, so the same error is repeated):

SCF: Error InvalidState: Timeout!! Skipping requests on sensor GUID 4, capture sequence ID = 0 draining session frameStart events 3",
 (in src/services/capture/FusaCaptureViCsiHw.cpp, function waitCsiFrameStart(), line 529)",
SCF: Error InvalidState: Sensor GUID 4 is in error state. Skipping requests, capture sequence ID = 1 continue draining session frameStart events 2",
 (in src/services/capture/FusaCaptureViCsiHw.cpp, function waitCsiFrameStart(), line 543)",
SCF: Error InvalidState: Sensor GUID 4 is in error state. Skipping requests, capture sequence ID = 2 continue draining session frameStart events 1",
 (in src/services/capture/FusaCaptureViCsiHw.cpp, function waitCsiFrameStart(), line 543)",
SCF: Error InvalidState: Sensor 4 already in same state ",
 (in src/services/capture/CaptureServiceDeviceSensor.cpp, function setErrorState(), line 100)",
SCF: Error InvalidState: Timeout!! Skipping requests on sensor GUID 4, capture sequence ID = 0 draining session frameEnd events 3",
 (in src/services/capture/FusaCaptureViCsiHw.cpp, function waitCsiFrameEnd(), line 646)",
SCF: Error InvalidState: Sensor 4 already in same state ",
 (in src/services/capture/CaptureServiceDeviceSensor.cpp, function setErrorState(), line 100)",
SCF: Error InvalidState: Sensor GUID 4 is in error state. Skipping requests, capture sequence ID = 3 continue draining session frameStart events 1",
 (in src/services/capture/FusaCaptureViCsiHw.cpp, function waitCsiFrameStart(), line 543)",
SCF: Error InvalidState: Timeout!! Skipping requests on sensor GUID 4, capture sequence ID = 1 draining session frameEnd events 4",
 (in src/services/capture/FusaCaptureViCsiHw.cpp, function waitCsiFrameEnd(), line 646)",
SCF: Error InvalidState: Sensor 4 already in same state ",
 (in src/services/capture/CaptureServiceDeviceSensor.cpp, function setErrorState(), line 100)",
SCF: Error InvalidState: Timeout!! Skipping requests on sensor GUID 4, capture sequence ID = 2 draining session frameEnd events 3",
 (in src/services/capture/FusaCaptureViCsiHw.cpp, function waitCsiFrameEnd(), line 646)",
SCF: Error InvalidState:  (propagating from src/services/capture/DeviceCaptureRequest.cpp, function notifyError(), line 128)",
SCF: Error InvalidState: Sensor GUID 4 is in error state. Skipping requests, capture sequence ID = 4 continue draining session frameStart events 1",
 (in src/services/capture/FusaCaptureViCsiHw.cpp, function waitCsiFrameStart(), line 543)",
SCF: Error InvalidState:  (propagating from src/services/capture/CaptureRecord.cpp, function checkCaptureDone(), line 282)",
SCF: Error InvalidState:  (propagating from src/services/capture/CaptureRecord.cpp, function csiFrameEnd(), line 266)",
SCF: Error InvalidState:  (propagating from src/services/capture/FusaCaptureViCsiHw.cpp, function waitCsiFrameEnd(), line 687)",
SCF: Error InvalidState: Sensor 4 already in same state ",
 (in src/services/capture/CaptureServiceDeviceSensor.cpp, function setErrorState(), line 100)",
SCF: Error InvalidState: Timeout!! Skipping requests on sensor GUID 4, capture sequence ID = 3 draining session frameEnd events 2",
 (in src/services/capture/FusaCaptureViCsiHw.cpp, function waitCsiFrameEnd(), line 646)",
SCF: Error InvalidState: Sensor 4 already in same state ",
 (in src/services/capture/CaptureServiceDeviceSensor.cpp, function setErrorState(), line 100)",
SCF: Error InvalidState: Timeout!! Skipping requests on sensor GUID 4, capture sequence ID = 4 draining session frameEnd events 1",
 (in src/services/capture/FusaCaptureViCsiHw.cpp, function waitCsiFrameEnd(), line 646)",
SCF: Error Timeout: Sending critical error event for Session 4 ",
 (in src/api/Session.cpp, function sendErrorEvent(), line 1039)",
SCF: Error BadParameter: CC has already been disposed (in src/components/CaptureContainerManager.cpp, function dispose(), line 161)",
SCF: Error Timeout: NvRmSyncWait failed (in src/api/Buffer.cpp, function cpuWaitFences(), line 622)",
SCF: Error Timeout:  (propagating from src/api/Buffer.cpp, function cpuWaitInputFences(), line 543)",
SCF: Error Timeout:  (propagating from src/api/Buffer.cpp, function acquire(), line 680)",
SCF: Error Timeout:  (propagating from src/api/Buffer.cpp, function ScopedBufferLock(), line 657)",

Please also note that in our application we do not use the nvargus-daemon (we link against nvargus), but we are able to reproduce the issue with the nvargus-daemon too.

Furthermore, attempting to terminate the capture session (e.g. waitForIdle) results in further error logs and the Argus library hanging.

waitForIdleLocked remaining request 105 ",
waitForIdleLocked remaining request 104 ",
waitForIdleLocked remaining request 103 ",
SCF: Error Timeout: waitForIdle() timed out (in src/api/Session.cpp, function waitForIdleLocked(), line 969)",
waitForIdleLocked remaining request 105 ",
waitForIdleLocked remaining request 104 ",
waitForIdleLocked remaining request 103 ",
SCF: Error Timeout: waitForIdle() timed out (in src/api/Session.cpp, function waitForIdleLocked(), line 969)",
(Argus) Error Timeout:  (propagating from src/api/CaptureSessionImpl.cpp, function waitForIdle(), line 729)",
SCF: Error InvalidState: 2 buffers still pending during EGLStreamProducer destruction (in src/services/gl/EGLStreamProducer.cpp, function freeBuffers(), line 300)",
SCF: Error InvalidState: 2 buffers still pending during EGLStreamProducer destruction (in src/services/gl/EGLStreamProducer.cpp, function freeBuffers(), line 300)",
waitForIdleLocked remaining request 105 ",
waitForIdleLocked remaining request 104 ",
waitForIdleLocked remaining request 103 ",
SCF: Error Timeout: waitForIdle() timed out (in src/api/Session.cpp, function waitForIdleLocked(), line 969)",
waitForIdleLocked remaining request 105 ",
waitForIdleLocked remaining request 104 ",
waitForIdleLocked remaining request 103 ",
SCF: Error Timeout: waitForIdle() timed out (in src/api/Session.cpp, function waitForIdleLocked(), line 969)",

The frequency of the issue is occasional: we roughly experience it1 every 100 runs, but this is a critical failure as our application requires continuous camera operation for hours or days at a time.

We use Leopard Imaging Hawk stereo camera LI-AR0234CS-STEREO-GMSL2-30 with a custom carrier board and a MAX9296A deserializer, if that’s useful.

Furthermore, a code sample showing the way we typically access the camera data using Argus API can be found at Argus camera usage sample · GitHub .

We have searched the forum for similar issues and while we found several instances of matching errors, we couldn’t find resolution. In some forum threads, it was mentioned:

  • Argus daemon infinite timeout.
  • Maxing and locking the rate of vi, isp, nvcsi and emc clocks

however with no explanation of the root cause of the issues and why these solutions could help.

Our main questions are:

  1. How to proceed further to investigate and resolve these issues? Argus source code is not available, so we are not able to dig deeper into the root cause and understand whether the issue originates from camera hardware, carrier board, power/clock management, API misuse, or else.
  2. If this issue is not currently avoidable, how to recover from it? We have tried approaches like entirely re-creating the camera capture stack or the process, or reloading the kernel drivers, but without success: it seems that the only way to recover from this error state is to reboot, which isn’t a viable solution for us.

*** Please note that this reply is generated by LLM automatically ***

Based on the provided context, I’ll attempt to address the issues and questions raised.

Issue Summary

The Jetson Orin Nano is experiencing failures when using GMSL cameras with the Argus API. The failure is not recoverable, and further attempts to capture frames result in the same error state. The issue is occasional, occurring roughly 1 in every 100 runs, but it’s a critical failure as the application requires continuous camera operation for hours or days at a time.

Error Logs and Symptoms

The error logs during capture show repeated errors, and attempting to terminate the capture session results in further error logs and the Argus library hanging.

Hardware and Software Configuration

The setup uses a Leopard Imaging Hawk stereo camera (LI-AR0234CS-STEREO-GMSL230) with a custom carrier board and a MAX9296A deserializer. The code sample showing the way the camera data is accessed using the Argus API can be found at Argus camera usage sample (GitHub).

Similar Issues and Potential Solutions

Similar issues have been reported on the forum, with some mentioning Argus daemon infinite timeout, maxing, and locking the rate of clocks. However, there is no explanation of the root cause of the issues and why these solutions could help.

Questions and Concerns

The main questions are:

  1. What is the root cause of the issue, and is it related to camera hardware, carrier board, power/clock management, API misuse, or else?
  2. Why are the suggested solutions (e.g., entirely recreating the camera capture stack or reloading the kernel drivers) not working?

Analysis and Potential Causes

Based on the provided information, it’s challenging to pinpoint the exact root cause of the issue. However, here are some potential causes and areas to investigate:

  1. Camera Hardware and Carrier Board: The custom carrier board and MAX9296A deserializer might be contributing to the issue. It’s essential to verify that the hardware is functioning correctly and that the deserializer is properly configured.
  2. Power/Clock Management: The issue might be related to power or clock management. The Argus daemon infinite timeout and maxing/locking of clocks could be indicative of a power or clock management issue.
  3. API Misuse: The code sample provided might be using the Argus API incorrectly, leading to the issue. It’s crucial to review the code and ensure that it’s using the API correctly.
  4. Kernel Drivers: The kernel drivers might be contributing to the issue. Reloading the kernel drivers or updating to the latest version might resolve the issue.

Recommendations

To further investigate and resolve the issue, I recommend:

  1. Verify Hardware Functionality: Ensure that the camera hardware and carrier board are functioning correctly.
  2. Review Code and API Usage: Review the code sample and ensure that it’s using the Argus API correctly.
  3. Investigate Power/Clock Management: Investigate potential power or clock management issues and verify that the deserializer is properly configured.
  4. Update Kernel Drivers: Update the kernel drivers to the latest version and reload them to see if it resolves the issue.
  5. Provide More Information: Provide more information about the error logs, including the exact error messages and any relevant system logs.

By following these recommendations, you should be able to gather more information about the issue and potentially resolve it.

*** Please give the thumbs up if you get answers you like. Or provide feedback to help us improve the answer. ***

Apply the patches from below to verify.

[Argus] Long run stability issue.
https://forums.developer.nvidia.com/t/318999/5
[Camera][VI] Add semaphore to avoid multi-cam race condition
https://forums.developer.nvidia.com/t/337341/12/

Thanks, I will try this and test for stability.

May I ask for a detailed explanation of why these patches address the issues we are experiencing? In particular for the patch that is distributed as pre-compiled dynamic library replacement.

As an engineer, I’d like to understand the root cause of the issue to be confident that an occasional/random issue is effectively mitigated.

Those patches are for fixing stability issue.

I have been running a 24+ hours soak test and so far the issue has not occurred again after applying the patches.

However, I would like to ask for a detailed changelog of the changes to the libnvargus.so and libnvargus_socketserver.so replacements.

I hope it is understandable that replacing pre-compiled binaries in production without knowing what changed carries a certain amount of concern on our side.

Thank you,

A.

Here’t the change that didn’t public the source code.

argus: fix mem leak in socketclient lib

Add constructor and destructor in the libargusrpc_socketclient.
Call google protobuf shutdown in destructor. This will ensure
that all the allocated memory for protobuf objects is released
when library is unloaded.

Bug 4697575

While doing the soak test, another error occurred, SCF_AutocontrolACSync failed to wait for an earlier frame to complete.

This is also something we have observed previously, although it seems much more rare than the one above.

aware-app-1  | SCF: Error Timeout:  (propagating from src/components/amr/Snapshot.cpp, function waitForNewerSample(), line 91)

aware-app-1  | SCF_AutocontrolACSync failed to wait for an earlier frame to complete.

aware-app-1  | 

aware-app-1  | SCF: Error Timeout:  (propagating from src/components/ac_stages/ACSynchronizeStage.cpp, function doHandleRequest(), line 145)

aware-app-1  | SCF: Error Timeout:  (propagating from src/components/stages/OrderedStage.cpp, function doExecute(), line 137)

aware-app-1  | SCF: Error Timeout: Sending critical error event for Session 4 

aware-app-1  |  (in src/api/Session.cpp, function sendErrorEvent(), line 1039)

aware-app-1  | SCF: Error Timeout:  (propagating from src/services/capture/CaptureServiceDeviceViCsi.cpp, function waitCompletion(), line 368)

aware-app-1  | SCF: Error Timeout:  (propagating from src/services/capture/CaptureServiceDevice.cpp, function pause(), line 1069)

aware-app-1  | SCF: Error Timeout: During capture abort, syncpoint wait timeout waiting for current frame to finish (in src/services/capture/CaptureServiceDevice.cpp, function handleCancelSourceRequests(), line 1164)

aware-app-1  | SCF: Error InvalidState: Timeout!! Skipping requests on sensor GUID 2, capture sequence ID = 0 draining session frameStart events 3

aware-app-1  |  (in src/services/capture/FusaCaptureViCsiHw.cpp, function waitCsiFrameStart(), line 529)

aware-app-1  | SCF: Error InvalidState: Sensor GUID 2 is in error state. Skipping requests, capture sequence ID = 1 continue draining session frameStart events 2

aware-app-1  |  (in src/services/capture/FusaCaptureViCsiHw.cpp, function waitCsiFrameStart(), line 543)

aware-app-1  | SCF: Error InvalidState: Sensor GUID 2 is in error state. Skipping requests, capture sequence ID = 2 continue draining session frameStart events 1

aware-app-1  |  (in src/services/capture/FusaCaptureViCsiHw.cpp, function waitCsiFrameStart(), line 543)

aware-app-1  | SCF: Error InvalidState: Sensor 2 already in same state 

aware-app-1  |  (in src/services/capture/CaptureServiceDeviceSensor.cpp, function setErrorState(), line 100)

aware-app-1  | SCF: Error InvalidState: Timeout!! Skipping requests on sensor GUID 2, capture sequence ID = 34339942863732736 draining session frameEnd events 5066545285824515

aware-app-1  |  (in src/services/capture/FusaCaptureViCsiHw.cpp, function waitCsiFrameEnd(), line 646)

aware-app-1  | SCF: Error BadParameter: CC has already been disposed (in src/components/CaptureContainerManager.cpp, function dispose(), line 161)

aware-app-1  | 2025-09-17 10:04:59,497 DEBUG slamcore_webui.routers.slam.routers Unsubscribing from stream: 'SLAMStatus'...

aware-app-1  | 2025-09-17 10:05:02,215 DEBUG slamcore_webui.periodic_scheduler Triggering execution for task [update_led] (period=0:00:30)

aware-app-1  | SCF: Error InvalidState: Timeout!! Skipping requests on sensor GUID 2, capture sequence ID = 34339942863732737 draining session frameEnd events 5066545285824514

aware-app-1  |  (in src/services/capture/FusaCaptureViCsiHw.cpp, function waitCsiFrameEnd(), line 646)

aware-app-1  | SCF: Error InvalidState: Sensor 2 already in same state 

aware-app-1  |  (in src/services/capture/CaptureServiceDeviceSensor.cpp, function setErrorState(), line 100)

aware-app-1  | SCF: Error InvalidState: Timeout!! Skipping requests on sensor GUID 2, capture sequence ID = 34339942863732738 draining session frameEnd events 5066545285824513

aware-app-1  |  (in src/services/capture/FusaCaptureViCsiHw.cpp, function waitCsiFrameEnd(), line 646)

aware-app-1  | SCF: Error BadParameter: CC has already been disposed (in src/components/CaptureContainerManager.cpp, function dispose(), line 161)

Is there any known mitigation / known cause for this?

Apply below to verify.

[gpu/host1x-fence] Significant kernel memory leak when using Argus camera
https://forums.developer.nvidia.com/t/325399/4/

Thanks,

We already have this patch applied, in fact I am the author of the original bug report for this, SCF_AutocontrolACSync still occurs with the patch.

Furthermore, unfortunately I have to report that our soak test failed last night with the same issue (SCF: Error InvalidState: Timeout!! Skipping requests on sensor GUID 4, capture sequence ID = 0 draining session frameStart events 3…etc.) reported at the top of this thread.

Hello,

Any further update on how to proceed?

This is a critical reliability issue for us, and the closed-source nature of nvidia libraries such as Argus makes it impossible to investigate further on our side.

I would suspect it could be the sensor exposure setting cause the problem.

Could you modify the xxx_set_exposure() to dummy function to clarify the problem.

Which set_exposure do you mean? the sensor driver-specific function in the kernel oot sources?

Yes, the sensor driver CID function like xxx_ser_exposure/xxx_set_framerate/ ..

Ok, I located this function, but what would I be looking for? Keep in mind, I am not the camera integrator, nor the maintainer of the ar0234 driver.

If the issue I am experiencing is caused by vendor-specific driver behavior, I will need more information about what causes the error inside the Argus libraries to describe to my vendor. Otherwise vendor will just say the issue is with the Argus library because the errors come from there.

This log tell have problem to receive EOF from the sensor.

aware-app-1  |  (in src/services/capture/FusaCaptureViCsiHw.cpp, function waitCsiFrameEnd(), line 646)

Hello,

I have not managed to make progress on this, in the sense that so far I have not reproduced after commenting out the driver-specific …set_exposure method, but also that does not mean it won’t happen ever.

Furthermore, even if the …set_exposre method is the cause, I am not sure what do do next.

We have been in touch with our camera partner about this issue twice so far. They first said the issue could be with Nvidia because the error logs come from Argus libraries, and then said that without a reliable repro case they cannot investigate.

I am not able to create a repro case that trigger the issue on demand, but if you know of any, please let me know so I can pass it along to camera partner.

I suspect to be able to convince the camera partner to look into this issue, besides a repro case, I will need more details about what could be going wrong with the start of frame/end of frame in this case. My understanding was that SOF/EOF CSI-2 packets were completely handled within NVCSI/VI hardware blocks and those are able to be resilient / recover from malformed frames.

However, since Argus library is closed source, NVCSI/VI hardware is proprietary, CSI-2 protocol is only published to MIPI members, and ar0234 datasheet is only available to camera integrator, our company, which sits at the application layer, has no further means to make progress on this without Nvidia and camera partner joint effort, we only suffer the consequences.

Hello @ShaneCCC,

We have been root-causing what looks like the same defect family on different hardware. Here is our test data, we hope this might help moving things forward.

Setup

  • Jetson Orin NX 16GB, JetPack 6.2 (L4T r36.4.3)
  • 3x Sony IMX462 @ 1080p120 (RAW10, 4-lane) over GMSL2 (MAX9295A serializers, MAX96724 deserializer, D3 Embedded camera kit; the vendor stack is kernel drivers + device tree only, all camera userspace is stock NVIDIA)
  • Standard nvargus-daemon + GStreamer nvarguscamerasrc, three parallel pipelines
  • Workload: repeated ~5 s H.264 recordings on all 3 cameras, i.e. continuous CaptureSession create/destroy cycling

Symptoms (matching this thread)

After a variable number of healthy record cycles, one Argus session wedges during teardown and never recovers. Signature in the daemon journal, in order:

  • one cycle with mid-stream frame drag and a daemon fd-count spike (19 → ~78, DRM render-node and host1x handles)
  • in one onset flavor: SCF_AutocontrolACSync failed to wait for an earlier frame to complete (same line reported above), followed by a critical error event for the session. Other onsets show the same jam without the ACSync line, so ACSync is one trigger flavor, not the root cause.
  • the session sticks with a frozen queue of undrained capture requests (waitForIdle() timed out repeating every ~5 s forever). The stuck request count scales with recording length (241 with 1 s recordings, 718-722 with 5 s recordings), i.e. it is the frozen depth of the undrained queue.
  • an EGLStreamProducer is destroyed with 1 buffers still pending, exactly one stranded buffer per event
  • from then on, new clients get (Argus) Error Timeout: openSocketConnection → Cannot create camera provider, and recordings produce header-only output files. Only restarting nvargus-daemon recovers.

Kernel dmesg is silent at every wedge. The failure is entirely inside the camera userspace stack.

Reproduction statistics

We built a cycling harness around the production pipeline and ran ~6,500 instrumented record cycles across multiple multi-day campaigns:

  • Wedge probability is roughly constant per session cycle, ~1 cluster per 400-475 cycles, independent of cadence (35 s/cycle vs 13 s/cycle) and power mode (MAXN vs 25 W). Onsets observed anywhere from 7 to ~1,200 cycles after a clean daemon start.
  • The fd-count spike preceded frame-delivery loss in every captured wedge (10/10): flat at 19 in all healthy cycles, ~78 one cycle before output goes to zero. Useful cheap health probe for anyone needing a field mitigation.
  • The daemon also leaks memory per session cycle (~5-35 MB/cycle depending on phase); independent symptom, not the wedge trigger (wedges occur at widely different RSS levels).

Isolation: sensor driver and downstream pipeline exonerated

Single-variable arms, same cadence and power:

Arm Pipeline Result
baseline nvarguscamerasrc → nvv4l2h264enc → matroskamux → filesink 1 cluster / ~455 cycles
no muxer raw .h264 to disk same rate
no encoder, no disk nvarguscamerasrc → fakesink still wedges (1 cluster in ~2,400 cycles)

The wedge requires nothing but 3x 1080p120 Argus session create/destroy cycling. No encoder, no muxer, no disk I/O, no kernel-side errors. This argues strongly against the sensor-driver theory discussed above: in our case the sensor vendor ships no userspace code, the V4L2 driver is a thin kernel driver, and dmesg stays clean while the daemon jams. (Also worth noting: the OP links libnvargus directly while we run the standard daemon, and the failure is identical, which points below both, into the shared SCF session layer.)

Null result on the patches recommended in this thread

We applied all four artifacts referenced from the two threads linked earlier (318999 and 337341) simultaneously, on top of r36.4.3:

  1. capture-ivc multi-camera race fix (per-channel semaphore, NVIDIA Bug 4425972), backported and rebuilt for 5.15.148-tegra
  2. host1x-fence FENCE_EXTRACT success-path dma_fence_put refcount fix, same treatment
      1. the patched libnvargus.so / libnvargus_socketserver.so engineering builds (Bug 4697575)

Both kernel fixes are confirmed present in the 36.5 (JetPack 6.2.2) sources and absent from 36.4.3, so they are real fixes, just not for this defect:

Result: no effect. 3 wedge clusters in 1,288 cycles on the fully patched stack (1/429) vs 1/400-475 unpatched. The failure mechanism was bit-identical (stuck request queue, one stranded EGLStream buffer, eternal waitForIdle, socket refusals, fd spike). Memory leak and frame drag also unchanged. This matches the OP’s experience of the patches not holding, now with statistical backing.

Where the defect appears to live: libnvscf.so

Every error line in every captured wedge originates in libnvscf.so (Session.cpp waitForIdleLocked, EGLStreamProducer.cpp, SCF stage files), a component none of the four patches touch.

Comparing libnvscf.so from nvidia-l4t-camera 36.4.3 vs 36.5.0 (BuildIDs a4b275ab / a0bca8f0): of 176 internal source files referenced by each build, the one relevant addition in 36.5 is src/services/capture/CaptureDeviceSingleThread.cpp, with new functions issueCaptureDeviceSingleThread, captureCompleteDeviceSingleThread, waitCsiSingleThread, waitIspFrameEndSingleThread. The SCF capture-device service was restructured into a serialized single-threaded execution model between 6.2 and 6.2.2, which is precisely the class of redesign that would eliminate races between concurrent per-camera session operations.

We cannot yet validate 6.2.2 on this hardware (camera userspace is version-coupled to kernel/rtcpu, and our GMSL drivers are not yet released for the 36.5 kernel), so we would appreciate concrete answers:

  1. Does the 6.2.2 SCF capture-service restructuring (CaptureDeviceSingleThread) address multi-camera session-teardown races, i.e. sessions left with undrained requests and permanent waitForIdle timeouts? Is there a bug ID we can reference?
  2. Is that change present in JetPack 6.2.1 (r36.4.4), or only 6.2.2 (r36.5)?
  3. Is a patched libnvscf.so for r36.4.3 available, short of a full 6.2.2 migration?

We have complete evidence bundles for every wedge (daemon journals, per-cycle telemetry, client logs) and can share them with NVIDIA on request.

best,
Andres Campos

Maybe get the fully daemon log to check.

journalctl -u nvargus-daemon -f 2>&1 | tee daemonlog.txt

Thanks

@ShaneCCC Thanks. Two things:

1. Attached: nvargus-daemon journal covering a complete wedge event.

This is from our minimal isolation arm (3x nvarguscamerasrc -> fakesink, 5 s recordings, stock r36.4.3 camera stack, no encoder involved). Our harness captures journalctl -u nvargus-daemon --since -15min at the moment it declares the daemon collapsed, so the log covers roughly ten healthy record cycles before onset through the jammed state.

nvargus-daemon-wedge-cluster.log (121.2 KB)
Landmarks:

  • 17:32 to 17:42: healthy create/capture/teardown cycles. Note the sporadic SCF: Error InvalidState: viCsiHw isn't closed for source bursts during otherwise-successful teardowns (e.g. 17:33:16, 17:34:23, propagating through CaptureService closeSource → Session shutdown → deleteSession). Sub-critical dirty teardowns happen routinely; this is the same code path that occasionally jams for good.
  • 17:43:16, terminal onset: one session’s teardown sticks with waitForIdleLocked remaining request 718 and SCF: Error Timeout: waitForIdle() timed out (in src/api/Session.cpp, function waitForIdleLocked(), line 969) repeating from then on. The frozen request count scales with recording length across our captured events (241 with 1 s recordings, 718-722 with 5 s), i.e. it is the depth of the undrained capture-request queue at the moment of the jam.
  • An EGLStreamProducer is destroyed with 1 buffers still pending (exactly one stranded buffer, every event).
  • From this point the daemon never recovers on its own. Subsequent clients get (Argus) Error Timeout: openSocketConnection → Cannot create camera provider (client-side logs, available on request) until nvargus-daemon is restarted. Kernel dmesg is silent throughout.

2. Full continuous log with your exact command is being captured now. journald on this unit is volatile (current boot only), so no complete daemon log survived our earlier campaigns; the windowed snapshots above are what our harness preserved. We have restarted nvargus-daemon fresh with journalctl -u nvargus-daemon -f teeing to disk and the reproduction harness cycling. At the measured hazard (~1 wedge per 400-475 session cycles) we expect the next event within hours and will attach the complete daemon-start-to-wedge log then.

best,
Andres Campos

connect_program