Deepstream 9.1 filtering corrupted frames ìn the nvmultiurisrcbin

  1. Hardware Platform (Jetson / GPU): GPU — NVIDIA A10 (HPE DL380 Gen10, Ubuntu, running inside Docker)
  2. DeepStream Version: 9.1.0
  3. JetPack Version (valid for Jetson only): n/a
  4. TensorRT Version: 10.14.1.48 (cuda13.0)
  5. NVIDIA GPU Driver Version (valid for GPU only): 595.71.05
  6. Issue Type (questions, new requirements, bugs): question (possibly bug)
  7. How to reproduce the issue?: Custom Python app built on DeepStream Service Maker. Pipeline: nvmultiurisrcbin (RTSP over TCP, ~33 thermal Hikvision DS-2TD H.264/H.265 streams, 1280x720 @ ~25 fps) → nvinfer (custom RF-DETR parser, batch-size 35, 576×576 TensorRT engine). Some streams intermittently deliver heavily smeared/corrupted decoded frames (macroblock artifacts across most of the frame, see attached screenshots) that persist until the next clean I-frame. The corruption is already present in the decoded NV12 surface; nvinfer runs on these frames and produces false detections. Config files and GST_DEBUG log can be attached on request.
  8. Requirement details (new requirement): n/a

Issue: Some streams intermittently produce heavily smeared/corrupted decoded frames macroblock artifacts across most of the frame (screenshot attached), consistent with reference-frame corruption after packet loss or a bad keyframe. The corruption persists over multiple frames until the next clean I-frame. nvinfer still runs on these frames and produces (false) detections.

I normally only see this kind of corruption when a camera is starting up (a “warm-up” phase after the RTSP connection is established). In my pipeline, however, the batch keeps running continuously and the connection should not be dropped: as far as I understand, nvmultiurisrcbin does not repeatedly re-request the RTSP stream the session stays open the whole time. So I’m unsure whether the corruption is already present in the data going into the pipeline (camera/network side), or whether the pipeline itself (decoder / batching) is causing it.

Questions:

  1. Are there configurations I can apply within my pipeline (nvmultiurisrcbin, nvv4l2decoder, nvstreammux) to prevent this?
  2. Is it possible to filter these frames out before they reach nvinfer? For example, can I read the decoded frame data (or decoder/RTP metadata) and determine that a frame is already corrupt before it enters the batch?

  1. How are nvmultiurisrcbin and nvinfer’s properties set? what do you mean by "The corruption persists over multiple frames until the next clean I-frame. "? after I-frame, will corruption issue still persist?
  2. when the app is running, could you use “nvidia-smi dmon >1.log” to get three minutes of resource monitoring log? then compress and share the file.
  3. if using minimal “nvmultiurisrcbin-> tiler->render sink” pipeline, will corruption issue still persist?

Thanks for looking into this. Answers to your three points below;

1. Properties, and what “persists until the next I-frame” means

nvmultiurisrcbin, set with g_object_set() at build time, with USE_NEW_NVSTREAMMUX=yes in the environment:

max-batch-size=35  width=1280  height=720  batched-push-timeout=40000
config-file-path=nvstreammux_config.txt
    [property] algorithm-type=1  adaptive-batching=1  max-same-source-frames=1
               max-fps-control=0  overall-max-fps=25/1  overall-min-fps=10/1
live-source=1  cudadec-memtype=0  drop-frame-interval=0  dec-skip-frames=0
leaky=2  max-size-buffers=3  num-extra-surfaces=1  low-latency-mode=1
drop-on-latency=1  latency=100  select-rtp-protocol=4 (TCP)  udp-buffer-size=2097152
disable-audio=1  rtsp-reconnect-interval=10  rtsp-reconnect-attempts=-1
drop-pipeline-eos=1

So how i observe the damage frame is trough a detection/trigger. it waits a certain amount of frames before triggering the decoder to render a frame with the metadata from the batch frame.

nvinfer:
network-mode=2 (FP16) batch-size=35 process-mode=1 interval=0 gie-unique-id=1
model-color-format=0 net-scale-factor=0.0173520735727919 offsets=123.675;116.28;103.53
cluster-mode=4 pre-cluster-threshold=0.2
parse-bbox-func-name=NvDsInferParseRFDETR (custom parser, RF-DETR large, 576x576 engine)

2:

dmon-20260909-100118.log.gz (2.6 KB)

Attached: dmon-20260909-100118.log, 3 minutes at 1 s interval with the full app running (nvidia-smi dmon -s pucvmet -d 1). Averages over the 172 samples:

3:**Minimal pipeline
**
nvmultiurisrcbin → nvmultistreamtiler (5x5, 3840x2160) → nvvideoconvert → nvv4l2h264enc → h264parse → matroskamux → filesink
and inspect the recording afterwards. 24 thermal sources (1280x720, 25 fps, H.264), same nvmultiurisrcbin properties as above, no nvinfer, 180 s per run, the app stopped during the run.

side note: I didn’t use smart recorder in the deepstream function because i found a bit obsolete because you can render a frame with gstreamer in the pipeline. nvds_obj_encode

from the log, decoder utilzation is not busy( 1–30%), GPU’s compute is heavily loaded(97–100%). Important: if inference cannot keep up with incoming frames, the pipeline will be affected (queues grow, latency rises, packets may be dropped with drop-on-latency=1).
To narrow down the issue, please test “nvmultiurisrcbin-> tiler-> render/filesink sink” pipline first, please refer to the following two cmds. one is for render sink, the other is for filesink.

gst-launch-1.0 -e nvmultiurisrcbin properties=xxx \
  uri-list="rtsp://127.0.0.1:8554/stream0,rtsp://127.0.0.1:8554/stream0" \
  sensor-id-list="cam_001,cam_002" \
  ! nvmultistreamtiler rows=1 columns=2 width=3840 height=2160 \
  !  nvvideoconvert ! nveglglessink sync=false
gst-launch-1.0 -e nvmultiurisrcbin properties=xxx  \
  uri-list="rtsp://127.0.0.1:8554/stream0,rtsp://127.0.0.1:8554/stream0" \
  sensor-id-list="cam_001,cam_002" \
  ! nvmultistreamtiler rows=1 columns=2 width=1920 height=1080 \
  !  nvvideoconvert ! 'video/x-raw(memory:NVMM),format=I420' ! nvv4l2h264enc bitrate=1000000 ! filesink location=test.264

If the minimal pipeline output (nvmultiurisrcbin → tiler → sink, no nvinfer) is still corrupted for a long time, the issue is likely on the stream/receive side (RTSP, network, or camera), not inference.
If the minimal pipeline output looks fine but corruption only appear in the full pipeline with nvinfer, it is more likely related to high GPU utilization / pipeline backpressure (inference cannot keep up, queues build, packets drop or get delayed).

Additionally, do all streams show corruption, or only some streams intermittently?
To rule out per-source receive/decode issues, please test each RTSP camera individually using the pipeline below. Run the test on every camera, or at least on the ones where you observe corrupted frames / false detections.

export USE_NEW_NVSTREAMMUX=yes
gst-launch-1.0 -e \
  nvmultiurisrcbin \
    uri-list="rtsp://127.0.0.1:8554/stream0" \
    sensor-id-list="cam_01" \
    max-batch-size=1 \
    width=1280 height=720 \
    live-source=1 \
    batched-push-timeout=40000 \
    select-rtp-protocol=4 \
    disable-audio=1 \
    latency=100 \
    drop-on-latency=1 \
    rtsp-reconnect-interval=10 \
    rtsp-reconnect-attempts=-1 \
    dec-skip-frames=0 \
    drop-frame-interval=0 \
    low-latency-mode=1 \
    num-extra-surfaces=1 \
  !  nvvideoconvert  ! nveglglessink sync=false

If Step 1 still shows corruption, test the same RTSP URL with a minimal pipeline and a software decoder . This helps determine whether the issue is in Hareware decoder.

#264
 gst-launch-1.0 -e \
  rtspsrc location="rtsp://127.0.0.1:8554/stream0" protocols=tcp latency=100 \
  ! rtph264depay ! h264parse ! avdec_h264 ! nveglglessink sync=false
#265
 gst-launch-1.0 -e \
  rtspsrc location="rtsp://127.0.0.1:8554/stream0" protocols=tcp latency=100 \
  ! rtph265depay ! h265parse ! avdec_h265 ! nveglglessink sync=false

Hello @gebruiker702 ,

As an additional test, you could try disabling drop-on-latency and increasing the RTSP latency. With drop-on-latency enabled, late RTP packets may be discarded. For H.264/H.265 streams, losing data from a reference frame could result in corruption that persists until the next clean I-frame.

You could try:

drop-on-latency=0
latency=500

or temporarily increase the latency further on one of the streams where the issue is easy to reproduce. If the corruption disappears, that would suggest the issue is related to packet drops in the receive/jitterbuffer path.

Hope it helps!

Julian Camacho
Embedded SW Engineer at RidgeRun

Contact us: support@ridgerun.com
Developers wiki: https://developer.ridgerun.com/
Website: www.ridgerun.com

Thanks both. I have run the tests. Results below; they point at the full pipeline rather than the stream.

All streams or only some: only some, intermittently. Two cameras are hit far more often than the others. It sits on a mast that sways in the wind, so every frame has motion across the whole picture and its P-frames are large. But i observed the swaying of the cams without corruption in the test.

1. Properties

nvmultiurisrcbin, set with g_object_set() at build time, USE_NEW_NVSTREAMMUX=yes:

max-batch-size=35 width=1280 height=720 batched-push-timeout=40000
config-file-path=nvstreammux_config.txt
[property] algorithm-type=1 adaptive-batching=1 max-same-source-frames=1
max-fps-control=0 overall-max-fps=25/1 overall-min-fps=10/1
live-source=1 cudadec-memtype=0 drop-frame-interval=3 dec-skip-frames=0
leaky=2 max-size-buffers=3 num-extra-surfaces=1 low-latency-mode=1
drop-on-latency=1 latency=100 select-rtp-protocol=4 (TCP) udp-buffer-size=2097152
disable-audio=1 rtsp-reconnect-interval=10 rtsp-reconnect-attempts=-1
drop-pipeline-eos=1

nvinfer:

network-mode=2 (FP16) batch-size=35 process-mode=1 interval=0 gie-unique-id=1
model-color-format=0 net-scale-factor=0.0173520735727919 offsets=123.675;116.28;103.53
cluster-mode=4 pre-cluster-threshold=0.2
parse-bbox-func-name=NvDsInferParseRFDETR (custom parser, RF-DETR large, 576x576 engine)

Downstream of nvinfer there is only a fakesink (sync=0, async=0, qos=0); the app reads the metadata from a pad probe and never touches the pixels.

“Persists until the next I-frame”: the smearing starts somewhere inside a GOP and every following P-frame is predicted from the damaged reference, so the whole picture stays smeared until the next I-frame, then it is clean again. It does not survive a clean I-frame.

2. nvidia-smi dmon (attached: dmon-20260909-100118.log, 3 minutes at 1 s, full app running)

sm dec enc pwr pclk pviol gtemp fb
99.8 % 12 % (max 30 %) 5 % 149 W 962 MHz 99 % 81 °C 2964 MB

NVDEC has plenty of headroom. The SMs are at 100 % from nvinfer, the board sits on its 150 W limit the whole time and the clock is held around 960 MHz instead of the 1695 MHz boost.

3. Minimal pipeline, three variants, all clean

No display on the server, so instead of a render sink:

nvmultiurisrcbin -> nvmultistreamtiler (5x5, 3840x2160) -> nvvideoconvert -> nvv4l2h264enc -> h264parse -> matroskamux -> filesink

24 thermal sources (1280x720, 25 fps, H.264), no nvinfer, 180 s per run, the app stopped during the runs, USE_NEW_NVSTREAMMUX=yes. Three runs with the nvmultiurisrcbin properties above, differing only in:

run difference result
A production settings (leaky=2, max-size-buffers=3, drop-on-latency=1, latency=100) clean
B leaky=0, drop-on-latency=0, queue at default clean
C as B plus latency=500 (RidgeRun’s suggestion) clean

The swaying camera is clearly moving in all three recordings and never smears. I checked the recordings both by eye and with ffmpeg (about 4000 frames each, zero decode errors, sample frames inspected). The GST_DEBUG=2 logs of the three runs (attached, passwords removed) contain no rtspsrc, h264parse or decoder warnings during streaming; the only jitterbuffer line in 9 minutes is one “backward timestamps at server, schedule resync” in run B at 2:35, with no visible effect. Everything after 0:03:10 in each log is the shutdown. dmon during these runs, for contrast with the full app: sm 1 %, dec 9 % (max 78 %), enc 10 %, 88 W, pclk 1687 MHz, pviol 1 %.

So the stream, the network and the hardware decoder are fine, and I have not run the per-camera and software-decoder tests, since the multi-stream run with the production settings already comes out clean.

So far

The corruption only appears in the full pipeline with nvinfer, on a GPU whose SMs are at 100 %. That is the second of the two cases fanzh described: back-pressure, not the stream. What I found about the mechanism: in the DeepStream 9.1 container, a dot dump of the running nvurisrcbin shows that the leaky and max-size-buffers properties land on dec_que, the queue between h264parse and the decoder (rtspsrc → depay → h264parse → tee → dec_que → decodebin → queue → nvvideoconvert). When nvinfer stalls the chain, that queue fills up and, with leaky=2, discards encoded access units. One dropped P-frame gives exactly the reference-frame corruption in the screenshot, and the camera with the largest P-frames is hit first and worst. If I have misread where that queue sits, I would be glad to know.

Thanks for the update.
1. Scale the minimal pipeline to 35 sources
From section “3. Minimal pipeline, three variants, all clean”, it looks like corruption could not be reproduced with 24 RTSP sources in the minimal pipeline.
To rule out receive/decode issues more completely, please repeat the same minimal test with 35 RTSP sources, since your deployment target is 35 streams.

If the output remains clean at 35 sources, that would further confirm the stream/receive/decode path is fine and the issue is tied to the full pipeline under inference load.

2. Ways to reduce or avoid the corruption
Based on your findings (corruption only with the full pipeline when GPU SMs are saturated), please try the following:
Lower inference load

  • Set interval=1 on nvinfer — this skips one consecutive batch between inferences and reduces GPU load.
  • Reduce the number of active sources and update batch-size on both nvstreammux and nvinfer accordingly.

Avoid dropping encoded data (pre-decode)

  • drop-on-latency=0 and latency=500 on nvmultiurisrcbin — avoid discarding late RTP packets and allow more jitter buffering on the receive path
  • leaky=0 on nvmultiurisrcbin — maps to dec_que (queue between h264parse and decodebin); with no leaky mode, encoded frames are not dropped when the queue fills

Hello,

I add a similar behavior when the inference was not fast enough to match the FPS of the RTSP used.

If you want to solve it by using a leaky queue, you have to modify the queue after the decoder (src_que) so that all dropped buffers are fully-decoded.

You can use queue overrun signal to validate that the queue is the bottleneck.