Artifacts and Jitter on RTSP Input caused by Pipeline Back-pressure

• Hardware Platform (Jetson / GPU): Jetson
• DeepStream Version: 6.3
• JetPack Version (valid for Jetson only): 5.1

I am referencing this older topic (https://forums.developer.nvidia.com/t/jitter-noise-in-rtsp-input-frame/174288) which describes a problem that seems to persist in DeepStream 6.3.

When the downstream pipeline (inference + post-processing) is slower than the RTSP input framerate, severe visual artifacts (macro-blocking, smearing) appear on the decoded frames. It seems that when the pipeline exerts back-pressure on the nvv4l2decoder, the decoder state gets corrupted instead of handling the delay cleanly.

Current Observations: I have analyzed the behavior and found the following workarounds, none of which are sustainable for a production environment

  • Forcing I-Frames:
    • If I only decode I-Frames (intra-decode-enable):
      • The output stream drops to a low FPS, because I am effectively processing far fewer frames.

      • This avoids artifacts, but the system becomes too slow for real-time use.

    • If the camera sends only I-Frames:
      • The network bandwidth increases dramatically, since I-Frames are much heavier than P-Frames.

      • This is unacceptable for our deployment constraints.

  • Matching FPS (interval / drop-frame-interval): Manually tuning the interval in the GIE or using drop-frame-interval in the decoder to match the exact processing capability of the hardware resolves the artifacts. This is a brittle solution. Inference load varies, and hardcoding a drop interval results in a low-FPS output even when the system could handle more, or artifacts returning if the load spikes.

I am looking for a way to make the decoder sending good quality frames even if the downstream is busy, without corrupting the frames and causing visual artifacts.

Have you identify the bottleneck of the downstream elements? Is it the nvinfer(GIE) which is much slower than the original RTSP FPS?

The nvinfer “interval” property will not cause frame rate drop. And the “interval” property of nvinfer can be dynamically set. You may refer to the source code of gst-nvinfer.

Thanks for the reply.

To answer your question: Yes, nvstreammux and nvinfer are the main bottlenecks causing the latency.

I appreciate the hint regarding the dynamic interval in gst-nvinfer, and I will investigate the source code. However, this suggests that the only solution is to align the input and processing FPS.

My core question remains: Is it inevitable to have decoder artifacts when the pipeline back-pressure occurs?

I gathered some data using NVDS_ENABLE_COMPONENT_LATENCY_MEASUREMENT=1 on an Orin NX:

  • Decoder latency: ~2ms per frame.

  • Streammux/PGIE latency: ~300ms+ per batch.

The hardware is clearly capable of decoding the incoming RTSP stream at full speed. The issue seems to be how nvv4l2decoder handles the back-pressure from the slower downstream elements.

I have two hypotheses regarding the nvv4l2decoder behavior. Could you confirm which one is correct?

  1. Asynchronous (Decoupled): The decoder processes every incoming RTSP packet to maintain its internal Decoded Picture Buffer state, but only pushes the latest available NV12 buffer when the downstream element is ready.
    • If this were true, I should see frame drops (skipping), but clean images.
  2. Synchronous (Blocked): The downstream pressure literally blocks the decoder thread. When the decoder unblocks, it picks up the next available packet from the RTSP source. If that packet is a P-Frame but the decoder was forced to skip the previous Reference Frames (I/P) while blocked, the decoding becomes corrupted (artifacts).

My observations strongly point to Hypothesis #2. I would like your opinion on the mechanics of this issue

If the downstream is always slower than upstream, the upstream needs more and more extra buffering space to store the late data. The accumulated late data eventually exceeded the buffering space and be dropped. Seems in your case, this data dropping happens with the RTSP payload buffering. No component has unlimited buffering space. Either the decoded frames should be dropped, or the compressed encoded video data should be dropped.

I fully understand the logic you described: if the pipeline is slow, buffers fill up, and data must eventually be dropped.

To be clear: I actually want frame dropping to occur. My goal is to maintain real-time latency, so skipping frames is acceptable and expected. This is actually what is happening right now: even though I observe jitter and artifacts, I am not accumulating latency relative to the real-time stream.

I assume these buffers already exist within the DeepStream app structure?

I simply do not understand why the frames that are NOT dropped are of such bad quality.

The problem is what is being dropped. Currently, it seems the back-pressure is causing drops at the RTSP/Network layer (Encoded Data).

  • Because H.264/H.265 relies on temporal compression (P-frames referencing previous frames), dropping a packet before the decoder breaks the GOP structure.

  • This explains why the frames that do get processed are visually corrupted (artifacts/smearing): the decoder is trying to predict an image based on missing reference data.

Even with the “interval” property of nvinfer, the udp packets accumulated in the rtsp buffering may also exceed the limitation for the very short time when the nvinfer is working. You may try to set larger udp buffering size or try to use tcp instead of udp to receive rtsp payload.

Since the encoded payload size varies through time, the suitable buffering size and latency may vary too. It is not guaranteed.

Thank you for the suggestions regarding buffering.

To confirm, we are already configured for TCP-only streaming, and the back-pressure artifacts still occur.

  1. Could you please confirm if this artifact-generating behavior is considered a known or expected characteristic of the nvv4l2decoder when it receives back-pressure?
  2. To ensure frame quality, we are considering a pipeline design where the decoder is fully decoupled:
    • Force the decoder to process all encoded packets (generating clean NV12 frames).

    • Place a fixed-size, leaky queue immediately after the decoder.

    • The queue then always provides the latest clean NV12 frame to the downstream pipeline, dropping older frames if the buffer is full.

Does this logic align with the expected behavior of the DeepStream pipeline? Is this a logical way to ensure decoded frame quality while accepting necessary frame drops? Is it possible to do that in the deepstream default app ?

No

If you drop the decoded frames by yourself, it will not conflict with DeepStream itself.

DeepStream sample apps are all samples to demonstrate the usage of DeepStream components and interfaces. It is OK to customize your own apps according to your requirements.

I would like to conclude by expressing a general observation.

I believe it would be beneficial for the developer community if these forum threads could evolve towards deeper reflections on the underlying technologies (e.g., GStreamer, network interaction, RTSP) rather than strictly maintaining a focus limited only to DeepStream components.

DeepStream is introduced into a larger technical ecosystem and constantly interacts with other components that influence its performance. Limiting the reflections solely to DeepStream often restricts the depth of analysis required to resolve complex root causes like this one.

Thank you again for the time and answers provided throughout this discussion.

Sincerely.