DeepStream: Batching Not Occurring (Only 1 Frame per Batch Instead of 16) Causing FPS Drop with Multiple Streams

Please provide complete information as applicable to your setup.

• Hardware Platform (Jetson / GPU) : GPU (H100 NVL)
• DeepStream Version : 7.1
• JetPack Version (valid for Jetson only) :
• TensorRT Version : 10.3.0.26
• NVIDIA GPU Driver Version (valid for GPU only) : 580.82.07
• Issue Type( questions, new requirements, bugs)
• How to reproduce the issue ? (This is for bugs. Including which sample app is using, the configuration files content, the command line used and other details for reproducing)
• Requirement details( This is for new requirement. Including the module name-for which plugin or for which sample application, the function description)

Issue Summary:

  • Batching not occurring as configured (num_frames_in_batch always = 1)

  • Throughput drops when adding the 4th stream

  • Only one frame is being inferred at a time instead of a batch of 16


Description:

Hi NVIDIA Team,

We’ve developed an application using DeepStream that dynamically adds multiple input video/RTSP streams at 30-second intervals and publishes detected metadata to RabbitMQ.

Each stream runs at 30 FPS, and performance is stable up to 3 streams. However, when we add a 4th stream, throughput drops across all streams to around 24 FPS.

Upon debugging, we noticed that inference is happening for only one frame at a time, even though our configuration is intended to process a batch of 16 frames from 16 sources. The batch metadata consistently shows num_frames_in_batch = 1, indicating that batching isn’t working as expected.


Code snippet used for probing:

gst_buffer = info.get_buffer()

if not gst_buffer:

return Gst.PadProbeReturn.OK

batch_meta = pyds.gst_buffer_get_nvds_batch_meta(hash(gst_buffer))

num_frames_in_batch = batch_meta.num_frames_in_batch

print(f"Processing batch with {num_frames_in_batch} frames")

Sample output:

Processing batch with 1 frames

Processing batch with 1 frames


Expected Behavior:

We expect num_frames_in_batch to reflect up to 16 frames per batch, as configured in both nvinfer and nvstreammux. The inference model supports dynamic batching with:

  • Minimum batch size: 1

  • Maximum batch size: 16

However, the batch size remains fixed at 1 in runtime.


Configuration Details:

nvinfer Config:

[property]

gpu-id=0

onnx-file=/deepstream_app/src/deepstream/models/yolov5m.onnx

model-engine-file=/deepstream_app/src/deepstream/models/model_b16_gpu0_fp16.engine

batch-size=16

infer-dims=3;640;640

network-mode=2

num-detected-classes=80

interval=0

gie-unique-id=1

process-mode=1

network-type=0

cluster-mode=2

maintain-aspect-ratio=1

symmetric-padding=1

workspace-size=2048

parse-bbox-func-name=NvDsInferParseYolo

custom-lib-path=/DeepStream-Yolo/nvdsinfer_custom_impl_Yolo/libnvdsinfer_custom_impl_Yolo.so

engine-create-func-name=NvDsInferYoloCudaEngineGet

[class-attrs-all]

nms-iou-threshold=0.45

pre-cluster-threshold=0.25

topk=300

nvstreammux Configuration:

self.streammux = Gst.ElementFactory.make(“nvstreammux”, “stream-mux”)

self.streammux.set_property(“batch-size”, 16)

self.streammux.set_property(“width”, 1920)

self.streammux.set_property(“height”, 1080)

self.streammux.set_property(“batched-push-timeout”, 33333)

self.streammux.set_property(“live-source”, True)

self.pipeline.add(self.streammux)



Could you please help us identify:

  1. Why batching is not occurring (despite batch-size=16 in both streammux and nvinfer)?

  2. How to correctly configure the pipeline to process a batch of 16 frames and improve throughput across multiple input streams?

Can you show us the complete pipeline and all configurations with the elements in the pipeline?

Where did you put the probe function?

Have you monitored the GPU loading with “nvidia-smi dmon” while running your app?

1.WE have Probe on OSD in fakesink

2.We monitored nvidia-msi and GPU utilisation is not going beyond 50%

Pipeline Architecture:

Our application uses a dynamic pipeline architecture where:

  • A main pipeline consists of nvstreammux, nvinfer, and nvtracker ,demux components

  • Source bins and their associated components (including fakesinks) are created dynamically and linked to the main pipeline at runtime when streams are added.

  • Expected Behavior: GPU utilization should scale appropriately with the workload, especially when processing multiple streams with inference and tracking operations.

Actual Behavior: GPU utilization plateaus at ~50% regardless of the number of active streams or processing load.

Code Snippets:

1. Main Pipeline Initialization:

python

self.pipeline = Gst.Pipeline()

self.streammux = Gst.ElementFactory.make(“nvstreammux”, “stream-mux”)

self.streammux.set_property(“batch-size”, 16)

self.streammux.set_property(“width”, 1920)

self.streammux.set_property(“height”, 1080)

self.streammux.set_property(“batched-push-timeout”, 33333)

self.streammux.set_property(“live-source”, True)

self.streammux.set_property(“nvbuf-memory-type”, 0)

self.streammux.set_property(“attach-sys-ts”, True)

self.pipeline.add(self.streammux)

gie_configs = “/deepstream_app/src/deepstream/config/config_pgie_yolov5m.txt”

if not self.is_file_exists(gie_configs):

raise RuntimeError(f"File doesn't exists:{gie_configs}")

# Primary and secondary inference elements

self.gie_configs = [gie_configs]

self.gies =

previous_elm = self.streammux

for i, config in enumerate(self.gie_configs):

gie = Gst.ElementFactory.make("nvinfer", f"gie\_{i}")

if not gie:

    raise RuntimeError(f"Failed to create GIE element for config {config}")

gie.set_property("config-file-path", config)

gie.set_property("batch-size", 16)

self.pipeline.add(gie)

previous_elm.link(gie)

self.gies.append(gie)

previous_elm = gie

self.tracker = Gst.ElementFactory.make(“nvtracker”, “tracker”)

if not self.tracker:

sys.stderr.write(" Unable to create tracker \\n")

self.tracker.set_property(“tracker-width”, 640)

self.tracker.set_property(“tracker-height”, 640)

self.tracker.set_property(“gpu_id”, 0)

self.tracker.set_property(

"ll-lib-file",

"/opt/nvidia/deepstream/deepstream/lib/libnvds_nvmultiobjecttracker.so",

)

self.tracker.set_property(

"ll-config-file",

Config.app_configs.get(

    "TRACKER_CONFIG",

    "/opt/nvidia/deepstream/deepstream/samples/configs/deepstream-app/config_tracker_NvSORT.yml",

),

)

self.pipeline.add(self.tracker)

previous_elm.link(self.tracker)

self.demux = Gst.ElementFactory.make(“nvstreamdemux”, “stream-demux”)

self.pipeline.add(self.demux)

self.tracker.link(self.demux)

# Pre-create request pads on demux for potential sources

self.demux_src_pads = [

self.demux.get_request_pad(f"src\_{i}") for i in range(self.max_sources)

]

2. Dynamic Source Addition Method:

python

def add_source(

self,

uri: str,

device_id: str,

alert_type_id: str,

rtsp_output_width: int = 640,

rtsp_output_height: int = 640,

) → int:

"""

Add a new source to the pipeline and create an RTSP mount for it.

"""

if self.pipeline.get_state(1).state != Gst.State.PLAYING:

    raise RuntimeError(

        "Pipeline is not running. Start the pipeline before adding sources."

    )



if uri is None:

    raise RuntimeError("No uri is provided")

    if uri.startswith("file:///"):

        if not re.fullmatch(r"file:///\[^ \]+", uri):

            raise RuntimeError(f"Invalid or empty file URI: {uri}")

        file_path = uri\[7:\]

        if not os.path.isfile(file_path):

            raise RuntimeError(f"File does not exist: {file_path}")

    elif (

        self.check_rtsp_link(uri) is False

        or not uri.startswith("rtsp://")

    ):

        raise RuntimeError(f"Invalid RTSP link: {uri}")



spot, uuid, is_fresh = self.spot_manager.acquire(device_id, alert_type_id)

if spot is None:

    raise RuntimeError("No available spots for new source")



*# Create and link source bin*

if uri.startswith("rtsp://") or uri.startswith("https://"):

    src_bin = self.source_bin_factory.create_source_bin(uuid, uri, "rtspsrc")

else:

    src_bin = self.source_bin_factory.create_source_bin(uuid, uri, "nvurisrcbin")



self.pipeline.add(src_bin)

src_bin.set_state(Gst.State.PAUSED)

src_bin.get_state(Gst.CLOCK_TIME_NONE)

src_pad = src_bin.get_static_pad("src")



if is_fresh:

    mux_pad = self.streammux.get_request_pad(f"sink\_{spot}")

else:

    mux_pad = self.streammux.get_static_pad(f"sink\_{spot}")

   

if not mux_pad:

    raise RuntimeError(

        f"Failed to get request pad sink\_{uuid} — maybe not released?"

    )



self.pad_to_index\[src_pad\] = spot

src_pad.link(mux_pad)

self.sources\[spot\] = src_bin



*# Create queue and processing thread for this source*

if not hasattr(self, "source_queues"):

    self.source_queues = {}

if not hasattr(self, "source_threads"):

    self.source_threads = {}



self.logger.info(f"\[DEBUG\] Creating queue and thread for source {uri} with {spot}")

source_queue = Queue()

self.source_queues\[spot\] = source_queue

processing_thread = threading.Thread(

    target=self.\_processing_worker_loop, args=(source_queue, spot), daemon=False

)

self.source_threads\[spot\] = processing_thread

processing_thread.start()



*# Build per stream output branch*

self.\_setup_fakesink_sink(spot, uuid)



src_bin.sync_state_with_parent()

src_bin.set_state(Gst.State.PLAYING)

3. Fakesink Setup for Metadata Probing:

python

def _setup_fakesink_sink(self, index: int, uuid: str):

"""

Setup fakesink branch for metadata extraction from a specific stream.

This branch probes detection metadata from OSD without video output.

"""

*# Create minimal pipeline for metadata extraction*

conv = Gst.ElementFactory.make("nvvideoconvert", f"conv_rabbitmq\_{uuid}")

capsfilter = Gst.ElementFactory.make("capsfilter", f"capsfilter_rabbitmq\_{uuid}")

osd = Gst.ElementFactory.make("nvdsosd", f"osd_rabbitmq\_{uuid}")

fakesink = Gst.ElementFactory.make("fakesink", f"fakesink\_{uuid}")



if not conv or not capsfilter or not osd or not fakesink:

    print(f"\[ERROR\] Failed to create elements for fakesink {uuid}")

    return



*# Configure OSD*

osd.set_property("display-bbox", 1)

osd.set_property("display-mask", 0)



*# Configure fakesink*

fakesink.set_property("sync", 1)

fakesink.set_property("async", False)



conv.set_property("nvbuf-memory-type", int(pyds.NVBUF_MEM_CUDA_UNIFIED))

capsfilter.set_property("caps", Gst.Caps.from_string("video/x-raw(memory:NVMM), format=RGBA"))



self.pipeline.add(conv)

self.pipeline.add(capsfilter)

self.pipeline.add(osd)

self.pipeline.add(fakesink)



conv.sync_state_with_parent()

capsfilter.sync_state_with_parent()

osd.sync_state_with_parent()

fakesink.sync_state_with_parent()



*# Link: demux -> conv -> capsfilter -> osd -> fakesink*

self.demux_src_pads\[index\].link(conv.get_static_pad("sink"))

conv.link(capsfilter)

capsfilter.link(osd)

osd.link(fakesink)



print("Added probe buffer")

osd.get_static_pad("sink").add_probe(

    Gst.PadProbeType.BUFFER, self.conv_pad_buffer_probe, 0

)



self.branches\[index\] = {

    "osd": osd,

    "sink": fakesink,

    "type": "rabbitmq"

}

After summarizing your pieces of descriptions, I think your probe function get only one frame in the batch is absolutely right. Batches only exist after nvstreammux and before nvmultistreamtiler or nvstreamdemux. Your probe function is added after nvstreamdemux, there is no real “batch” but just a batch meta data struct here. Please check with DeepStream document to understand what the “batch” actually is.

Where and how did you do this? Is the FPS the same if you remove this part from your app?

There is no update from you for a period, assuming this is not an issue anymore. Hence we are closing this topic. If need further support, please open a new one. Thanks.