Please provide complete information as applicable to your setup.
• Hardware Platform (Jetson / GPU) : GPU (H100 NVL)
• DeepStream Version : 7.1
• JetPack Version (valid for Jetson only) :
• TensorRT Version : 10.3.0.26
• NVIDIA GPU Driver Version (valid for GPU only) : 580.82.07
• Issue Type( questions, new requirements, bugs)
• How to reproduce the issue ? (This is for bugs. Including which sample app is using, the configuration files content, the command line used and other details for reproducing)
• Requirement details( This is for new requirement. Including the module name-for which plugin or for which sample application, the function description)
Issue Summary:
-
Batching not occurring as configured (num_frames_in_batch always = 1)
-
Throughput drops when adding the 4th stream
-
Only one frame is being inferred at a time instead of a batch of 16
Description:
Hi NVIDIA Team,
We’ve developed an application using DeepStream that dynamically adds multiple input video/RTSP streams at 30-second intervals and publishes detected metadata to RabbitMQ.
Each stream runs at 30 FPS, and performance is stable up to 3 streams. However, when we add a 4th stream, throughput drops across all streams to around 24 FPS.
Upon debugging, we noticed that inference is happening for only one frame at a time, even though our configuration is intended to process a batch of 16 frames from 16 sources. The batch metadata consistently shows num_frames_in_batch = 1, indicating that batching isn’t working as expected.
Code snippet used for probing:
gst_buffer = info.get_buffer()
if not gst_buffer:
return Gst.PadProbeReturn.OK
batch_meta = pyds.gst_buffer_get_nvds_batch_meta(hash(gst_buffer))
num_frames_in_batch = batch_meta.num_frames_in_batch
print(f"Processing batch with {num_frames_in_batch} frames")
Sample output:
Processing batch with 1 frames
Processing batch with 1 frames
Expected Behavior:
We expect num_frames_in_batch to reflect up to 16 frames per batch, as configured in both nvinfer and nvstreammux. The inference model supports dynamic batching with:
-
Minimum batch size: 1
-
Maximum batch size: 16
However, the batch size remains fixed at 1 in runtime.
Configuration Details:
nvinfer Config:
[property]
gpu-id=0
onnx-file=/deepstream_app/src/deepstream/models/yolov5m.onnx
model-engine-file=/deepstream_app/src/deepstream/models/model_b16_gpu0_fp16.engine
batch-size=16
infer-dims=3;640;640
network-mode=2
num-detected-classes=80
interval=0
gie-unique-id=1
process-mode=1
network-type=0
cluster-mode=2
maintain-aspect-ratio=1
symmetric-padding=1
workspace-size=2048
parse-bbox-func-name=NvDsInferParseYolo
custom-lib-path=/DeepStream-Yolo/nvdsinfer_custom_impl_Yolo/libnvdsinfer_custom_impl_Yolo.so
engine-create-func-name=NvDsInferYoloCudaEngineGet
[class-attrs-all]
nms-iou-threshold=0.45
pre-cluster-threshold=0.25
topk=300
nvstreammux Configuration:
self.streammux = Gst.ElementFactory.make(“nvstreammux”, “stream-mux”)
self.streammux.set_property(“batch-size”, 16)
self.streammux.set_property(“width”, 1920)
self.streammux.set_property(“height”, 1080)
self.streammux.set_property(“batched-push-timeout”, 33333)
self.streammux.set_property(“live-source”, True)
self.pipeline.add(self.streammux)
Could you please help us identify:
-
Why batching is not occurring (despite batch-size=16 in both streammux and nvinfer)?
-
How to correctly configure the pipeline to process a batch of 16 frames and improve throughput across multiple input streams?