# Multiple rtsp with multiple nvinfer

**URL:** <https://forums.developer.nvidia.com/t/multiple-rtsp-with-multiple-nvinfer/372009>\
**Category:** DeepStream SDK\
**Tags:** jetson, deepstream\
**Created:** [June 2, 2026, 11:59am UTC](https://forums.developer.nvidia.com/t/multiple-rtsp-with-multiple-nvinfer/372009 "2026-06-02T11:59:54Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![hritik.shah](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@hritik.shah](https://forums.developer.nvidia.com/u/hritik.shah)\
**Post date:** [June 2, 2026, 11:59am UTC](https://forums.developer.nvidia.com/t/multiple-rtsp-with-multiple-nvinfer/372009/1 "2026-06-02T11:59:54Z")

</div>

**Hardware Platform:** Jetson Orin Nano 8GB  
**DeepStream Version:** 7.1  
**JetPack Version:** 6.1 (L4T R36.4)  
**TensorRT Version:** 10.3  
**Issue Type:** Optimization Question

i have multiple rtsp sources and multiple services(cv models) , my goal is to maximize the number of cameras and services i can run on a single jetson

this is what i am doing:  
rtspsrc → nvv4l2decoder → tee ─┬──-\> service\_mux\_A (batch=N) → nvinfer\_A → probe  
└──-\> service\_mux\_B (batch=N) → nvinfer\_B → probe

- 1 decoder per camera (shared across all services on that camera)

- 1 nvinfer per service type (shared across all cameras)

- nvstreammux `batch-size` capped at 1 (TRT engines built with `maxBatchSize=1`)

- CPU RGBA output from nvvideoconvert (no NVMM) to avoid VIC exhaustion

- New camera hot-adds a decoder and connects its tee to all active service muxes simultaneously

can this be optimized to get a better throughput in any way???

---

<div class="post-metadata">

**Author:** ![MarkusHoHo](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/markushoho/32/233145_2.png) [@MarkusHoHo](https://forums.developer.nvidia.com/u/MarkusHoHo)\
**Post date:** [June 2, 2026, 12:02pm UTC](https://forums.developer.nvidia.com/t/multiple-rtsp-with-multiple-nvinfer/372009/2 "2026-06-02T12:02:01Z")

</div>

Hello @hritik.shah!

Based on the title and content of your topic, it looks like it may receive better visibility and feedback in a different category. We took the liberty of moving it for you.

If this was an incorrect assessment, please send me a direct message.

_Disclaimer: this moderation suggestion and message were generated with AI assistance._

---

<div class="post-metadata">

**Author:** ![Fiona.Chen](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/fiona.chen/32/15508_2.png) [@Fiona.Chen](https://forums.developer.nvidia.com/u/Fiona.Chen)\
**Post date:** [June 3, 2026, 10:12am UTC](https://forums.developer.nvidia.com/t/multiple-rtsp-with-multiple-nvinfer/372009/6 "2026-06-03T10:12:26Z")

</div>

> [@hritik.shah](#):
>
> - 1 decoder per camera (shared across all services on that camera)
> - 1 nvinfer per service type (shared across all cameras)
> - nvstreammux `batch-size` capped at 1 (TRT engines built with `maxBatchSize=1`)
> - CPU RGBA output from nvvideoconvert (no NVMM) to avoid VIC exhaustion
> - New camera hot-adds a decoder and connects its tee to all active service muxes simultaneously

Do you have multiple RTSP streams(cameras) to be added to the ineference pipeline dynamically? The sample /opt/nvidia/deepstream/deepstream/sources/apps/sample\_apps/deepstream-server can handle such case. Do you know the maximum number of the streams(cameras)?

What is the relationship bwteeen the two models “nvinfer\_A” and “nvinfer\_B”? If the two models will both infer all input video streams, only one “nvstreammux” is enough.

Where and how do you implement “CPU RGBA output from nvvideoconvert (no NVMM)” in the pipeline?

What do you do in the “probe”？

What do you mean by “get a better throughput”? The FPS value?

---

<div class="post-metadata">

**Author:** ![hritik.shah](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@hritik.shah](https://forums.developer.nvidia.com/u/hritik.shah)\
**Post date:** [June 4, 2026, 4:44am UTC](https://forums.developer.nvidia.com/t/multiple-rtsp-with-multiple-nvinfer/372009/7 "2026-06-04T04:44:11Z")

</div>

yes i have multiple rtsp streams to be added dynamically.

i dont know the maximum number of streams, thats what i want to know, i want the maximum possible streams.

both models are different, and the number of streams both of them infer is also dynamic , i add / remove streams during runtime

this is how i use nvvideoconvert  
rtspsrc → nvv4l2decoder(num-extra-surfaces=0) → tee  
tee → queue(leaky, max=2) → nvstreammux(batch=N) → nvinfer → nvvideoconvert(compute-hw=1)  
 → capsfilter(video/x-raw,RGBA)

better throughput is number of streams i can add

also:

1. how does batch\_size in the infer config matter to the number of streams the engine can handle?
2. i had an issue of “failed in mem copy”, which got fixed by using copy-hw=1, scaling-compute-hw=1, is this a correct fix?
3. then i rebuilt the engines with batch\_size=32, and on test , upto 32-33 rtsp streams got added on the same infer, then i got this error “libnvrm\_gpu.so: NvRmGpuLibOpen failed, error=6”, so is batch\_size the maximum number of streams i can add? if yes can i change the batch\_size at runtime without having to rebuild the engine pre-run?

---

<div class="post-metadata">

**Author:** ![Fiona.Chen](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/fiona.chen/32/15508_2.png) [@Fiona.Chen](https://forums.developer.nvidia.com/u/Fiona.Chen)\
**Post date:** [June 4, 2026, 6:33am UTC](https://forums.developer.nvidia.com/t/multiple-rtsp-with-multiple-nvinfer/372009/8 "2026-06-04T06:33:30Z")

</div>

From the DeepStream pipeline view, the maximum streams number it can support depends on the slowest part in the pipeline. You need to find out the bottleneck in your pipeline by yourself.

1. Jetson Orin Nano hardware decoder capability is listed in [Jetson AGX Orin for Next-Gen Robotics | NVIDIA](https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-orin/)
2. The model TRT engine performance with different batch size can be measured by the TensorRT tool “trtexec”
3. What did you do with the RGBA data after “nvvideoconvert” ?

> [@hritik.shah](#):
>
> both models are different, and the number of streams both of them infer is also dynamic , i add / remove streams during runtime

Can you elaborate it clearly? The streams will be added/removed dynamically, but we want to know whether the two models inference on exactly the same streams at the same moment. E.G. when there are 5 streams added to the pipeline, will the model A inference on stream 1,2,3 while model B inference on stream 3, 4, 5? Or both model A and model B will infer on stream 1,2,3,4,5?

> [@hritik.shah](#):
>
> how does batch\_size in the infer config matter to the number of streams the engine can handle?

The nvinfer batch size is the TensorRT model engine batch size. If your model is built to batch size 32 engine, that means the engine can infer at most 32 frames at one time. If you build the batch size 1 model engine, you need to infer 32 times with the engine for 32 frames. Most models we have tried show that to infer 32 frames with batch size 32 engine for one time is faster than infer 32 frames with batch size 1 engine for 32 times. We don’t know about your models, you may need to measure the models by yourself.

> [@hritik.shah](#):
>
> i had an issue of “failed in mem copy”, which got fixed by using copy-hw=1, scaling-compute-hw=1, is this a correct fix?

It works.

> [@hritik.shah](#):
>
> so is batch\_size the maximum number of streams i can add?

No. I think I have explained the maximum number of streams depends on your pipeline.

> [@hritik.shah](#):
>
> can i change the batch\_size at runtime without having to rebuild the engine pre-run?

No.

---

<div class="post-metadata">

**Author:** ![hritik.shah](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@hritik.shah](https://forums.developer.nvidia.com/u/hritik.shah)\
**Post date:** [June 4, 2026, 6:51am UTC](https://forums.developer.nvidia.com/t/multiple-rtsp-with-multiple-nvinfer/372009/9 "2026-06-04T06:51:34Z")

</div>

> [@Fiona.Chen](#):
>
> Can you elaborate it clearly?

yes both models inference on exactly the same streams at the same moment, both model A and model B will infer on stream 1,2,3,4,5

> [@Fiona.Chen](#):
>
> No. I think I have explained the maximum number of streams depends on your pipeline.

yes but i got this error “libnvrm\_gpu.so: NvRmGpuLibOpen failed, error=6” exactly when 33rd stream is added with the batch\_size=32 engine for both the models tested seperately , and the ram didnt actually exhaust, my models are Yolov11n and RF-DETRs .

1. what does this error mean “libnvrm\_gpu.so: NvRmGpuLibOpen failed, error=6”
2. if i load the next streams after 32 streams on another infer of the same model , will that work and increase the number of streams?
3. how will the config parameter “interval” change the infer in my usecase, and what is the best suggested interval , considering all my streams are running at 20fps

---

<div class="post-metadata">

**Author:** ![Fiona.Chen](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/fiona.chen/32/15508_2.png) [@Fiona.Chen](https://forums.developer.nvidia.com/u/Fiona.Chen)\
**Post date:** [June 5, 2026, 3:07am UTC](https://forums.developer.nvidia.com/t/multiple-rtsp-with-multiple-nvinfer/372009/10 "2026-06-05T03:07:23Z")

</div>

> [@hritik.shah](#):
>
> yes both models inference on exactly the same streams at the same moment, both model A and model B will infer on stream 1,2,3,4,5

Please use only one nvstreammux for your case.

> [@hritik.shah](#):
>
> what does this error mean “libnvrm\_gpu.so: NvRmGpuLibOpen failed, error=6”

It seems the fd exhaust. Please try “ulimit -n 4096”

> [@hritik.shah](#):
>
> if i load the next streams after 32 streams on another infer of the same model , will that work and increase the number of streams?

Do you mean to add another pipeline? I think I have said the pipeline capability is decided by the slowest part, if the second pipeline shares the same resources, nothing will be changed.

> [@hritik.shah](#):
>
> how will the config parameter “interval” change the infer in my usecase, and what is the best suggested interval , considering all my streams are running at 20fps

The “interval” parameter is to skip the inference on some batches. If the bottleneck is the GPU loading of your models, it may help to improve the throughput. Every model is different, the different batch size TensorRT engines for the same model are different. The same model runs on different GPUs are different. The value is decided by your pipeline and GPU loading, you need to measure it by yourself.

---

<div class="post-metadata">

**Author:** ![yingliu](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/yingliu/32/136703_2.png) [@yingliu](https://forums.developer.nvidia.com/u/yingliu)\
**Post date:** [June 23, 2026, 5:56am UTC](https://forums.developer.nvidia.com/t/multiple-rtsp-with-multiple-nvinfer/372009/11 "2026-06-23T05:56:05Z")

</div>

There is no update from you for a period, assuming this is not an issue anymore. Hence we are closing this topic. If need further support, please open a new one. Thanks.

---

<div class="post-metadata">

**Author:** ![system](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/system/32/68080_2.png) [@system](https://forums.developer.nvidia.com/u/system)\
**Post date:** [July 7, 2026, 5:56am UTC](https://forums.developer.nvidia.com/t/multiple-rtsp-with-multiple-nvinfer/372009/12 "2026-07-07T05:56:39Z")

</div>

This topic was automatically closed 14 days after the last reply. New replies are no longer allowed.
