Hi,
You can launch trtexec in the different consoles for parallel inference.
$ /usr/src/tensorrt/bin/trtexec --onnx=/usr/src/tensorrt/data/mnist/mnist.onnx --useDLACore=0 --allowGPUFallback
$ /usr/src/tensorrt/bin/trtexec --onnx=/usr/src/tensorrt/data/mnist/mnist.onnx --useDLACore=1 --allowGPUFallback
If you are finding a multi-thread inference example, please check the below comment:
We create two engines for the same model with the corresponding context to make it parallel.
Thanks