# Question about TensorRT reproducibility on different architectures

**URL:** <https://forums.developer.nvidia.com/t/question-about-tensorrt-reproducibility-on-different-architectures/189129>\
**Category:** TensorRT\
**Created:** [September 13, 2021, 4:45pm UTC](https://forums.developer.nvidia.com/t/question-about-tensorrt-reproducibility-on-different-architectures/189129 "2021-09-13T16:45:31Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![nk\_white](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@nk\_white](https://forums.developer.nvidia.com/u/nk_white)\
**Post date:** [September 13, 2021, 4:45pm UTC](https://forums.developer.nvidia.com/t/question-about-tensorrt-reproducibility-on-different-architectures/189129/1 "2021-09-13T16:45:31Z")

</div>

## Description

I have a general question about the expectation of TensorRT engine files.

I used Transfer Learning Toolkit/TAO Toolkit on a dGPU to train a custom YOLOv3 model. I iterated through several pruning and retraining steps until I had a model that worked really well on my training set and test sets. I then created a .etlt file and finally a TensorRT .engine file and ran inference on a test set on the dGPU and was happy with the results. Then, I moved the model to the Jetson Nano by creating a .engine file on the Nano, using the same tlt-converter command line as I did with the dGPU. However, the inference results - as measured by eye (bounding boxes in Deepstream 5.1) were much worse than on the dGPU. I don’t know of an easy way to just do inference on the Nano using a .engine file to test the results quantitatively. But my question is this: **Should TensorRT engines produce the same results on different platforms?** If so, then I will continue to try to figure out how to quantitatively measure inference on the TLT YOLOv3 model outside of Deepstream 5.1. If not, then how does one select the proper model on the dGPU to deploy to edge devices?

## Environment

**TensorRT Version** : 7.2.1  
**GPU Type** : Tesla V100  
**Nvidia Driver Version** : 460.32.03  
**CUDA Version** :  
**CUDNN Version** :  
**Operating System + Version** : CentOS7  
**Python Version (if applicable)**:  
**TensorFlow Version (if applicable)**:  
**PyTorch Version (if applicable)**:  
**Baremetal or Container (if container which image + tag)**:

## Relevant Files

Please attach or include links to any models, data, files, or scripts necessary to reproduce your issue. (Github repo, Google Drive, Dropbox, etc.)

## Steps To Reproduce

Please include:

- Exact steps/commands to build your repro
- Exact steps/commands to run your repro
- Full traceback of errors encountered

---

<div class="post-metadata">

**Author:** ![spolisetty](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/spolisetty/32/14047_2.png) [@spolisetty](https://forums.developer.nvidia.com/u/spolisetty)\
**Post date:** [September 15, 2021, 5:54am UTC](https://forums.developer.nvidia.com/t/question-about-tensorrt-reproducibility-on-different-architectures/189129/2 "2021-09-15T05:54:58Z")

</div>

Hi @nk_white,

TensorRT engines do not produce the same results on different platforms. Please refer following post to know more details on your query.

> [@Is TensorRT inference deterministic/reproducibile?](https://forums.developer.nvidia.com/t/is-tensorrt-inference-deterministic-reproducibile/159955/2):
>
> Hi @daniel.widmann, If you are using same engine with same input, TensorRT should be deterministic. However I don’t think engine building is supposed to be deterministic as tactics are chosen based on observed runtime. If you’re outputting your log with info level, you should be able to compare tactic selection between the two engines. Since different tactics/kernels could change order of operations, you would expect floating point differences. You can refer to the below link. Thanks!

Thank you.

---

<div class="post-metadata">

**Author:** ![nk\_white](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@nk\_white](https://forums.developer.nvidia.com/u/nk_white)\
**Post date:** [September 16, 2021, 10:54pm UTC](https://forums.developer.nvidia.com/t/question-about-tensorrt-reproducibility-on-different-architectures/189129/3 "2021-09-16T22:54:30Z")

</div>

Thank you for your reply. From the explanation, I interpret “floating point differences” to be qualitatively equal. The prediction performance I’m seeing is not qualitatively equal so I will continue on in an effort to figure out why my model is performing so poorly in DeepStream 5.1.

---

<div class="post-metadata">

**Author:** ![TomNVIDIA](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/tomnvidia/32/14181_2.png) [@TomNVIDIA](https://forums.developer.nvidia.com/u/TomNVIDIA)\
**Post date:** [October 12, 2021, 7:47pm UTC](https://forums.developer.nvidia.com/t/question-about-tensorrt-reproducibility-on-different-architectures/189129/6 "2021-10-12T19:47:16Z")

</div>


