# Run to run variation with TensorRT

**URL:** <https://forums.developer.nvidia.com/t/run-to-run-variation-with-tensorrt/226526>\
**Category:** TensorRT\
**Tags:** tensorrt\
**Created:** [August 30, 2022, 4:21pm UTC](https://forums.developer.nvidia.com/t/run-to-run-variation-with-tensorrt/226526 "2022-08-30T16:21:08Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![DML](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@DML](https://forums.developer.nvidia.com/u/DML)\
**Post date:** [August 30, 2022, 4:21pm UTC](https://forums.developer.nvidia.com/t/run-to-run-variation-with-tensorrt/226526/1 "2022-08-30T16:21:08Z")

</div>

Description: Run to run variation with TensorRT

Environment:  
NVIDIA Release: 22.07  
NVIDIA TensorRT Version: 8.4.1  
NVIDIA Driver Version: 515.43.04  
CUDA Version: 11.7  
NVIDIA GPU: NVIDIA Tesla T4  
Docker Image: [nvcr.io/nvidia/tensorrt:22.07-py3](http://nvcr.io/nvidia/tensorrt:22.07-py3)

To get the inference data using trtexec. there are two steps involved.

1. Build a TRT engines from a model
2. Get a inference performance metrics by loading TRT engines

I see less than 1% variation once TRT engines are built from a model and perform inference stage multiple times.  
If I perform step 1 and step 2 multiple times for the same model with same configs, I see variation up to 3% in inference throughput. Is it normal?

If I build a TRT engine for the same model multiple times (with same configs) then should trtexec generate a TRT engine with same size?

---

<div class="post-metadata">

**Author:** ![spolisetty](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/spolisetty/32/14047_2.png) [@spolisetty](https://forums.developer.nvidia.com/u/spolisetty)\
**Post date:** [September 2, 2022, 11:17am UTC](https://forums.developer.nvidia.com/t/run-to-run-variation-with-tensorrt/226526/4 "2022-09-02T11:17:48Z")

</div>

Hi,

The builder times kernels to find the fastest, and sometimes if the timings are close between two different precisions, due to timing noise the builder may choose differently on different runs. So 3% variation in runtime is not necessarily unusual, and possibly some variation in engine size.

Thank you.
