# LPRNet training and deployment

**URL:** <https://forums.developer.nvidia.com/t/lprnet-training-and-deployment/295332>\
**Category:** TAO Toolkit\
**Created:** [June 6, 2024, 5:22am UTC](https://forums.developer.nvidia.com/t/lprnet-training-and-deployment/295332 "2024-06-06T05:22:07Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![mainak1](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@mainak1](https://forums.developer.nvidia.com/u/mainak1)\
**Post date:** [June 6, 2024, 5:22am UTC](https://forums.developer.nvidia.com/t/lprnet-training-and-deployment/295332/1 "2024-06-06T05:22:07Z")

</div>

Please provide the following information when requesting support.

• Hardware (T4/V100/Xavier/Nano/etc) : rtx 4090/t4(6.3 DS) and Nano(6.0 DS)  
• Network Type (Detectnet\_v2/Faster\_rcnn/Yolo\_v4/LPRnet/Mask\_rcnn/Classification/etc) : LPRNet

While training how to save weights based on best metrics (like accuracy) rather than after every 5 epochs?

I have trained a model using tao toolkit and the weight files are as `.hdf5` format. How to convert this to `etlt` so as to deploy in the pipeline?

I generated the `onnx` file from `hdf5`, but when I use that `onnx` in DS pipeline to generate the `engine` file error pops up:

```auto
Using file: ./models/anpr_config.yml
0:00:00.583616331 221656 0x560ff81ece40 INFO nvinfer gstnvinfer.cpp:682:gst_nvinfer_logger:<secondary-infer-engine2> NvDsInferContext[UID 3]: Info from NvDsInferContextImpl::buildModel() <nvdsinfer_context_impl.cpp:2002> [UID = 3]: Trying to create engine from model files
[libprotobuf ERROR google/protobuf/text_format.cc:298] Error parsing text-format onnx2trt_onnx.ModelProto: 1:1: Invalid control characters encountered in text.
[libprotobuf ERROR google/protobuf/text_format.cc:298] Error parsing text-format onnx2trt_onnx.ModelProto: 2:11: Invalid control characters encountered in text.
[libprotobuf ERROR google/protobuf/text_format.cc:298] Error parsing text-format onnx2trt_onnx.ModelProto: 2:17: Already saw decimal point or exponent; can't have another one.
[libprotobuf ERROR google/protobuf/text_format.cc:298] Error parsing text-format onnx2trt_onnx.ModelProto: 2:13: Message type "onnx2trt_onnx.ModelProto" has no field named "keras2onnx".
ERROR: [TRT]: ModelImporter.cpp:688: Failed to parse ONNX model from file: /home/mainak/ms/C++/anpr_kp/models/lprnet_epoch-024.onnx
ERROR: ../nvdsinfer/nvdsinfer_model_builder.cpp:315 Failed to parse onnx file
ERROR: ../nvdsinfer/nvdsinfer_model_builder.cpp:971 failed to build network since parsing model errors.
ERROR: ../nvdsinfer/nvdsinfer_model_builder.cpp:804 failed to build network.
0:00:02.577464454 221656 0x560ff81ece40 ERROR nvinfer gstnvinfer.cpp:676:gst_nvinfer_logger:<secondary-infer-engine2> NvDsInferContext[UID 3]: Error in NvDsInferContextImpl::buildModel() <nvdsinfer_context_impl.cpp:2022> [UID = 3]: build engine file failed
0:00:02.615730998 221656 0x560ff81ece40 ERROR nvinfer gstnvinfer.cpp:676:gst_nvinfer_logger:<secondary-infer-engine2> NvDsInferContext[UID 3]: Error in NvDsInferContextImpl::generateBackendContext() <nvdsinfer_context_impl.cpp:2108> [UID = 3]: build backend context failed
0:00:02.615758235 221656 0x560ff81ece40 ERROR nvinfer gstnvinfer.cpp:676:gst_nvinfer_logger:<secondary-infer-engine2> NvDsInferContext[UID 3]: Error in NvDsInferContextImpl::initialize() <nvdsinfer_context_impl.cpp:1282> [UID = 3]: generate backend failed, check config file settings
0:00:02.616255703 221656 0x560ff81ece40 WARN nvinfer gstnvinfer.cpp:898:gst_nvinfer_start:<secondary-infer-engine2> error: Failed to create NvDsInferContext instance
0:00:02.616263316 221656 0x560ff81ece40 WARN nvinfer gstnvinfer.cpp:898:gst_nvinfer_start:<secondary-infer-engine2> error: Config file path: /home/mainak/ms/C++/anpr_kp/models/lpr_config_sgie_us.yml, NvDsInfer Error: NVDSINFER_CONFIG_FAILED
Running...
ERROR from element secondary-infer-engine2: Failed to create NvDsInferContext instance
Error details: gstnvinfer.cpp(898): gst_nvinfer_start (): /GstPipeline:ANPR-pipeline/GstNvInfer:secondary-infer-engine2:
Config file path: /home/mainak/ms/C++/anpr_kp/models/lpr_config_sgie_us.yml, NvDsInfer Error: NVDSINFER_CONFIG_FAILED
Returned, stopping playback
Deleting pipeline
Disconnecting MQTT Client
Destroying MQTT Client

```

@Morganh Any suggestion on this is highly appreciated

---

<div class="post-metadata">

**Author:** ![Morganh](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/morganh/32/9748_2.png) [@Morganh](https://forums.developer.nvidia.com/u/Morganh)\
**Post date:** [June 7, 2024, 12:54am UTC](https://forums.developer.nvidia.com/t/lprnet-training-and-deployment/295332/3 "2024-06-07T00:54:50Z")

</div>

> [@mainak1](#):
>
> I have trained a model using tao toolkit and the weight files are as `.hdf5` format. How to convert this to `etlt` so as to deploy in the pipeline?
> 
> I generated the `onnx` file from `hdf5`

Can you follow [https://github.com/NVIDIA/tao\_tutorials/blob/main/notebooks/tao\_launcher\_starter\_kit/lprnet/lprnet.ipynb](https://github.com/NVIDIA/tao_tutorials/blob/main/notebooks/tao_launcher_starter_kit/lprnet/lprnet.ipynb) to generate onnx file? See the “Deploy!” section.

---

<div class="post-metadata">

**Author:** ![mainak1](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@mainak1](https://forums.developer.nvidia.com/u/mainak1)\
**Post date:** [June 7, 2024, 5:05am UTC](https://forums.developer.nvidia.com/t/lprnet-training-and-deployment/295332/4 "2024-06-07T05:05:48Z")

</div>

I followed it to get the `onnx`. However, when deployed then I get the above said error. Also I might retrain hence, I wanted to know how to get `tlt` from `hdf5` and `etlt` from `onnx`?

---

<div class="post-metadata">

**Author:** ![Morganh](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/morganh/32/9748_2.png) [@Morganh](https://forums.developer.nvidia.com/u/Morganh)\
**Post date:** [June 7, 2024, 3:18pm UTC](https://forums.developer.nvidia.com/t/lprnet-training-and-deployment/295332/5 "2024-06-07T15:18:53Z")

</div>

> [@mainak1](#):
>
> `ERROR: [TRT]: ModelImporter.cpp:688: Failed to parse ONNX model from file:`

Can you use Netron to open this onnx file successfully?

---

<div class="post-metadata">

**Author:** ![mainak1](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@mainak1](https://forums.developer.nvidia.com/u/mainak1)\
**Post date:** [June 10, 2024, 8:40am UTC](https://forums.developer.nvidia.com/t/lprnet-training-and-deployment/295332/6 "2024-06-10T08:40:13Z")

</div>

> [@mainak1](#):
>
> While training how to save weights based on best metrics (like accuracy) rather than after every 5 epochs?

ANy help on this?? @Morganh

---

<div class="post-metadata">

**Author:** ![mainak1](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@mainak1](https://forums.developer.nvidia.com/u/mainak1)\
**Post date:** [June 10, 2024, 9:14am UTC](https://forums.developer.nvidia.com/t/lprnet-training-and-deployment/295332/7 "2024-06-10T09:14:21Z")

</div>

> [@Morganh](#):
>
> Can you use Netron to open this onnx file successfully?

I’m able to integrate it.

---

<div class="post-metadata">

**Author:** ![Morganh](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/morganh/32/9748_2.png) [@Morganh](https://forums.developer.nvidia.com/u/Morganh)\
**Post date:** [June 11, 2024, 4:09pm UTC](https://forums.developer.nvidia.com/t/lprnet-training-and-deployment/295332/8 "2024-06-11T16:09:43Z")

</div>

> [@mainak1](#):
>
> While training how to save weights based on best metrics (like accuracy) rather than after every 5 epochs?

You can set it in [https://github.com/NVIDIA/tao\_tutorials/blob/main/notebooks/tao\_launcher\_starter\_kit/lprnet/specs/tutorial\_spec.txt#L25](https://github.com/NVIDIA/tao_tutorials/blob/main/notebooks/tao_launcher_starter_kit/lprnet/specs/tutorial_spec.txt#L25).

---

<div class="post-metadata">

**Author:** ![mainak1](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@mainak1](https://forums.developer.nvidia.com/u/mainak1)\
**Post date:** [June 12, 2024, 5:36am UTC](https://forums.developer.nvidia.com/t/lprnet-training-and-deployment/295332/9 "2024-06-12T05:36:13Z")

</div>

> [@Morganh](#):
>
> You can set it in [tao\_tutorials/notebooks/tao\_launcher\_starter\_kit/lprnet/specs/tutorial\_spec.txt at main · NVIDIA/tao\_tutorials · GitHub](https://github.com/NVIDIA/tao_tutorials/blob/main/notebooks/tao_launcher_starter_kit/lprnet/specs/tutorial_spec.txt#L25).

Ok. But `checkpoint_interval` takes `int` values. How can I save on best metrics like (accuracy)

---

<div class="post-metadata">

**Author:** ![Morganh](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/morganh/32/9748_2.png) [@Morganh](https://forums.developer.nvidia.com/u/Morganh)\
**Post date:** [June 12, 2024, 7:54am UTC](https://forums.developer.nvidia.com/t/lprnet-training-and-deployment/295332/10 "2024-06-12T07:54:44Z")

</div>

> [@mainak1](#):
>
> How can I save on best metrics like (accuracy)

The result folder will save the model which has best accuracy.

---

<div class="post-metadata">

**Author:** ![mainak1](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@mainak1](https://forums.developer.nvidia.com/u/mainak1)\
**Post date:** [June 12, 2024, 8:04am UTC](https://forums.developer.nvidia.com/t/lprnet-training-and-deployment/295332/11 "2024-06-12T08:04:40Z")

</div>

So like if `checkpoint_interval = 5`, then the which checkpoint will be saved? the 5th one for the best among last 5 epochs?

---

<div class="post-metadata">

**Author:** ![Morganh](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/morganh/32/9748_2.png) [@Morganh](https://forums.developer.nvidia.com/u/Morganh)\
**Post date:** [June 12, 2024, 8:08am UTC](https://forums.developer.nvidia.com/t/lprnet-training-and-deployment/295332/12 "2024-06-12T08:08:30Z")

</div>

This `checkpoint_interval` means running evaluation every 5 epochs.  
You can set it to 1. Then run evaluation every 1 epoch.  
Then the best model can be found from all the epochs.

---

<div class="post-metadata">

**Author:** ![mainak1](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@mainak1](https://forums.developer.nvidia.com/u/mainak1)\
**Post date:** [June 12, 2024, 8:29am UTC](https://forums.developer.nvidia.com/t/lprnet-training-and-deployment/295332/13 "2024-06-12T08:29:33Z")

</div>

@Morganh  
One final thing before I close the topic. I’m trying to retrain the model:

`tao model lprnet train --gpus=1 --gpu_index=0 -e /workspace/tao-experiments/lprnet/specs/tutorial_spec.txt -k nvidia_tlt -r /workspace/tao-experiments/lprnet/experiment_dir_unpruned -m /workspace/tao-experiments/lprnet/experiment_dir_unpruned/weights/lprnet_epoch-070.hdf5 --initial_epoch 70`

the following error comes up:

```auto
INFO: Log file already exists at /workspace/tao-experiments/lprnet/experiment_dir_unpruned/status.json
INFO: Merging specification from /workspace/tao-experiments/lprnet/specs/tutorial_spec.txt
INFO: Loading pretrained weights. This may take a while...
INFO: Training was interrupted
INFO: Training was interrupted.
Execution status: PASS
2024-06-12 07:57:59,796 [TAO Toolkit] [INFO] nvidia_tao_cli.components.docker_handler.docker_handler 363: Stopping container.

```

However, the $KEY is set to `nvidia-_tlt`. What key to use in order to retrain `.hdf5` weight? Any suggestion on this? I’m trainig on ec2 instance

---

<div class="post-metadata">

**Author:** ![Morganh](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/morganh/32/9748_2.png) [@Morganh](https://forums.developer.nvidia.com/u/Morganh)\
**Post date:** [June 12, 2024, 8:32am UTC](https://forums.developer.nvidia.com/t/lprnet-training-and-deployment/295332/14 "2024-06-12T08:32:12Z")

</div>

> [@mainak1](#):
>
> However, the $KEY is set to `nvidia-_tlt`. What key to use in order to retrain `.hdf5` weight?

According to [GPU-optimized AI, Machine Learning, & HPC Software | NVIDIA NGC | NVIDIA NGC](https://catalog.ngc.nvidia.com/orgs/nvidia/teams/tao/models/lprnet).  
The key is: nvidia\_tlt

---

<div class="post-metadata">

**Author:** ![mainak1](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@mainak1](https://forums.developer.nvidia.com/u/mainak1)\
**Post date:** [June 12, 2024, 8:33am UTC](https://forums.developer.nvidia.com/t/lprnet-training-and-deployment/295332/15 "2024-06-12T08:33:54Z")

</div>

> [@mainak1](#):
>
> the following error comes up:
> 
> ```auto
> INFO: Log file already exists at /workspace/tao-experiments/lprnet/experiment_dir_unpruned/status.json
> INFO: Merging specification from /workspace/tao-experiments/lprnet/specs/tutorial_spec.txt
> INFO: Loading pretrained weights. This may take a while...
> INFO: Training was interrupted
> INFO: Training was interrupted.
> Execution status: PASS
> 2024-06-12 07:57:59,796 [TAO Toolkit] [INFO] nvidia_tao_cli.components.docker_handler.docker_handler 363: Stopping container.
> 
> ```

@Morganh  
Any suggestion on the above error?

---

<div class="post-metadata">

**Author:** ![Morganh](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/morganh/32/9748_2.png) [@Morganh](https://forums.developer.nvidia.com/u/Morganh)\
**Post date:** [June 12, 2024, 9:12am UTC](https://forums.developer.nvidia.com/t/lprnet-training-and-deployment/295332/16 "2024-06-12T09:12:07Z")

</div>

> [@mainak1](#):
>
> ```auto
> INFO: Training was interrupted
> INFO: Training was interrupted.
> 
> ```

There is not error. The info usually means your key is wrong.  
As mentioned above, please use `nvidia_tlt`.

---

<div class="post-metadata">

**Author:** ![mainak1](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@mainak1](https://forums.developer.nvidia.com/u/mainak1)\
**Post date:** [June 12, 2024, 9:17am UTC](https://forums.developer.nvidia.com/t/lprnet-training-and-deployment/295332/17 "2024-06-12T09:17:20Z")

</div>

> [@Morganh](#):
>
> As mentioned above, please use `nvidia_tlt`.

I’ve set it:

```auto
tao model lprnet train --gpus=1 --gpu_index=0 -e /workspace/tao-experiments/lprnet/specs/tutorial_spec.txt -k nvidia_tlt -r /workspace/tao-experiments/lprnet/experiment_dir_unpruned -m /workspace/tao-experiments/lprnet/experiment_dir_unpruned/weights/lprnet_epoch-070.hdf5 --initial_epoch 70

```

Still the error persists! I’m retraining against the `hdf5` file not `tlt`. Will the key change in that case?

---

<div class="post-metadata">

**Author:** ![Morganh](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/morganh/32/9748_2.png) [@Morganh](https://forums.developer.nvidia.com/u/Morganh)\
**Post date:** [June 12, 2024, 9:21am UTC](https://forums.developer.nvidia.com/t/lprnet-training-and-deployment/295332/18 "2024-06-12T09:21:20Z")

</div>

> [@mainak1](#):
>
> `-m /workspace/tao-experiments/lprnet/experiment_dir_unpruned/weights/lprnet_epoch-070.hdf5 `

How did you train and get `experiment_dir_unpruned/weights/lprnet_epoch-070.hdf5`? Is there any key when you run previous training?

---

<div class="post-metadata">

**Author:** ![mainak1](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@mainak1](https://forums.developer.nvidia.com/u/mainak1)\
**Post date:** [June 12, 2024, 9:22am UTC](https://forums.developer.nvidia.com/t/lprnet-training-and-deployment/295332/19 "2024-06-12T09:22:55Z")

</div>

Yes the for training I used the below command:

```auto
tao model lprnet train --gpus=1 --gpu_index=0 -e /workspace/tao-experiments/lprnet/specs/tutorial_spec.txt -k nvidia_tlt -r /workspace/tao-experiments/lprnet/experiment_dir_unpruned -m /workspace/tao-experiments/lprnet/pretrained_lprnet_baseline18/lprnet_vtrainable_v1.0/us_lprnet_baseline18_trainable.tlt

```

where `-k nvidia_tlt`

---

<div class="post-metadata">

**Author:** ![Morganh](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/morganh/32/9748_2.png) [@Morganh](https://forums.developer.nvidia.com/u/Morganh)\
**Post date:** [June 12, 2024, 9:23am UTC](https://forums.developer.nvidia.com/t/lprnet-training-and-deployment/295332/20 "2024-06-12T09:23:51Z")

</div>

> [@mainak1](#):
>
> -m /workspace/tao-experiments/lprnet/experiment\_dir\_unpruned/weights/lprnet\_epoch-070.hdf5

Please try not set `-k` since you are retraining with new model.

---

<div class="post-metadata">

**Author:** ![mainak1](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@mainak1](https://forums.developer.nvidia.com/u/mainak1)\
**Post date:** [June 12, 2024, 9:26am UTC](https://forums.developer.nvidia.com/t/lprnet-training-and-deployment/295332/21 "2024-06-12T09:26:37Z")

</div>

Now, I’m using:

```auto
tao model lprnet train --gpus=1 --gpu_index=0 -e /workspace/tao-experiments/lprnet/specs/tutorial_spec.txt -r /workspace/tao-experiments/lprnet/experiment_dir_unpruned -m /workspace/tao-experiments/lprnet/experiment_dir_unpruned/weights/lprnet_epoch-070.hdf5 --initial_epoch 70

```

still the error persists!!

[Next page](https://forums.developer.nvidia.com/t/lprnet-training-and-deployment/295332.md?page=2)
