# Error in TAO-Toolkit while training

**URL:** <https://forums.developer.nvidia.com/t/error-in-tao-toolkit-while-training/215066>\
**Category:** TAO Toolkit\
**Created:** [May 20, 2022, 12:14pm UTC](https://forums.developer.nvidia.com/t/error-in-tao-toolkit-while-training/215066 "2022-05-20T12:14:23Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![2656780992](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@2656780992](https://forums.developer.nvidia.com/u/2656780992)\
**Post date:** [May 20, 2022, 12:14pm UTC](https://forums.developer.nvidia.com/t/error-in-tao-toolkit-while-training/215066/1 "2022-05-20T12:14:24Z")

</div>

Please provide the following information when requesting support.

• Hardware (T4/V100/Xavier/Nano/etc) 3090  
• Network Type ActionRecognitionNet  
• TLT Version [nvcr.io/nvidia/tlt-streamanalytics:v3.0-py3](http://nvcr.io/nvidia/tlt-streamanalytics:v3.0-py3)  
• Training spec file(If have, please share here)  
[train\_rgb\_3d\_finetune.yaml](https://forums.developer.nvidia.com/uploads/short-url/ncgUn640FezBL6hzMPIebwQcuLC.yaml) (761 Bytes)

• How to reproduce the issue ? (This is for errors. Please share the command line and the detailed log here.)I am trying to train action recognition net inside TAO toolkit container by following the nvidia blog for ActionRecognitionNet.

I have started the container using the following command in my personal machine:

docker run --name fan-tlt --runtime=nvidia -it -v /var/run/docker.sock:/var/run/docker.sock -v /media/userdata/fanyl/tlt/:/home -p 8888:8888 -w /home [nvcr.io/nvidia/tlt-streamanalytics:v3.0-py3](http://nvcr.io/nvidia/tlt-streamanalytics:v3.0-py3) /bin/bash

Inside this, I was able to successfully follow the jupyter notebook as mentioned in the blog up till the training part. When i run the following command  
print(“Train RGB only model with PTM”)  
!tao action\_recognition train   
-e $SPECS\_DIR/train\_rgb\_3d\_finetune.yaml   
-r $RESULTS\_DIR/rgb\_3d\_ptm   
-k $KEY   
model\_config.rgb\_pretrained\_model\_path=$RESULTS\_DIR/pretrained/actionrecognitionnet\_vtrainable\_v1.0/resnet18\_3d\_rgb\_hmdb5\_32.tlt   
#ognition train   
model\_config.rgb\_pretrained\_num\_classes=5  
print(“Train RGB only model with PTM”)

!tao action\_recognition train \

```
              -e $SPECS_DIR/train_rgb_3d_finetune.yaml \

              -r $RESULTS_DIR/rgb_3d_ptm \

              -k $KEY \

              model_config.rgb_pretrained_model_path=$RESULTS_DIR/pretrained/actionrecognitionnet_vtrainable_v1.0/resnet18_3d_rgb_hmdb5_32.tlt \

              #ognition train \

              model_config.rgb_pretrained_num_classes=5

```

I am getting error:

Train RGB only model with PTM  
2022-05-20 11:58:15,350 [INFO] root: Registry: [‘[nvcr.io](http://nvcr.io)’]  
2022-05-20 11:58:15,418 [INFO] tlt.components.instance\_handler.local\_instance: Running command in container: [nvcr.io/nvidia/tao/tao-toolkit-pyt:v3.21.11-py3](http://nvcr.io/nvidia/tao/tao-toolkit-pyt:v3.21.11-py3)  
2022-05-20 11:58:15,541 [WARNING] tlt.components.docker\_handler.docker\_handler:  
Docker will run the commands as root. If you would like to retain your  
local host permissions, please add the “user”:“UID:GID” in the  
DockerOptions portion of the “/root/.tao\_mounts.json” file. You can obtain your  
users UID and GID by using the “id -u” and “id -g” commands on the  
terminal.  
ERROR: The indicated experiment spec file `/home/tlt-experiments/action_recognition_net/host/specs/train_rgb_3d_finetune.yaml` doesn’t exist!  
2022-05-20 11:58:18,148 [INFO] tlt.components.docker\_handler.docker\_handler: Stopping container.

I can make sure that the file path exists. Why is this mistake?

---

<div class="post-metadata">

**Author:** ![Morganh](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/morganh/32/9748_2.png) [@Morganh](https://forums.developer.nvidia.com/u/Morganh)\
**Post date:** [May 20, 2022, 4:22pm UTC](https://forums.developer.nvidia.com/t/error-in-tao-toolkit-while-training/215066/2 "2022-05-20T16:22:01Z")

</div>

> [@2656780992](#):
>
> ERROR: The indicated experiment spec file `/home/tlt-experiments/action_recognition_net/host/specs/train_rgb_3d_finetune.yaml` doesn’t exist!

Since you are running command “!tao action\_recognition train xxx” in notebook, that means tao-launcher is used. It is necessary to set correct ~/.tao\_mounts.json to map the files.

I find that you are triggering tao docker as below  
`docker run --name fan-tlt --runtime=nvidia -it -v /var/run/docker.sock:/var/run/docker.sock -v /media/userdata/fanyl/tlt/:/home -p 8888:8888 -w /home [nvcr.io/nvidia/tlt-streamanalytics:v3.0-py3](http://nvcr.io/nvidia/tlt-streamanalytics:v3.0-py3) /bin/bash`

You can directly run training without tao-launcher and jupyter notebook, i.e.,  
`#` action\_recognition train xxx

---

<div class="post-metadata">

**Author:** ![2656780992](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@2656780992](https://forums.developer.nvidia.com/u/2656780992)\
**Post date:** [May 21, 2022, 1:09am UTC](https://forums.developer.nvidia.com/t/error-in-tao-toolkit-while-training/215066/3 "2022-05-21T01:09:48Z")

</div>

Hello, after reading your reply, I have two questions for you to reply to.

The first question is: how to set the correct ~ / tao\_ mounts. JSON file. I added specs to the example\_ Dir path, but the error still exists. My ~ / tao\_ mounts. json is as follows:

# Mapping up the local directories to the TAO docker.

import json  
import os  
mounts\_ file = os. path. expanduser(“~/.tao\_mounts.json”)  
tlt\_ configs = {  
“Mounts”:[  
{  
“source”: os. environ[“HOST\_DATA\_DIR”],  
“destination”: “/data”  
},  
{  
“source”: os. environ[“HOST\_SPECS\_DIR”],  
“destination”: “/specs”  
},  
{  
“source”: os. environ[“HOST\_RESULTS\_DIR”],  
“destination”: “/results”  
},  
{  
“source”: os. path. expanduser(“~/.cache”),  
“destination”: “/root/.cache”  
},  
{  
“source”: os. environ[“SPECS\_DIR”],  
“destination”: “/specs”  
}  
],  
“DockerOptions”: {  
“shm\_size”: “16G”,  
“ulimits”: {  
“memlock”: -1,  
“stack”: 67108864  
}  
}  
}  
with open(mounts\_file, “w”) as mfile:  
json. dump(tlt\_configs, mfile, indent=4)

# The second question is: how to use action\_ recognition train xxx

When I run directly, the error is reported as follows:

root@c67f146de16e :/home/tlt-experiments/action\_ recognition\_ net# action\_ recognition train -e $SPECS\_ DIR/train\_ rgb\_ 3d\_ finetune. yaml -r $RESULTS\_ DIR/rgb\_ 3d\_ ptm -k $KEY model\_ config. rgb\_ pretrained\_ model\_ path=$RESULTS\_ DIR/pretrained/actionrecognitionnet\_ vtrainable\_ v1. 0/resnet18\_ 3d\_ rgb\_ hmdb5\_ 32.tlt model\_ config. rgb\_ pretrained\_ num\_ classes=5

bash: action\_ recognition: command not found

---

<div class="post-metadata">

**Author:** ![2656780992](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@2656780992](https://forums.developer.nvidia.com/u/2656780992)\
**Post date:** [May 21, 2022, 1:35am UTC](https://forums.developer.nvidia.com/t/error-in-tao-toolkit-while-training/215066/4 "2022-05-21T01:35:23Z")

</div>

This is a supplement to the second question:

root@c67f146de16e :/home/tlt-experiments/action\_ recognition\_ net# whereis action\_ recognition  
action\_ recognition:

How to add an action\_ Recognition command?

---

<div class="post-metadata">

**Author:** ![Morganh](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/morganh/32/9748_2.png) [@Morganh](https://forums.developer.nvidia.com/u/Morganh)\
**Post date:** [May 21, 2022, 2:23am UTC](https://forums.developer.nvidia.com/t/error-in-tao-toolkit-while-training/215066/5 "2022-05-21T02:23:39Z")

</div>

> [@2656780992](#):
>
> root@c67f146de16e :/home/tlt-experiments/action\_ recognition\_ net# whereis action\_ recognition  
> action\_ recognition:
> 
> How to add an action\_ Recognition command?

See below info.

```auto
$ tao info --verbose
Configuration of the TAO Toolkit Instance

dockers:
        nvidia/tao/tao-toolkit-tf:
                v3.21.11-tf1.15.5-py3:
                        docker_registry: nvcr.io
                        tasks:
                                1. augment
                                2. bpnet
                                3. classification
                                4. dssd
                                5. emotionnet
                                6. efficientdet
                                7. fpenet
                                8. gazenet
                                9. gesturenet
                                10. heartratenet
                                11. lprnet
                                12. mask_rcnn
                                13. multitask_classification
                                14. retinanet
                                15. ssd
                                16. unet
                                17. yolo_v3
                                18. yolo_v4
                                19. yolo_v4_tiny
                                20. converter
                v3.21.11-tf1.15.4-py3:
                        docker_registry: nvcr.io
                        tasks:
                                1. detectnet_v2
                                2. faster_rcnn
        nvidia/tao/tao-toolkit-pyt:
                v3.21.11-py3:
                        docker_registry: nvcr.io
                        tasks:
                                1. speech_to_text
                                2. speech_to_text_citrinet
                                3. text_classification
                                4. question_answering
                                5. token_classification
                                6. intent_slot_classification
                                7. punctuation_and_capitalization
                                8. spectro_gen
                                9. vocoder
                                10. action_recognition
        nvidia/tao/tao-toolkit-lm:
                v3.21.08-py3:
                        docker_registry: nvcr.io
                        tasks:
                                1. n_gram
format_version: 2.0
toolkit_version: 3.21.11
published_date: 11/08/2021

```

The action\_recognition network is in [nvcr.io/nvidia/tao/tao-toolkit-pyt:v3.21.11-py3](http://nvcr.io/nvidia/tao/tao-toolkit-pyt:v3.21.11-py3)

So, you need to trigger docker [nvcr.io/nvidia/tao/tao-toolkit-pyt:v3.21.11-py3](http://nvcr.io/nvidia/tao/tao-toolkit-pyt:v3.21.11-py3).

That’s why officially we mentioned in the user guide that it is recommended to use tao-launcher to trigger tao docker instead of "docker run xxx ".

For ~/.tao\_mounts.json, please refer to [TAO Launcher — Tao Toolkit](https://docs.nvidia.com/tao/tao-toolkit/text/tao_launcher.html#running-the-launcher)

---

<div class="post-metadata">

**Author:** ![2656780992](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@2656780992](https://forums.developer.nvidia.com/u/2656780992)\
**Post date:** [May 21, 2022, 1:01pm UTC](https://forums.developer.nvidia.com/t/error-in-tao-toolkit-while-training/215066/6 "2022-05-21T13:01:59Z")

</div>

> [@2656780992](#):
>
> mounts\_ file = os. path. expanduser(“~/.tao\_mounts.json”)

Hello,

For ~/. tao\_ mounts. json, I refer to Tao toolkit launcher - Tao toolkit 3.22.02 documentation. However, I found that the local machine path cannot be found in the docker container, and it can be used when I use the path in the container as the local machine path. I want to know how to configure the address.

The second problem is the error when running outside the container:

root@ubuntu -MS-7B94:/media/userdata/fanyl/tlt/tlt-experiments/action\_ recognition\_ net1# tao action\_ recognition train -e /media/userdata/fanyl/tlt/tlt-experiments/action\_ recognition\_ net1/specs/train\_ rgb\_ 3d\_ finetune. yaml -r /media/userdata/fanyl/tlt/tlt-experiments/action\_ recognition\_ net1/results/rgb\_ 3d\_ ptm -k Zm9xNHQ0ajE1YjI5aGJiNzU4OTZtcDhxdDY6YjhhMTE4NGEtYWJmNi00MGU0LWIxNjAtNmYyNjg2N2JlYjUy model\_ config. rgb\_ pretrained\_ model\_ path=/media/userdata/fanyl/tlt/tlt-experiments/action\_ recognition\_ net1/results/pretrained/actionrecognitionnet\_ vtrainable\_ v1. 0/resnet18\_ 3d\_ rgb\_ hmdb5\_ 32.tlt model\_ config. rgb\_ pretrained\_ num\_ classes=5

~/. tao\_ mounts. json wasn’t found. Falling back to obtain mount points and docker configs from ~/. tlt\_ mounts. json.

Please note that this will be deprecated going forward.

2022-05-21 20:56:42,644 [INFO] root: Registry: [‘[nvcr.io](http://nvcr.io)’]

2022-05-21 20:56:42,704 [INFO] tlt. components. instance\_ handler. local\_ instance: Running command in container: nvcr. io/nvidia/tao/tao-toolkit-pyt:v3. 21.11-py3

2022-05-21 20:56:42,825 [INFO] root: No mount points were found in the /root/. tlt\_ mounts. json file.

2022-05-21 20:56:42,825 [WARNING] tlt. components. docker\_ handler. docker\_ handler:

Docker will run the commands as root. If you would like to retain your

local host permissions, please add the “user”:“UID:GID” in the

DockerOptions portion of the “/root/.tlt\_mounts.json” file. You can obtain your

users UID and GID by using the “id -u” and “id -g” commands on the

terminal.

ERROR: The indicated experiment spec file `/media/userdata/fanyl/tlt/tlt-experiments/action_ recognition_ net1/specs/train_ rgb_ 3d_ finetune. yaml` doesn’t exist!

I want to know how to configure it outside the container tao\_ mounts. json。

---

<div class="post-metadata">

**Author:** ![Morganh](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/morganh/32/9748_2.png) [@Morganh](https://forums.developer.nvidia.com/u/Morganh)\
**Post date:** [May 22, 2022, 1:55am UTC](https://forums.developer.nvidia.com/t/error-in-tao-toolkit-while-training/215066/7 "2022-05-22T01:55:55Z")

</div>

You can create ~/.tao\_mounts.json.  
Then follow tao user guide to setup correct mapping.

---

<div class="post-metadata">

**Author:** ![2656780992](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@2656780992](https://forums.developer.nvidia.com/u/2656780992)\
**Post date:** [May 22, 2022, 10:47am UTC](https://forums.developer.nvidia.com/t/error-in-tao-toolkit-while-training/215066/8 "2022-05-22T10:47:12Z")

</div>

According to your suggestion, I have solved the above problem. However, an error is still reported when running again. The details are as follows:

> ubuntu@ubuntu-MS-7B94:/media/userdata/fanyl/cv\_samples\_vv1.3.0/action\_recognition\_net1$ tao action\_recognition train -e /specs/train\_rgb\_3d\_finetune.yaml -r /results/rgb\_3d\_ptm -k Zm9xNHQ0ajE1YjI5aGJiNzU4OTZtcDhxdDY6YjhhMTE4NGEtYWJmNi00MGU0LWIxNjAtNmYyNjg2N2JlYjUy model\_config.rgb\_pretrained\_model\_path=/results/pretrained/actionrecognitionnet\_vtrainable\_v1.0/resnet18\_3d\_rgb\_hmdb5\_32.tlt model\_config.rgb\_pretrained\_num\_classes=5  
> 2022-05-22 18:42:36,345 [INFO] root: Registry: [‘[nvcr.io](http://nvcr.io)’]  
> 2022-05-22 18:42:36,409 [INFO] tlt.components.instance\_handler.local\_instance: Running command in container: [nvcr.io/nvidia/tao/tao-toolkit-pyt:v3.21.11-py3](http://nvcr.io/nvidia/tao/tao-toolkit-pyt:v3.21.11-py3)  
> 2022-05-22 18:42:36,524 [WARNING] tlt.components.docker\_handler.docker\_handler:  
> Docker will run the commands as root. If you would like to retain your  
> local host permissions, please add the “user”:“UID:GID” in the  
> DockerOptions portion of the “/home/ubuntu/.tao\_mounts.json” file. You can obtain your  
> users UID and GID by using the “id -u” and “id -g” commands on the  
> terminal.  
> /home/jenkins/agent/workspace/tlt-pytorch-main-nightly/cv/action\_recognition/scripts/train.py:76: UserWarning:  
> ‘train\_rgb\_3d\_finetune.yaml’ is validated against ConfigStore schema with the same name.  
> This behavior is deprecated in Hydra 1.1 and will be removed in Hydra 1.2.  
> See [https://hydra.cc/docs/next/upgrades/1.0\_to\_1.1/automatic\_schema\_matching](https://hydra.cc/docs/next/upgrades/1.0_to_1.1/automatic_schema_matching) for migration instructions.  
> loading trained weights from /results/pretrained/actionrecognitionnet\_vtrainable\_v1.0/resnet18\_3d\_rgb\_hmdb5\_32.tlt  
> Error executing job with overrides: [‘output\_dir=/results/rgb\_3d\_ptm’, ‘encryption\_key=Zm9xNHQ0ajE1YjI5aGJiNzU4OTZtcDhxdDY6YjhhMTE4NGEtYWJmNi00MGU0LWIxNjAtNmYyNjg2N2JlYjUy’, ‘model\_config.rgb\_pretrained\_model\_path=/results/pretrained/actionrecognitionnet\_vtrainable\_v1.0/resnet18\_3d\_rgb\_hmdb5\_32.tlt’, ‘model\_config.rgb\_pretrained\_num\_classes=5’]  
> Traceback (most recent call last):  
> File “/opt/conda/lib/python3.8/site-packages/hydra/\_internal/utils.py”, line 211, in run\_and\_report  
> return func()  
> File “/opt/conda/lib/python3.8/site-packages/hydra/\_internal/utils.py”, line 368, in   
> lambda: hydra.run(  
> File “/opt/conda/lib/python3.8/site-packages/hydra/\_internal/hydra.py”, line 110, in run  
> \_ = ret.return\_value  
> File “/opt/conda/lib/python3.8/site-packages/hydra/core/utils.py”, line 233, in return\_value  
> raise self.\_return\_value  
> File “/opt/conda/lib/python3.8/site-packages/hydra/core/utils.py”, line 160, in run\_job  
> ret.return\_value = task\_function(task\_cfg)  
> File “/home/jenkins/agent/workspace/tlt-pytorch-main-nightly/cv/action\_recognition/scripts/train.py”, line 70, in main  
> File “/home/jenkins/agent/workspace/tlt-pytorch-main-nightly/cv/action\_recognition/scripts/train.py”, line 22, in run\_experiment  
> File “/home/jenkins/agent/workspace/tlt-pytorch-main-nightly/cv/action\_recognition/model/pl\_ar\_model.py”, line 29, in **init**  
> File “/home/jenkins/agent/workspace/tlt-pytorch-main-nightly/cv/action\_recognition/model/pl\_ar\_model.py”, line 36, in \_build\_model  
> File “/home/jenkins/agent/workspace/tlt-pytorch-main-nightly/cv/action\_recognition/model/build\_nn\_model.py”, line 76, in build\_ar\_model  
> File “/home/jenkins/agent/workspace/tlt-pytorch-main-nightly/cv/action\_recognition/model/ar\_model.py”, line 88, in get\_basemodel3d  
> File “/home/jenkins/agent/workspace/tlt-pytorch-main-nightly/cv/action\_recognition/model/ar\_model.py”, line 23, in load\_pretrained\_weights  
> File “/home/jenkins/agent/workspace/tlt-pytorch-main-nightly/cv/action\_recognition/utils/common\_utils.py”, line 22, in patch\_decrypt\_checkpoint  
> File “/home/jenkins/agent/workspace/tlt-pytorch-main-nightly/tlt\_utils/checkpoint\_encryption.py”, line 26, in decrypt\_checkpoint  
> \_pickle.UnpicklingError: invalid load key, ‘\xf6’.  
> During handling of the above exception, another exception occurred:  
> Traceback (most recent call last):  
> File “/home/jenkins/agent/workspace/tlt-pytorch-main-nightly/cv/action\_recognition/scripts/train.py”, line 76, in   
> File “/home/jenkins/agent/workspace/tlt-pytorch-main-nightly/cv/super\_resolution/scripts/configs/hydra\_runner.py”, line 99, in wrapper  
> File “/opt/conda/lib/python3.8/site-packages/hydra/\_internal/utils.py”, line 367, in \_run\_hydra  
> run\_and\_report(  
> File “/opt/conda/lib/python3.8/site-packages/hydra/\_internal/utils.py”, line 251, in run\_and\_report  
> assert mdl is not None  
> AssertionError  
> 2022-05-22 18:42:42,903 [INFO] tlt.components.docker\_handler.docker\_handler: Stopping container.  
> ubuntu@ubuntu-MS-7B94:/media/userdata/fanyl/cv\_samples\_vv1.3.0/action\_recognition\_net1$

Please tell me what I should do?

---

<div class="post-metadata">

**Author:** ![Morganh](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/morganh/32/9748_2.png) [@Morganh](https://forums.developer.nvidia.com/u/Morganh)\
**Post date:** [May 23, 2022, 8:12am UTC](https://forums.developer.nvidia.com/t/error-in-tao-toolkit-while-training/215066/9 "2022-05-23T08:12:05Z")

</div>

This kind of error is usually due to wrong mapping setting.

Please check the ~/.tao\_mounts.json.

Please note that all the path in the commandline should be the path inside the docker.  
The path is defined in ~/.tao\_mounts.json.

Or you directly login the docker and run tasks.  
$ tao action\_recognition  
then,  
`#` action\_recognition train xxx

---

<div class="post-metadata">

**Author:** ![2656780992](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@2656780992](https://forums.developer.nvidia.com/u/2656780992)\
**Post date:** [May 25, 2022, 9:36am UTC](https://forums.developer.nvidia.com/t/error-in-tao-toolkit-while-training/215066/10 "2022-05-25T09:36:58Z")

</div>

After reconfiguring the mapping file, I rerun and the following error is reported:

ubuntu@ubuntu-MS-7B94:~$ tao action\_recognition train -e /workspace/tlt-experiments/like/specs/train\_rgb\_3d\_finetune.yaml -r /workspace/tlt-experiments/results/rgb\_3d\_ptm -k $KEY model\_config.rgb\_pretrained\_model\_path=/workspace/tlt-experiments/results/pretrained/actionrecognitionnet\_vtrainable\_v1.0/resnet18\_3d\_rgb\_hmdb5\_32.tlt model\_config.rgb\_pretrained\_num\_classes=5  
2022-05-25 17:31:18,152 [INFO] root: Registry: [‘[nvcr.io](http://nvcr.io)’]  
2022-05-25 17:31:18,226 [INFO] tlt.components.instance\_handler.local\_instance: Running command in container: [nvcr.io/nvidia/tao/tao-toolkit-pyt:v3.21.11-py3](http://nvcr.io/nvidia/tao/tao-toolkit-pyt:v3.21.11-py3)  
2022-05-25 17:31:18,345 [WARNING] tlt.components.docker\_handler.docker\_handler:  
Docker will run the commands as root. If you would like to retain your  
local host permissions, please add the “user”:“UID:GID” in the  
DockerOptions portion of the “/home/ubuntu/.tao\_mounts.json” file. You can obtain your  
users UID and GID by using the “id -u” and “id -g” commands on the  
terminal.  
mismatched input ‘=’ expecting   
See [https://hydra.cc/docs/next/advanced/override\_grammar/basic](https://hydra.cc/docs/next/advanced/override_grammar/basic) for details

Set the environment variable HYDRA\_FULL\_ERROR=1 for a complete stack trace.  
2022-05-25 17:31:23,170 [INFO] tlt.components.docker\_handler.docker\_handler: Stopping container.

What is the reason?  
Here is my Tao information：

tao info  
Configuration of the TAO Toolkit Instance  
dockers: [‘nvidia/tao/tao-toolkit-tf’, ‘nvidia/tao/tao-toolkit-pyt’, ‘nvidia/tao/tao-toolkit-lm’]  
format\_version: 2.0  
toolkit\_version: 3.22.02  
published\_date: 02/28/2022

---

<div class="post-metadata">

**Author:** ![2656780992](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@2656780992](https://forums.developer.nvidia.com/u/2656780992)\
**Post date:** [May 28, 2022, 2:20pm UTC](https://forums.developer.nvidia.com/t/error-in-tao-toolkit-while-training/215066/11 "2022-05-28T14:20:31Z")

</div>

I tried again and still reported an error. If there are relevant examples, please send them to me.

> print(“Train RGB only model with PTM”)  
> !tao action\_recognition train   
> -e $SPECS\_DIR/train\_rgb\_3d\_finetune.yaml   
> -r $RESULTS\_DIR/rgb\_3d\_ptm   
> -k $KEY   
> model\_config.rgb\_pretrained\_model\_path=$RESULTS\_DIR/pretrained/actionrecognitionnet\_vtrainable\_v1.0/resnet18\_3d\_rgb\_hmdb5\_32.tlt   
> model\_config.rgb\_pretrained\_num\_classes=5

> Train RGB only model with PTM  
> 2022-05-28 21:46:46,098 [INFO] root: Registry: [‘[nvcr.io](http://nvcr.io)’]  
> 2022-05-28 21:46:46,174 [INFO] tlt.components.instance\_handler.local\_instance: Running command in container: [nvcr.io/nvidia/tao/tao-toolkit-pyt:v3.21.11-py3](http://nvcr.io/nvidia/tao/tao-toolkit-pyt:v3.21.11-py3)  
> 2022-05-28 21:46:46,214 [WARNING] tlt.components.docker\_handler.docker\_handler:  
> Docker will run the commands as root. If you would like to retain your  
> local host permissions, please add the “user”:“UID:GID” in the  
> DockerOptions portion of the “/home/ubuntu/.tao\_mounts.json” file. You can obtain your  
> users UID and GID by using the “id -u” and “id -g” commands on the  
> terminal.  
> /home/jenkins/agent/workspace/tlt-pytorch-main-nightly/cv/action\_recognition/scripts/train.py:76: UserWarning:  
> ‘train\_rgb\_3d\_finetune.yaml’ is validated against ConfigStore schema with the same name.  
> This behavior is deprecated in Hydra 1.1 and will be removed in Hydra 1.2.  
> See [https://hydra.cc/docs/next/upgrades/1.0\_to\_1.1/automatic\_schema\_matching](https://hydra.cc/docs/next/upgrades/1.0_to_1.1/automatic_schema_matching) for migration instructions.  
> loading trained weights from /home/action\_recognition\_net1/results/pretrained/actionrecognitionnet\_vtrainable\_v1.0/resnet18\_3d\_rgb\_hmdb5\_32.tlt  
> Error executing job with overrides: [‘output\_dir=/home/action\_recognition\_net1/results/rgb\_3d\_ptm’, ‘encryption\_key=Zm9xNHQ0ajE1YjI5aGJiNzU4OTZtcDhxdDY6YjhhMTE4NGEtYWJmNi00MGU0LWIxNjAtNmYyNjg2N2JlYjUy’, ‘model\_config.rgb\_pretrained\_model\_path=/home/action\_recognition\_net1/results/pretrained/actionrecognitionnet\_vtrainable\_v1.0/resnet18\_3d\_rgb\_hmdb5\_32.tlt’, ‘model\_config.rgb\_pretrained\_num\_classes=5’]  
> Traceback (most recent call last):  
> File “/opt/conda/lib/python3.8/site-packages/hydra/\_internal/utils.py”, line 211, in run\_and\_report  
> return func()  
> File “/opt/conda/lib/python3.8/site-packages/hydra/\_internal/utils.py”, line 368, in   
> lambda: hydra.run(  
> File “/opt/conda/lib/python3.8/site-packages/hydra/\_internal/hydra.py”, line 110, in run  
> \_ = ret.return\_value  
> File “/opt/conda/lib/python3.8/site-packages/hydra/core/utils.py”, line 233, in return\_value  
> raise self.\_return\_value  
> File “/opt/conda/lib/python3.8/site-packages/hydra/core/utils.py”, line 160, in run\_job  
> ret.return\_value = task\_function(task\_cfg)  
> File “/home/jenkins/agent/workspace/tlt-pytorch-main-nightly/cv/action\_recognition/scripts/train.py”, line 70, in main  
> File “/home/jenkins/agent/workspace/tlt-pytorch-main-nightly/cv/action\_recognition/scripts/train.py”, line 22, in run\_experiment  
> File “/home/jenkins/agent/workspace/tlt-pytorch-main-nightly/cv/action\_recognition/model/pl\_ar\_model.py”, line 29, in **init**  
> File “/home/jenkins/agent/workspace/tlt-pytorch-main-nightly/cv/action\_recognition/model/pl\_ar\_model.py”, line 36, in \_build\_model  
> File “/home/jenkins/agent/workspace/tlt-pytorch-main-nightly/cv/action\_recognition/model/build\_nn\_model.py”, line 76, in build\_ar\_model  
> File “/home/jenkins/agent/workspace/tlt-pytorch-main-nightly/cv/action\_recognition/model/ar\_model.py”, line 88, in get\_basemodel3d  
> File “/home/jenkins/agent/workspace/tlt-pytorch-main-nightly/cv/action\_recognition/model/ar\_model.py”, line 23, in load\_pretrained\_weights  
> File “/home/jenkins/agent/workspace/tlt-pytorch-main-nightly/cv/action\_recognition/utils/common\_utils.py”, line 22, in patch\_decrypt\_checkpoint  
> File “/home/jenkins/agent/workspace/tlt-pytorch-main-nightly/tlt\_utils/checkpoint\_encryption.py”, line 26, in decrypt\_checkpoint  
> \_pickle.UnpicklingError: invalid load key, ‘\xf6’.

During handling of the above exception, another exception occurred:

Traceback (most recent call last):  
File “/home/jenkins/agent/workspace/tlt-pytorch-main-nightly/cv/action\_recognition/scripts/train.py”, line 76, in   
File “/home/jenkins/agent/workspace/tlt-pytorch-main-nightly/cv/super\_resolution/scripts/configs/hydra\_runner.py”, line 99, in wrapper  
File “/opt/conda/lib/python3.8/site-packages/hydra/\_internal/utils.py”, line 367, in \_run\_hydra  
run\_and\_report(  
File “/opt/conda/lib/python3.8/site-packages/hydra/\_internal/utils.py”, line 251, in run\_and\_report  
assert mdl is not None  
AssertionError  
2022-05-28 21:46:57,602 [INFO] tlt.components.docker\_handler.docker\_handler: Stopping container.

---

<div class="post-metadata">

**Author:** ![2656780992](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@2656780992](https://forums.developer.nvidia.com/u/2656780992)\
**Post date:** [May 29, 2022, 2:57am UTC](https://forums.developer.nvidia.com/t/error-in-tao-toolkit-while-training/215066/12 "2022-05-29T02:57:07Z")

</div>

I directly log in to docker and run the task。  
Executed a command：

> ubuntu@ubuntu-MS-7B94:~$ tao action\_recognition

> root@ab0479f7c057:/workspace/tlt/samples# action\_recognition train -e /home/action\_recognition\_net1/specs/train\_rgb\_3d\_finetune.yaml -r /home/action\_recognition\_net1/results/rgb\_3d\_ptm -k Zm9xNHQ0ajE1YjI5aGJiNzU4OTZtcDhxdDY6YjhhMTE4NGEtYWJmNi00MGU0LWIxNjAtNmYyNjg2N2JlYjUymodel\_config.rgb\_pretrained\_model\_path=/home/action\_recognition\_net1/results/pretrained/actionrecognitionnet\_vtrainable\_v1.0/resnet18\_3d\_rgb\_hmdb5\_32.tlt model\_config.rgb\_pretrained\_num\_classes=5

Here I encountered a new error, which is roughly as follows:

> mismatched input ‘=’ expecting   
> See [https://hydra.cc/docs/next/advanced/override\_grammar/basic](https://hydra.cc/docs/next/advanced/override_grammar/basic) for details

Set the environment variable HYDRA\_FULL\_ERROR=1 for a complete stack trace.

I didn’t find ~/.tao\_mounts.json in the container. Is it a mapping problem?

---

<div class="post-metadata">

**Author:** ![Morganh](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/morganh/32/9748_2.png) [@Morganh](https://forums.developer.nvidia.com/u/Morganh)\
**Post date:** [May 30, 2022, 3:14pm UTC](https://forums.developer.nvidia.com/t/error-in-tao-toolkit-while-training/215066/13 "2022-05-30T15:14:01Z")

</div>

> [@2656780992](#):
>
> I didn’t find ~/.tao\_mounts.json in the container. Is it a mapping problem?

I am afraid yes. Please try to run with terminal instead of notebook.  
And for debugging, you can login the docker run command.  
$ tao action\_recognition  
then inside the docker  
`#` action\_recognition train xxx

---

<div class="post-metadata">

**Author:** ![Morganh](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/morganh/32/9748_2.png) [@Morganh](https://forums.developer.nvidia.com/u/Morganh)\
**Post date:** [May 31, 2022, 1:25am UTC](https://forums.developer.nvidia.com/t/error-in-tao-toolkit-while-training/215066/14 "2022-05-31T01:25:08Z")

</div>

> [@2656780992](#):
>
> I didn’t find ~/.tao\_mounts.json in the container. Is it a mapping problem?

Firstly, please create ~/.tao\_mounts.json and set it correctly.  
Then please try to run with terminal instead of notebook.  
And for debugging, you can login the docker run command.  
$ tao action\_recognition  
then inside the docker  
`#` action\_recognition train xxx

---

<div class="post-metadata">

**Author:** ![yingliu](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/yingliu/32/136703_2.png) [@yingliu](https://forums.developer.nvidia.com/u/yingliu)\
**Post date:** [July 6, 2022, 6:34am UTC](https://forums.developer.nvidia.com/t/error-in-tao-toolkit-while-training/215066/15 "2022-07-06T06:34:43Z")

</div>

**There is no update from you for a period, assuming this is not an issue anymore.  
Hence we are closing this topic. If need further support, please open a new one.  
Thanks**

---

<div class="post-metadata">

**Author:** ![system](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/system/32/68080_2.png) [@system](https://forums.developer.nvidia.com/u/system)\
**Post date:** [July 20, 2022, 6:35am UTC](https://forums.developer.nvidia.com/t/error-in-tao-toolkit-while-training/215066/16 "2022-07-20T06:35:01Z")

</div>

This topic was automatically closed 14 days after the last reply. New replies are no longer allowed.
