Trouble Running Multiple Isaac Sim Processes

Isaac Sim Version

4.5.0
4.2.0
4.1.0
4.0.0
4.5.0
2023.1.1
2023.1.0-hotfix.1
Other (please specify):

Operating System

Ubuntu 22.04
Ubuntu 20.04
Windows 11
Windows 10
Other (please specify):

GPU Information

  • Model: RTX A6000 x 4
  • Driver Version: 560.35.03

Topic Description

Detailed Description

I’m experimenting with generating synthetic data with Isaac Sim and Replicator. My standalone script runs fine on a single GPU and a single process. However, I would like to utilize all four GPUs to speed up data generation. I tried to start four processes of Isaac Sim, each utilizing one GPU, with the commands:

./python.sh my_script.py --/renderer/multiGpu/enable=false --/renderer/activeGpu=<gpu_id>

The first two processes start just fine and output data as expected. The third and fourth processes would hang with these in the terminal:

[12.048s] Simulation App Starting
2025-04-21 19:41:24 [12,098ms] [Warning] [omni.kvdb.plugin] Disabling key-value database because another kit process is locking it
[12.197s] app ready
2025-04-21 19:41:24 [12,374ms] [Warning] [rtx.scenedb.plugin] SceneDbContext : TLAS limit buffer size 7374781440
2025-04-21 19:41:24 [12,374ms] [Warning] [rtx.scenedb.plugin] SceneDbContext : TLAS limit : valid false, within: false
2025-04-21 19:41:24 [12,374ms] [Warning] [rtx.scenedb.plugin] SceneDbContext : TLAS limit : decrement: 167690, decrement size: 7301033856
2025-04-21 19:41:24 [12,374ms] [Warning] [rtx.scenedb.plugin] SceneDbContext : New limit 9748724 (slope: 439, intercept: 13179904)
2025-04-21 19:41:24 [12,374ms] [Warning] [rtx.scenedb.plugin] SceneDbContext : TLAS limit buffer size 4287352704
2025-04-21 19:41:24 [12,374ms] [Warning] [rtx.scenedb.plugin] SceneDbContext : TLAS limit : valid true, within: true
2025-04-21 19:41:24 [12,569ms] [Warning] [omni.usd-abi.plugin] No setting was found for '/rtx-defaults-transient/meshlights/forceDisable'
2025-04-21 19:41:24 [12,626ms] [Warning] [omni.usd-abi.plugin] No setting was found for '/rtx-defaults/post/dlss/execMode'

and the “Simulation App Startup Complete” would not show up, with no error message.

Additional Information

What I’ve Tried

I’ve tried swapping the order of running the processes. Seems like it is consistent that there can only be two Isaac Sim running, and the third one would always hang as described. I’ve also made each process write to a different folder to avoid IO conflicts.

Additional Context

In my script, I create two render products for a set of stereo cameras and render left and right images.
The first thing I tried was running a single Isaac Sim process using four GPUs. There is a weird bug where the right half of the right image would be replaced by the right half of the left image.
If I run my script with --/renderer/multiGpu/maxGpuCount=2, the bug would be resolved, but it would not utilize the rest of the 2 GPUs.
Then I tried to run one process per GPU and found that and ran into the hanging of the third process issue as described above.

Does this issue occur only for a specific script? Can you reproduce it by running the standalone example script ~/isaacsim/standalone_examples/tutorials/getting_started.py?

Could you share the full logs from running the first, second, and third process?

The error messages is telling you that you are maxed out on shader cache for the materials database. This is a hard limit. Kit is not designed to run like this.

If you want to run multiple instances on one machine, use Containers. Not direct desktop launches.

Here is a good tip. Rather than trying to run four kit instances, each using only 1 GPU, reverse that and run 1 kit instance, one simulation, and let it use all four GPUs. Much more efficient. You forget that running multiple kit apps, takes multiple cpu threads, multiple calls to the same system memory, multiple calls the same hard drive. It’s much more stress of the system.

Put all your GPU power into finishing ONE task at 4x speed, that 4 GPUs struggling to all work on different tasks.

Thank you for your reply. I have reproduced the issue with getting_started.py. However, I had to make two changes to run three instances simultaneously: 1. change headless to True; 2. increase simulation iteration from 3 to 100.

Here are the logs.
kit_20250422_164236.log (797.7 KB)
kit_20250422_164214.log (839.5 KB)
kit_20250422_164042.log (837.5 KB)

The third process seems to wait forever for some operation:

...
2025-04-22 20:48:27 [351,312ms] [Info] [gpu.foundation.plugin] *** Waiting for RtPso async group async compilation: 155 seconds so far
2025-04-22 20:48:32 [356,312ms] [Info] [gpu.foundation.plugin] *** Waiting for RtPso async group async compilation: 160 seconds so far
2025-04-22 20:48:37 [361,312ms] [Info] [gpu.foundation.plugin] *** Waiting for RtPso async group async compilation: 165 seconds so far

Thank you for your reply.

I get these SceneDbContext warnings even if I run a single process. I have posted the logs to the reply above. In short, I saw some suspicious output:

[Info] [gpu.foundation.plugin] *** Waiting for RtPso async group async compilation: 165 seconds so far

Could you confirm that this is still a shader cache maxed out problem?

Regarding efficiency of running 1 task with 4 GPUs vs 4 tasks with 1 GPU each, according to this documentation, running 1 task with multiple GPUs is not as efficient: Isaac Sim Performance Optimization Handbook — Isaac Sim Documentation

When rendering 2 720p cameras with 2 GPUs, we saw a speed up of 72% to 89% compared to single GPU performance, but using 4 GPUs yielded only 61 - 81% improvement.

Although the kit might not be designed to run this way, I think there are potential advantages if running multiple instances is made possible.

Finally, I am running into another issue when I run 1 task with 4 GPUs. That is actually the main reason I was experimenting with a multi-process approach. I’ve made another post about the details of that issue here:

Let me explain a little more on that efficiency post. It does say only 720p. That is very low resolution. Most outputs these days are at least 4k, 2160p. It does say right below that, that high resolution frames benefit greatly from multiple gpus.

Also something major which you might have missed. None of those tips are talking about multiple instances of kit. They all mean 1 instance of kit with multiple viewports. 2 viewports for 2 gpus, 4 viewports for 4 gpus, and so on. So again, look at running ONE kit instance, with 4 viewports, each one assigned to an individual gpu (which actually happens automatically) and then you can render at 100% per viewport, per gpu.

“but our 8 camera with 4 GPUs test scaled even better with an overall speedup of 271% - 281%.” - This is referring to ONE kit session with 8 cameras. Not 8 kit sessions. That would be a nightmare.

Thank you very much for your detailed explanation!

Let’s say I have 8 GPUs, but I only need to render 1 720p image per data point. The best approach would be running 1 kit application, in which I can run 8 parallel simulations like a grid in the stage, then have 1 camera for each simulation?

When you say Simulation, are you talking about actually simulating something, or just rendering out to replicator or Movie Capture? Remember that a lot of simulation happens in the cpu, not just the gpu. Asking a single cpu to do 8 complex things at once is not a good idea. As mentioned above, people conflate 8 GPUs, with having 8 times the power of a machine with 1 gpu. It’s like a car with 8 engines. Yes it has a lot of power, but there is a rate of return far lower than 8x speed. How fast can you go with one car, one driver, one road? You see what I am saying?

So you are rendering out just low res 720p images from Replicator? Why so low resolution?

Are you rendering with the realtime rtx renderer or path tracing? A 720p image should take literally 1 second to render in realtime. In path tracing mode, with 8 GPUs, maybe 5 seconds.

It’s hardly worth even splitting up a job like that. How many images do you have to render? 1000, 10,000 etc ? Honestly 8 GPUs are not much good here. It may be overkill. Sometimes more is less.

If had 10,000 frames to render at 720p, each frame taking say 5 seconds a frame, even on my single gpu, I would just let it crank over night.

If you wanted it to go faster, you are better off with splitting up the job over a basic 8 machine render farm. 8 machines each running a basic single 3090 l, would EASILY outperform a 8xGPU machine. We like to think everything scales linearly, but it doesn’t. After 2x, the cpu is on fire, your memory is maxed, the hard drive is screaming. The only and best way to fully use 8 GPUs is on a massive massive single super high resolution frame in path tracing mode, all working together.

Do you own this 8 way GPU machine or you are renting it? I personally think it’s better to have 8 machines with single cards, or at least 4 machines with dual cards, than one monster machine. There comes a point when more power in a single machine just cannot be used efficiently in parallel in realtime.

If I were you, consider taking out 6 of those cards, and sticking them in three more cheap basic machines, 2 cards per machine, and you would get more parallel work done.

Thanks for the suggestions!

My task is to spawn objects in a procedural way, run physics simulation, and then render images for the settled scene. If I understand it correctly, without enabling GPU physics, the scene setup and simulation will run on CPU, and the rendering will run on GPU. In that case, there will be times where the GPUs are idle waiting for the CPU simulation, as well as CPU idle waiting for GPU rendering to finish. Before Isaac Sim I used to run this SDG task with Blender, where there exists a simple solution–run multiple Blender instances executing the same SDG script, and let system handle the scheduling of resources, so at least one of CPU/GPU/Disk is utilized near 100%.

I have 4 GPU and 8 GPU machines for neural network training, therefore it is not practical to distribute the GPUs to multiple machines. I believe this would be the case for most SDG use cases. Also for network training, I think 720p resolution is reasonable for most network tasks.

So in that case then, that leaves you will two solutions.

Solution one, you should just run one version of kit and let it rip through the work with all the gpu’s running, assuming you are using path tracing, which you have not indicated yet. If you are using real-time, you will not see much difference. If you are feeling good about that, try two kit sessions and split the work, and see if you are getting the same overall time saving, when doing that math. I doubt it. You will see a significant drop off.

Solution two, If you want “true” parallelization, you have to switch to using Linux and containers or K8s and assign each gpu to each container.

This is not really on Omniverse problem. This problem exists with any computer hardware and software. You either put all 8 horses to pull one carriage or you put 4 horses each on two carriages. At the end of the day, the next result is the same.