# Avoid synchronization in optixLaunch

**URL:** https://forums.developer.nvidia.com/t/avoid-synchronization-in-optixlaunch/221458
**Category:** OptiX
**Created:** [July 21, 2022, 11:45am UTC](https://forums.developer.nvidia.com/t/avoid-synchronization-in-optixlaunch/221458 "2022-07-21T11:45:12Z")
**Posts on this page:** 6
**Page:** 1

<div class="post-metadata">

### Author: ![wenzel.jakob](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/wenzel.jakob/32/31120_2.png) [@wenzel.jakob](https://forums.developer.nvidia.com/u/wenzel.jakob)
#### Post date: [July 21, 2022, 11:45am UTC](https://forums.developer.nvidia.com/t/avoid-synchronization-in-optixlaunch/221458/1 "2022-07-21T11:45:12Z")

</div>

Dear all,

I am currently profiling an application using NSight Systems to remove CPU\<-\>GPU synchronization points and improve performance. While doing this, I realized that `optixLaunch` appears to call `cuStreamSynchronize` internally (nsight systems can be set up to capture the backtrace of CUDA synchronization calls, and `optixTrace` shows up as cause).

My code submits work to a custom CUDA stream that has been created in non-blocking mode, and that includes the “optixLaunch” call. I am _not_ generating concurrent OptiX launches, what I want is simply that the CPU can run ahead while the GPU is busy to generate the launch needed for the next frame. However, this all falls apart if `optixLaunch` then does further synchronization.

Ideas?

Thanks,  
Wenzel

---

<div class="post-metadata">

### Author: ![wenzel.jakob](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/wenzel.jakob/32/31120_2.png) [@wenzel.jakob](https://forums.developer.nvidia.com/u/wenzel.jakob)
#### Post date: [July 21, 2022, 12:02pm UTC](https://forums.developer.nvidia.com/t/avoid-synchronization-in-optixlaunch/221458/2 "2022-07-21T12:02:44Z")

</div>

Another potentially synchronization-related issue that I observe is that `optixAccelBuild` generates various “NVIDIA internal” kernels, but then ends up stuck for a long time in a `cuMemcpyAsync` operation. What seems to be happening here is that the acceleration structure build is actually synchronizing with the CPU that is now waiting for a preceding rendering step to finish.

 ![Screenshot from 2022-07-21 14-01-48](https://global.discourse-cdn.com/nvidia/original/3X/5/2/521eab902ce07d5ce3b4889dc69641ff39d2fa31.png)

---

<div class="post-metadata">

### Author: ![wenzel.jakob](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/wenzel.jakob/32/31120_2.png) [@wenzel.jakob](https://forums.developer.nvidia.com/u/wenzel.jakob)
#### Post date: [July 21, 2022, 12:05pm UTC](https://forums.developer.nvidia.com/t/avoid-synchronization-in-optixlaunch/221458/3 "2022-07-21T12:05:57Z")

</div>

In case it matters, this is on x86\_64 linux (Ubuntu 20.04) with driver 515.48.08.

---

<div class="post-metadata">

### Author: ![dhart](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/dhart/32/14043_2.png) [@dhart](https://forums.developer.nvidia.com/u/dhart)
#### Post date: [July 21, 2022, 5:37pm UTC](https://forums.developer.nvidia.com/t/avoid-synchronization-in-optixlaunch/221458/4 "2022-07-21T17:37:18Z")

</div>

Hi @wenzel.jakob!

So looking at the code for `optixLaunch()`, I do see one call to `cuStreamSychnorize()` that happens only when validation mode debug exceptions are enabled. Is that the case here, do you have validation mode enabled? (Validation mode intentionally serializes OptiX launches in order to facilitate debugging and rule out all the difficult async problems.)

–  
David.

---

<div class="post-metadata">

### Author: ![dhart](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/dhart/32/14043_2.png) [@dhart](https://forums.developer.nvidia.com/u/dhart)
#### Post date: [July 21, 2022, 5:50pm UTC](https://forums.developer.nvidia.com/t/avoid-synchronization-in-optixlaunch/221458/5 "2022-07-21T17:50:02Z")

</div>

For the `optixAccelBuild()`, that sounds different. I’m just thinking out loud… is the device pointer used for the memcpy also used in any preceding stream API calls? Is there any possibility that the host buffer could be paged out when the call is made? I’ll ask the CUDA team what reasons a `cudaMemcpyAsync()` call might stall or synchronize.

–  
David.

---

<div class="post-metadata">

### Author: ![dhart](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/dhart/32/14043_2.png) [@dhart](https://forums.developer.nvidia.com/u/dhart)
#### Post date: [July 21, 2022, 5:54pm UTC](https://forums.developer.nvidia.com/t/avoid-synchronization-in-optixlaunch/221458/6 "2022-07-21T17:54:10Z")

</div>

Perhaps this thread is relevant? Robert notes that multiple memory operations will serialize due to PCI rules, and the rest of the thread has some good hints too I think: [cudaMemcpyAsync - #2 by Robert\_Crovella](https://forums.developer.nvidia.com/t/cudamemcpyasync/38061/2)

–  
David.
