# \_\_syncthreads() code execution hangs

**URL:** <https://forums.developer.nvidia.com/t/syncthreads-code-execution-hangs/1495>\
**Category:** CUDA Programming and Performance\
**Created:** [September 18, 2007, 8:17am UTC](https://forums.developer.nvidia.com/t/syncthreads-code-execution-hangs/1495 "2007-09-18T08:17:16Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![sicb0161](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@sicb0161](https://forums.developer.nvidia.com/u/sicb0161)\
**Post date:** [September 18, 2007, 8:17am UTC](https://forums.developer.nvidia.com/t/syncthreads-code-execution-hangs/1495/1 "2007-09-18T08:17:16Z")

</div>

Hello,

I am computing the dot product, similar to the example (nvidia projects).

```auto
// Tree - like reduction

	if (thx < i){

  for(int stride = i / 2; stride > 0; stride >>= 1){

  	__syncthreads();

  	//shared_h[thx] += shared_h[stride + thx];

  }

	}

```

In my version the vector lengths must not be a power of two, so that I put the condition thx \< i, as the tree like reduction needs vector lengths equal to the power of two.

The problem is that the code hangs when the number of threads exceeds 16.

Why is that?

In the Programming Guide it says that

> [@](#):
>
> \_\_syncthreads() is allowed in conditional code but only of the conditional evaluates identically across the entire threads block…

I am not really sure what that means.

Thanks in advance.

Cem

---

<div class="post-metadata">

**Author:** ![MisterAnderson42](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@MisterAnderson42](https://forums.developer.nvidia.com/u/MisterAnderson42)\
**Post date:** [September 18, 2007, 12:51pm UTC](https://forums.developer.nvidia.com/t/syncthreads-code-execution-hangs/1495/2 "2007-09-18T12:51:18Z")

</div>

It means that the \_\_syncthreads() MUST be called by all threads in the entire block. You have the syncthreads inside an if, so some threads don’t get there. You can fix it by putting the if (thx \< i) inside the for loop.

---

<div class="post-metadata">

**Author:** ![sicb0161](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@sicb0161](https://forums.developer.nvidia.com/u/sicb0161)\
**Post date:** [September 18, 2007, 2:19pm UTC](https://forums.developer.nvidia.com/t/syncthreads-code-execution-hangs/1495/3 "2007-09-18T14:19:07Z")

</div>

> [@](#):
>
> It means that the \_\_syncthreads() MUST be called by all threads in the entire block. You have the syncthreads inside an if, so some threads don’t get there. You can fix it by putting the if (thx \< i) inside the for loop.
> 
> [snapback]252534[/snapback]

wonderfull,

this makes sense!

okay then i have another question:

how come that in some of the program code there is a conditional like

```auto
if(thx == 0){

 do sth.

}

```

like in the example of the scalar product. I thought that the order of how warps are executed is not determined.

thx in advance

---

<div class="post-metadata">

**Author:** ![MisterAnderson42](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@MisterAnderson42](https://forums.developer.nvidia.com/u/MisterAnderson42)\
**Post date:** [September 18, 2007, 2:34pm UTC](https://forums.developer.nvidia.com/t/syncthreads-code-execution-hangs/1495/4 "2007-09-18T14:34:25Z")

</div>

In the scalar product example, the entire block calculates only a single result. It’s bad practice to have multiple threads writing to the same memory location, so the if (thx == 0) is there to make sure that only one thread performs the memory write.

It is true that the order of warp execution is undefined, so the if (thx == 0) _could_ have race condition issues. In the scalarProd example, there has been a \_\_syncthreads() call to make sure all threads are caught up, and then accumResult is updated. Since thread 0 is writing the value from accumResult[0], there cannot be any race condition to access it since thread 0 also updated accumResult[0] a few lines of code up!

In any of the examples that use if (thx == 0), you should see syncthreads used in appropriate locations to prevent race conditions.

---

<div class="post-metadata">

**Author:** ![sicb0161](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@sicb0161](https://forums.developer.nvidia.com/u/sicb0161)\
**Post date:** [September 20, 2007, 4:26pm UTC](https://forums.developer.nvidia.com/t/syncthreads-code-execution-hangs/1495/5 "2007-09-20T16:26:25Z")

</div>

> [@](#):
>
> In the scalar product example, the entire block calculates only a single result. It’s bad practice to have multiple threads writing to the same memory location, so the if (thx == 0) is there to make sure that only one thread performs the memory write.
> 
> It is true that the order of warp execution is undefined, so the if (thx == 0) _could_ have race condition issues. In the scalarProd example, there has been a \_\_syncthreads() call to make sure all threads are caught up, and then accumResult is updated. Since thread 0 is writing the value from accumResult[0], there cannot be any race condition to access it since thread 0 also updated accumResult[0] a few lines of code up!
> 
> In any of the examples that use if (thx == 0), you should see syncthreads used in appropriate locations to prevent race conditions.
> 
> [snapback]252574[/snapback]

I see thx a lot !

I have another question:

I am writing a qr decomposition for smaller matrix sizes. I start to read from global memory to the processing and the try to write back after the calculation to the same global memory, as the qr factorization is an iterative process.

I have checked my calculation for a single iteration step and I get weird errors. But when I write to another global memory location, then there seems to be no calculation error. I have checked all the intermediate results and they are correct.

Is there any amount of time needed so I can write back to the same memory ?

Dont have a clue whats wrong !

thanks in advance,

Cem
