# CUDA Warp primitive behaviour question

**URL:** <https://forums.developer.nvidia.com/t/cuda-warp-primitive-behaviour-question/289385>\
**Category:** CUDA Programming and Performance\
**Created:** [April 12, 2024, 5:46am UTC](https://forums.developer.nvidia.com/t/cuda-warp-primitive-behaviour-question/289385 "2024-04-12T05:46:28Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![794906124](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@794906124](https://forums.developer.nvidia.com/u/794906124)\
**Post date:** [April 12, 2024, 5:46am UTC](https://forums.developer.nvidia.com/t/cuda-warp-primitive-behaviour-question/289385/1 "2024-04-12T05:46:28Z")

</div>

Hello all, Now I’m learning the warp level data exchange primitives. I have some question on the `mask` parameter.

I’m now conducting the intra-warp data exchange with following code, which matches my expected output well:

```auto
if ((1 << lane_id) & (uint32_t) 0x5a5a5a5a) {
        *(uint32_t*)(&data[0]) = __shfl_xor_sync(0xffffffff, *(uint32_t*)(&data[0]), 5);
    }

```

While I want to use the `mask` parameter to do the same, as:

```auto
*(uint32_t*)(&data[0]) = __shfl_xor_sync(0x5a5a5a5a, *(uint32_t*)(&data[0]), 5);

```

The result indicated that **ALL** threads in the warp participate the data exchange, while I only want the threads with `threadIdx.x % 8 == 1,3,4,6` to participate the data exchange (So I set the `mask` to be `0x5a5a5a5a (0x5a=0b 0101 1010)` which exactly indicate the 1, 3, 4 and 6 thread). How should I use the `mask` parameter to achieve my expected effect?

---

<div class="post-metadata">

**Author:** ![Robert\_Crovella](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/robert_crovella/32/14043_2.png) [@Robert\_Crovella](https://forums.developer.nvidia.com/u/Robert_Crovella)\
**Post date:** [April 12, 2024, 3:46pm UTC](https://forums.developer.nvidia.com/t/cuda-warp-primitive-behaviour-question/289385/2 "2024-04-12T15:46:07Z")

</div>

The mask doesn’t exclude threads. The mask guarantees convergence to at least the level specified in the mask. But the mask does not exclude threads “not selected” in the mask, nor does it prevent those “not selected” threads from participating in the op.

If you want only certain threads to participate, use a boolean if condition to select those threads, before running the primitive. Make sure your mask is consistent with your conditional selection of threads.

- Depending on results not supported by the mask is undefined behavior.
- Make sure that both source and necessary destination lanes are included in the mask as well as appropriately selected by your conditional code.

---

<div class="post-metadata">

**Author:** ![794906124](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@794906124](https://forums.developer.nvidia.com/u/794906124)\
**Post date:** [April 13, 2024, 2:29am UTC](https://forums.developer.nvidia.com/t/cuda-warp-primitive-behaviour-question/289385/3 "2024-04-13T02:29:18Z")

</div>

Thanks for your reply, guess it’s good to use my current solution.
