# What does mask mean in warp shuffle functions (\_\_shfl\_sync)

**URL:** https://forums.developer.nvidia.com/t/what-does-mask-mean-in-warp-shuffle-functions-shfl-sync/67697
**Category:** CUDA Programming and Performance
**Created:** [November 22, 2018, 8:41am UTC](https://forums.developer.nvidia.com/t/what-does-mask-mean-in-warp-shuffle-functions-shfl-sync/67697 "2018-11-22T08:41:49Z")
**Posts on this page:** 1
**Showing post:** 1

<div class="post-metadata">

### Author: ![jijn](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@jijn](https://forums.developer.nvidia.com/u/jijn)
#### Post date: [November 22, 2018, 8:41am UTC](https://forums.developer.nvidia.com/t/what-does-mask-mean-in-warp-shuffle-functions-shfl-sync/67697/1 "2018-11-22T08:41:49Z")

</div>

I am trying to understand the **mask** parameter in shuffle functions, e.g.

```auto
T __shfl_down_sync(unsigned mask, T var, unsigned int delta, int width=warpSize);

```

My understanding is that only threads indicated by 1 bits in ‘mask’ do the exchange. If mask = 0xffffffff, it means all threads in the warp will do the exchange. If mask = 0x00000000, no threads will do the exchange. However, I didn’t get the expected result. Following is the test code.

```auto
#include <stdio.h>
#include "cuda_runtime.h"

__global__ void kernel(double *a){
  double v = a[threadIdx.x];

  unsigned mask = 0x00000000;//0xffffffff;//0x000000ff;
  unsigned int offset = 4;
  v += __shfl_down_sync(mask, v, offset, 8);

  a[threadIdx.x] = v;
}

void main(){
  double *a, *a_d;
  a = (double*)calloc(32,sizeof(double));
  cudaMalloc((void **)&a_d,32*sizeof(double));
  for(int i=0;i<32;i++){ a[i]=i/4; }
  cudaMemcpy(a_d, a, 32*sizeof(double), cudaMemcpyHostToDevice);

  for(int i=0;i<32;i++){ printf("%2.0f ",a[i]); }
  printf("\n");

  kernel<<<1,32>>>(a_d);

  cudaMemcpy(a, a_d, 32*sizeof(double), cudaMemcpyDeviceToHost);
  for(int i=0;i<32;i++){ printf("%2.0f ",a[i]); }
}

```

I got the same results (all threads do the exchange) no matter which mask is used (mask=0xffffffff, mask=0x00000000, mask=0x000000ff).

I am wondering if it is a bug or my understanding of the mask parameter is wrong? I am using a Tesla V100 and CUDA 10.0.

Thanks.

---

_[View the full topic](https://forums.developer.nvidia.com/t/what-does-mask-mean-in-warp-shuffle-functions-shfl-sync/67697)._
