# Thrust::async::for\_each() with zip\_iterators

**URL:** <https://forums.developer.nvidia.com/t/thrust-for-each-with-zip-iterators/240977>\
**Category:** CUDA Programming and Performance\
**Created:** [January 29, 2023, 8:17pm UTC](https://forums.developer.nvidia.com/t/thrust-for-each-with-zip-iterators/240977 "2023-01-29T20:17:48Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![koltona](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@koltona](https://forums.developer.nvidia.com/u/koltona)\
**Post date:** [January 29, 2023, 8:17pm UTC](https://forums.developer.nvidia.com/t/thrust-for-each-with-zip-iterators/240977/1 "2023-01-29T20:17:48Z")

</div>

When I try to use thrust::async::for\_each() with zip iterators, the code [1] does not compile ([2]).

It does work however if I use the “sync” version thrust::for\_each() with zip iterators [3] or if I use two calls of thrust::async::for\_each() with normal iterators [4].

I am using hpc\_sdk/Linux\_x86\_64/22.3/ [5,6].

Thanks

====  
[1]

```auto
//test.cu
#include <thrust/device_vector.h>
#include <thrust/async/for_each.h>
#include <thrust/copy.h>
#include <iostream>
#include <chrono>

const int ds = 10000000;

int main(){
  thrust::device_vector<int> d_A(ds,1);
  thrust::device_vector<int> d_B(ds,2);
    
  auto t1 = std::chrono::steady_clock::now(); // Start timing     

    ///////////////
    #ifdef ZIPSYNC // this works
    
    auto e12 = thrust::for_each(
    thrust::make_zip_iterator(thrust::make_tuple(d_A.begin(),d_B.begin())), 
    thrust::make_zip_iterator(thrust::make_tuple(d_A.end(),d_B.end())), 
    [=] __device__ (auto &tup){
        if (thrust::get<0>(tup)==1) thrust::get<0>(tup)++;
        if (thrust::get<1>(tup)==2) thrust::get<1>(tup)++;
    }    
    );
    
    #elif defined(ZIPASYNC) // this does not compile
    
    auto e12 = thrust::async::for_each(
    thrust::make_zip_iterator(thrust::make_tuple(d_A.begin(),d_B.begin())), 
    thrust::make_zip_iterator(thrust::make_tuple(d_A.end(),d_B.end())), 
    [=] __device__ (auto &tup){
        if (thrust::get<0>(tup)==1) thrust::get<0>(tup)++;
        if (thrust::get<1>(tup)==2) thrust::get<1>(tup)++;
    }    
    );
    
    #else // this works
    auto e1 = thrust::async::for_each(d_A.begin(), d_A.end(),  
    [=] __device__ (auto &t){
      if (t==1) t++;
    }
    );
    auto e2 = thrust::async::for_each(d_B.begin(), d_B.end(), 
        [=] __device__ (auto &t){
          if (t==2) t++;
        }
    );
    #endif
  
    auto t2 = std::chrono::steady_clock::now();       
    cudaDeviceSynchronize();
    auto t3 = std::chrono::steady_clock::now();   
  
    thrust::copy_n(d_A.begin(), 5, std::ostream_iterator<int>(std::cout, ","));
    thrust::copy_n(d_B.begin(), 5, std::ostream_iterator<int>(std::cout, ","));
    std::cout << std::endl;
  
    std::cout
    << "before cudaDeviceSynchronize "
    << std::chrono::duration_cast<std::chrono::microseconds>(t2 - t1).count()
    << std::endl;
  
    std::cout
    << "after cudaDeviceSynchronize "
    << std::chrono::duration_cast<std::chrono::microseconds>(t3 - t1).count()
    << std::endl;
}

```

[2] nvcc --extended-lambda test.cu -DZIPASYNC  
[3] nvcc --extended-lambda test.cu -DZIPSYNC  
[4] nvcc --extended-lambda test.cu

[5]  
which nvcc  
/opt/nvidia/hpc\_sdk/Linux\_x86\_64/22.3/compilers/bin/nvcc

[6]  
nvcc --version  
nvcc: NVIDIA (R) Cuda compiler driver  
Copyright (c) 2005-2022 NVIDIA Corporation  
Built on Thu\_Feb\_10\_18:23:41\_PST\_2022  
Cuda compilation tools, release 11.6, V11.6.112  
Build cuda\_11.6.r11.6/compiler.30978841\_0

---

<div class="post-metadata">

**Author:** ![striker159](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@striker159](https://forums.developer.nvidia.com/u/striker159)\
**Post date:** [January 30, 2023, 5:46am UTC](https://forums.developer.nvidia.com/t/thrust-for-each-with-zip-iterators/240977/2 "2023-01-30T05:46:21Z")

</div>

> [@koltona](#):
>
> ```auto
> auto e12 = thrust::async::for_each(
> thrust::make_zip_iterator(thrust::make_tuple(d_A.begin(),d_B.begin())), 
> thrust::make_zip_iterator(thrust::make_tuple(d_A.end(),d_B.end())), 
> [=] __device__ (auto tup){
> if (thrust::get<0>(tup)==1) thrust::get<0>(tup)++;
> if (thrust::get<1>(tup)==2) thrust::get<1>(tup)++;
> }    
> );
> 
> ```

This seems to work. I removed the ampersand after auto.

---

<div class="post-metadata">

**Author:** ![koltona](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@koltona](https://forums.developer.nvidia.com/u/koltona)\
**Post date:** [January 30, 2023, 5:50pm UTC](https://forums.developer.nvidia.com/t/thrust-for-each-with-zip-iterators/240977/3 "2023-01-30T17:50:31Z")

</div>

Great. I confirm that removing the ampersand of the tuple works.

On one hand, I do not understand why the (auto &tup) seems to work for thrust::for\_each() [1] but not for thrust::async::for\_each(), the only difference being the async.

On the other hand, it is clear from the test that (auto &tup) is not only unnecessary but also prone to error in the async case. The ampersand is nevertheless necessary for normal iterators (i.e. not packed in a tuple through zip\_iterator), otherwise it compiles, but we do not modify the arrays.

So, Is it in general incorrect to use the ampersand with zip\_iterators?

[1]  
I have also checked in

> <https://github.com/NVIDIA/thrust/blob/master/examples/arbitrary_transformation.cu>

that if I use (Tuple &t) instead of (Tuple t) in arbitrary\_functor1 the code compiles and run ok.

---

<div class="post-metadata">

**Author:** ![system](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/system/32/68080_2.png) [@system](https://forums.developer.nvidia.com/u/system)\
**Post date:** [February 13, 2023, 5:50pm UTC](https://forums.developer.nvidia.com/t/thrust-for-each-with-zip-iterators/240977/4 "2023-02-13T17:50:55Z")

</div>

This topic was automatically closed 14 days after the last reply. New replies are no longer allowed.
