# Cuda Samples Scan Query

**URL:** <https://forums.developer.nvidia.com/t/cuda-samples-scan-query/34662>\
**Category:** CUDA Programming and Performance\
**Created:** [September 1, 2014, 6:10am UTC](https://forums.developer.nvidia.com/t/cuda-samples-scan-query/34662 "2014-09-01T06:10:30Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![sedona](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@sedona](https://forums.developer.nvidia.com/u/sedona)\
**Post date:** [September 1, 2014, 6:10am UTC](https://forums.developer.nvidia.com/t/cuda-samples-scan-query/34662/1 "2014-09-01T06:10:30Z")

</div>

Please can someone please straighten me out on the scan example in Cuda 6.0. I was expecting it to take some sequence of numbers and create a summed scan version, inter or intra-scan according to the input count N. But the sample actually seems to take a supplied number arrayLength, and then scans the input into lots of scanned segments of the original data, each arrayLength long. Am I reading this right, and if so, how is this helpful?

---

<div class="post-metadata">

**Author:** ![little\_jimmy](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@little\_jimmy](https://forums.developer.nvidia.com/u/little_jimmy)\
**Post date:** [September 1, 2014, 8:07am UTC](https://forums.developer.nvidia.com/t/cuda-samples-scan-query/34662/2 "2014-09-01T08:07:47Z")

</div>

it depends on whether you wish to scan or merely reduce: end up with a summed sequence, or just a sum

but more importantly, it depends on the input array length, and ensuring that it can fit on the device, given the device’s max applicable thread block dimension

for example, you can not fit a 5k element array on the device, without breaking it up into smaller sub sequences

it is generally very easy to move from scanned sub-sequences to a scanned sequence, if you want to scan instead of just reducing

---

<div class="post-metadata">

**Author:** ![Robert\_Crovella](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/robert_crovella/32/14043_2.png) [@Robert\_Crovella](https://forums.developer.nvidia.com/u/Robert_Crovella)\
**Post date:** [September 1, 2014, 12:57pm UTC](https://forums.developer.nvidia.com/t/cuda-samples-scan-query/34662/3 "2014-09-01T12:57:16Z")

</div>

The scanExclusiveLarge function can do batches (of the same size scans) in parallel. If you only want to do a single scan, pass 1 as the 3rd parameter (batchSize) and whatever is the length of your array as the 4th parameter (arrayLength). Refer to the scan.cu file. Depending on the size of your array, you would use different scan functions. These sizes are delineated in scan.cu such as MIN\_SHORT\_ARRAY\_SIZE and MIN\_LARGE\_ARRAY\_SIZE, etc.

---

<div class="post-metadata">

**Author:** ![sedona](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@sedona](https://forums.developer.nvidia.com/u/sedona)\
**Post date:** [September 1, 2014, 10:30pm UTC](https://forums.developer.nvidia.com/t/cuda-samples-scan-query/34662/4 "2014-09-01T22:30:17Z")

</div>

> [@txbob](#):
>
> The scanExclusiveLarge function can do batches (of the same size scans) in parallel. If you only want to do a single scan, pass 1 as the 3rd parameter (batchSize) and whatever is the length of your array as the 4th parameter (arrayLength). Refer to the scan.cu file. Depending on the size of your array, you would use different scan functions. These sizes are delineated in scan.cu such as MIN\_SHORT\_ARRAY\_SIZE and MIN\_LARGE\_ARRAY\_SIZE, etc.

Sadly there’s a factorRadix2 check against the arrayLength, which limits usefulness.

---

<div class="post-metadata">

**Author:** ![sedona](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@sedona](https://forums.developer.nvidia.com/u/sedona)\
**Post date:** [September 1, 2014, 10:41pm UTC](https://forums.developer.nvidia.com/t/cuda-samples-scan-query/34662/5 "2014-09-01T22:41:13Z")

</div>

> [@little\_jimmy](#):
>
> it depends on whether you wish to scan or merely reduce: end up with a summed sequence, or just a sum
> 
> but more importantly, it depends on the input array length, and ensuring that it can fit on the device, given the device’s max applicable thread block dimension
> 
> for example, you can not fit a 5k element array on the device, without breaking it up into smaller sub sequences
> 
> it is generally very easy to move from scanned sub-sequences to a scanned sequence, if you want to scan instead of just reducing

I think I’m mostly struggling with the batched array concept itself. Assuming a a large input set, I thought that we’d just get each block to scan its segment of the data, and then we launch subsequent kernels to update the L1 subscans with the global results. So why do we need the batching?

---

<div class="post-metadata">

**Author:** ![Robert\_Crovella](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/robert_crovella/32/14043_2.png) [@Robert\_Crovella](https://forums.developer.nvidia.com/u/Robert_Crovella)\
**Post date:** [September 2, 2014, 3:23am UTC](https://forums.developer.nvidia.com/t/cuda-samples-scan-query/34662/6 "2014-09-02T03:23:08Z")

</div>

You could pad your array up to the next size that satisfies the factorRadix2 check. And it is, after all, a sample code, not a production library.

If you’re just looking for a handy scan function, thrust and cub both have implementations that should be pretty flexible.
