# cache data in shared memory for subsequent calls

**URL:** <https://forums.developer.nvidia.com/t/cache-data-in-shared-memory-for-subsequent-calls/16630>\
**Category:** CUDA Programming and Performance\
**Created:** [May 14, 2010, 9:05pm UTC](https://forums.developer.nvidia.com/t/cache-data-in-shared-memory-for-subsequent-calls/16630 "2010-05-14T21:05:05Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ben\_Jiang](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@Ben\_Jiang](https://forums.developer.nvidia.com/u/Ben_Jiang)\
**Post date:** [May 14, 2010, 9:05pm UTC](https://forums.developer.nvidia.com/t/cache-data-in-shared-memory-for-subsequent-calls/16630/1 "2010-05-14T21:05:05Z")

</div>

I have this app which I beleive it’s best to cache data in shared memory for many many subsequent calls. The scenario is:

- the cache data never changes

- It will then have a long running process (loop forever) which just repeatly invoke the compute function from user input

Would someone verify if the below code will work? I am interested to know if s\_data will hold the right data across the first kernel call and the subsequential kernel calls…

Thanks in advance. Ben

The code follow:==========================

```auto
extern __shared__ s_data[];

__gloabl__ kernelLoadData(datapool){

		// copy data into s_data

	....

	threadIdx.x ...

}

__gloabl__ kernelCompute(signal){

	// access s_data and compute it with signal...

	....

}

void main(){

	// load data from disk:

	float* datapool;

	...

	// compute blockDim and thread count:

	dim3 n_block, block_size;

	// load once to each block:

	kernelLoadData<<<n_block, block_size>>>(datapool);

	float *signal;

	while(true){

		// wait for input

		....

		signal = ....	// some input func

		// compute from the input:

		kernelCompute<<<n_block, block_size>>>(signal);

		// read back from computed data...

	}

}

```

---

<div class="post-metadata">

**Author:** ![tmurray](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@tmurray](https://forums.developer.nvidia.com/u/tmurray)\
**Post date:** [May 14, 2010, 9:16pm UTC](https://forums.developer.nvidia.com/t/cache-data-in-shared-memory-for-subsequent-calls/16630/2 "2010-05-14T21:16:05Z")

</div>

The contents of shared memory at the beginning of any kernel call are undefined, so no, that will not work.

---

<div class="post-metadata">

**Author:** ![Jimmy\_Pettersson](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@Jimmy\_Pettersson](https://forums.developer.nvidia.com/u/Jimmy_Pettersson)\
**Post date:** [May 14, 2010, 11:24pm UTC](https://forums.developer.nvidia.com/t/cache-data-in-shared-memory-for-subsequent-calls/16630/3 "2010-05-14T23:24:44Z")

</div>

Perhaps you would want to do copy the data into constant memory space… If all threads of the warp are accessing the same element this might be a good option for you.

In your function main you just do a cudaMemcopyToSymbol(…) to a constant variable that has been declared in the global scope.

**constant** float read\_only\_array[length];

---

<div class="post-metadata">

**Author:** ![Ben\_Jiang](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@Ben\_Jiang](https://forums.developer.nvidia.com/u/Ben_Jiang)\
**Post date:** [May 25, 2010, 5:58pm UTC](https://forums.developer.nvidia.com/t/cache-data-in-shared-memory-for-subsequent-calls/16630/4 "2010-05-25T17:58:44Z")

</div>

Thanks, tmurray and Jimmy.

I have finished the work;). Anyway, what I used was: load the data into global memory and stay cached there. The data will then be read into shared memory for faster access. Both shared memory or constant memory are too small, since I have at least 20MB of float numbers. I’d need 16 GTX470 to load all of them in shared or constant memory;).

The bad news is I cannot cache the 20 MB data in shared memory, since each float will only be used once.

Anyhow, it works. Still way faster than CPU;), roughly by 20 times (40s-\>2s).

---

<div class="post-metadata">

**Author:** ![Ben\_Jiang](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@Ben\_Jiang](https://forums.developer.nvidia.com/u/Ben_Jiang)\
**Post date:** [May 25, 2010, 5:58pm UTC](https://forums.developer.nvidia.com/t/cache-data-in-shared-memory-for-subsequent-calls/16630/5 "2010-05-25T17:58:44Z")

</div>

Thanks, tmurray and Jimmy.

I have finished the work;). Anyway, what I used was: load the data into global memory and stay cached there. The data will then be read into shared memory for faster access. Both shared memory or constant memory are too small, since I have at least 20MB of float numbers. I’d need 16 GTX470 to load all of them in shared or constant memory;).

The bad news is I cannot cache the 20 MB data in shared memory, since each float will only be used once.

Anyhow, it works. Still way faster than CPU;), roughly by 20 times (40s-\>2s).
