# Cache data invalidation between kernel calls

**URL:** <https://forums.developer.nvidia.com/t/cache-data-invalidation-between-kernel-calls/26838>\
**Category:** CUDA Programming and Performance\
**Created:** [June 6, 2012, 4:37pm UTC](https://forums.developer.nvidia.com/t/cache-data-invalidation-between-kernel-calls/26838 "2012-06-06T16:37:48Z")\
**Posts on this page:** 1\
**Showing post:** 4

<div class="post-metadata">

**Author:** ![Gregory\_Diamos](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@Gregory\_Diamos](https://forums.developer.nvidia.com/u/Gregory_Diamos)\
**Post date:** [June 7, 2012, 12:51am UTC](https://forums.developer.nvidia.com/t/cache-data-invalidation-between-kernel-calls/26838/4 "2012-06-07T00:51:38Z")

</div>

Sorry, I think I misread your question, I thought you were asking how to make CPU writes visible to the GPU before finishing a kernel.

Hopefully this text from the PTX manual about the default cache policy answers you actual question:

“Cache at all levels, likely to be accessed again.  
The default load instruction cache operation is ld.ca, which allocates cache lines in all levels (L1  
and L2) with normal eviction policy. Global data is coherent at the L2 level, but multiple L1  
caches are not coherent for global data. If one thread stores to global memory via one L1 cache,  
and a second thread loads that address via a second L1 cache with ld.ca, the second thread may  
get stale L1 cache data, rather than the data stored by the first thread. The driver must  
invalidate global L1 cache lines between dependent grids of parallel threads. Stores by the first  
grid program are then correctly fetched by the second grid program issuing default ld.ca loads  
cached in L1.”

So only the L1s (not the L2) should be invalidated between dependent kernels. Also note that the  
L1s are write-through by default for global data:

“The default store instruction cache operation is st.wb, which writes back cache lines of coherent  
cache levels with normal eviction policy. Data stored to local per-thread memory is cached in L1  
and L2 with with write-back. However, sm\_20 does NOT cache global store data in L1 because  
multiple L1 caches are not coherent for global data. Global stores bypass L1, and discard any L1  
lines that match, regardless of the cache operation. Future GPUs may have globally-coherent L1  
caches, in which case st.wb could write-back global store data from L1.”

So the L1s are invalidated, but not written back (the L2 already has the most current value for global data,  
and local data is dead after the kernel finishes).

---

_[View the full topic](https://forums.developer.nvidia.com/t/cache-data-invalidation-between-kernel-calls/26838)._
