# Bypassing cache in Fermi

**URL:** <https://forums.developer.nvidia.com/t/bypassing-cache-in-fermi/18233>\
**Category:** CUDA Programming and Performance\
**Created:** [August 10, 2010, 4:27pm UTC](https://forums.developer.nvidia.com/t/bypassing-cache-in-fermi/18233 "2010-08-10T16:27:11Z")\
**Posts on this page:** 17\
**Page:** 1

<div class="post-metadata">

**Author:** ![trudger](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@trudger](https://forums.developer.nvidia.com/u/trudger)\
**Post date:** [August 10, 2010, 4:27pm UTC](https://forums.developer.nvidia.com/t/bypassing-cache-in-fermi/18233/1 "2010-08-10T16:27:11Z")

</div>

In the programming guide, I saw one option for avoiding use of L1 cache in the compiler options. Is there a way to specify this feature for a certain variable(array)?

---

<div class="post-metadata">

**Author:** ![vvolkov](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@vvolkov](https://forums.developer.nvidia.com/u/vvolkov)\
**Post date:** [August 11, 2010, 7:22am UTC](https://forums.developer.nvidia.com/t/bypassing-cache-in-fermi/18233/2 "2010-08-11T07:22:18Z")

</div>

> [@](#):
>
> In the programming guide, I saw one option for avoiding use of L1 cache in the compiler options. Is there a way to specify this feature for a certain variable(array)?

Don’t know if this is available in CUDA, but you can do it in PTX. See Chapter 8.7.5.1 in ptx\_isa\_2.1.pdf.

---

<div class="post-metadata">

**Author:** ![vvolkov](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@vvolkov](https://forums.developer.nvidia.com/u/vvolkov)\
**Post date:** [August 11, 2010, 7:22am UTC](https://forums.developer.nvidia.com/t/bypassing-cache-in-fermi/18233/3 "2010-08-11T07:22:18Z")

</div>

> [@](#):
>
> In the programming guide, I saw one option for avoiding use of L1 cache in the compiler options. Is there a way to specify this feature for a certain variable(array)?

Don’t know if this is available in CUDA, but you can do it in PTX. See Chapter 8.7.5.1 in ptx\_isa\_2.1.pdf.

---

<div class="post-metadata">

**Author:** ![Sulik](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@Sulik](https://forums.developer.nvidia.com/u/Sulik)\
**Post date:** [August 11, 2010, 6:20pm UTC](https://forums.developer.nvidia.com/t/bypassing-cache-in-fermi/18233/4 "2010-08-11T18:20:11Z")

</div>

Just use the “volatile” keyword on the variable.

---

<div class="post-metadata">

**Author:** ![Sulik](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@Sulik](https://forums.developer.nvidia.com/u/Sulik)\
**Post date:** [August 11, 2010, 6:20pm UTC](https://forums.developer.nvidia.com/t/bypassing-cache-in-fermi/18233/5 "2010-08-11T18:20:11Z")

</div>

Just use the “volatile” keyword on the variable.

---

<div class="post-metadata">

**Author:** ![auhgnist](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@auhgnist](https://forums.developer.nvidia.com/u/auhgnist)\
**Post date:** [August 17, 2010, 9:07am UTC](https://forums.developer.nvidia.com/t/bypassing-cache-in-fermi/18233/6 "2010-08-17T09:07:05Z")

</div>

> [@](#):
>
> Just use the “volatile” keyword on the variable.

This does not seem to work actually, at least for CUDA 3.1. I used volatile but the actual compiled PTX code uses only ‘ld.global’ which go through cache.

---

<div class="post-metadata">

**Author:** ![auhgnist](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@auhgnist](https://forums.developer.nvidia.com/u/auhgnist)\
**Post date:** [August 17, 2010, 9:07am UTC](https://forums.developer.nvidia.com/t/bypassing-cache-in-fermi/18233/7 "2010-08-17T09:07:05Z")

</div>

> [@](#):
>
> Just use the “volatile” keyword on the variable.

This does not seem to work actually, at least for CUDA 3.1. I used volatile but the actual compiled PTX code uses only ‘ld.global’ which go through cache.

---

<div class="post-metadata">

**Author:** ![laughingrice](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@laughingrice](https://forums.developer.nvidia.com/u/laughingrice)\
**Post date:** [August 23, 2010, 1:08pm UTC](https://forums.developer.nvidia.com/t/bypassing-cache-in-fermi/18233/8 "2010-08-23T13:08:34Z")

</div>

Just wondering,

I understand why keep some of the variables out of the cache (make sure to preserve cache locality for things we know we want to stay there)  
But why would we want to disable L1 cache all together?

---

<div class="post-metadata">

**Author:** ![laughingrice](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@laughingrice](https://forums.developer.nvidia.com/u/laughingrice)\
**Post date:** [August 23, 2010, 1:08pm UTC](https://forums.developer.nvidia.com/t/bypassing-cache-in-fermi/18233/9 "2010-08-23T13:08:34Z")

</div>

Just wondering,

I understand why keep some of the variables out of the cache (make sure to preserve cache locality for things we know we want to stay there)  
But why would we want to disable L1 cache all together?

---

<div class="post-metadata">

**Author:** ![Ailleur](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@Ailleur](https://forums.developer.nvidia.com/u/Ailleur)\
**Post date:** [August 27, 2010, 1:01pm UTC](https://forums.developer.nvidia.com/t/bypassing-cache-in-fermi/18233/10 "2010-08-27T13:01:06Z")

</div>

> [@](#):
>
> Just wondering,
> 
> I understand why keep some of the variables out of the cache (make sure to preserve cache locality for things we know we want to stay there)
> 
> But why would we want to disable L1 cache all together?

My own reason for that would be to measure the effect the cache has on my execution times. But I believe his question was in fact for a given variable.

I think for specific variables the only option that has been put forward is to use inline ptx instructions.

---

<div class="post-metadata">

**Author:** ![Ailleur](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@Ailleur](https://forums.developer.nvidia.com/u/Ailleur)\
**Post date:** [August 27, 2010, 1:01pm UTC](https://forums.developer.nvidia.com/t/bypassing-cache-in-fermi/18233/11 "2010-08-27T13:01:06Z")

</div>

> [@](#):
>
> Just wondering,
> 
> I understand why keep some of the variables out of the cache (make sure to preserve cache locality for things we know we want to stay there)
> 
> But why would we want to disable L1 cache all together?

My own reason for that would be to measure the effect the cache has on my execution times. But I believe his question was in fact for a given variable.

I think for specific variables the only option that has been put forward is to use inline ptx instructions.

---

<div class="post-metadata">

**Author:** ![MisterAnderson42](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@MisterAnderson42](https://forums.developer.nvidia.com/u/MisterAnderson42)\
**Post date:** [August 27, 2010, 1:58pm UTC](https://forums.developer.nvidia.com/t/bypassing-cache-in-fermi/18233/12 "2010-08-27T13:58:09Z")

</div>

You can also ready certain values with the texture cache to avoid polluting the L1 cache.

---

<div class="post-metadata">

**Author:** ![MisterAnderson42](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@MisterAnderson42](https://forums.developer.nvidia.com/u/MisterAnderson42)\
**Post date:** [August 27, 2010, 1:58pm UTC](https://forums.developer.nvidia.com/t/bypassing-cache-in-fermi/18233/13 "2010-08-27T13:58:09Z")

</div>

You can also ready certain values with the texture cache to avoid polluting the L1 cache.

---

<div class="post-metadata">

**Author:** ![Cliff\_Woolley](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/cliff_woolley/32/14043_2.png) [@Cliff\_Woolley](https://forums.developer.nvidia.com/u/Cliff_Woolley)\
**Post date:** [August 27, 2010, 10:25pm UTC](https://forums.developer.nvidia.com/t/bypassing-cache-in-fermi/18233/14 "2010-08-27T22:25:45Z")

</div>

> [@](#):
>
> I understand why keep some of the variables out of the cache (make sure to preserve cache locality for things we know we want to stay there)
> 
> But why would we want to disable L1 cache all together?

The most common case for wanting this is in applications with incoherent memory access patterns where the L1 cache doesn’t much help anyway. Memory fetches that go through L1 always result in 128-byte transactions, but accesses that skip L1 and access L2 directly can have smaller granularities, which can help reduce over-fetch in the case of scattered access.

–Cliff

---

<div class="post-metadata">

**Author:** ![Cliff\_Woolley](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/cliff_woolley/32/14043_2.png) [@Cliff\_Woolley](https://forums.developer.nvidia.com/u/Cliff_Woolley)\
**Post date:** [August 27, 2010, 10:25pm UTC](https://forums.developer.nvidia.com/t/bypassing-cache-in-fermi/18233/15 "2010-08-27T22:25:45Z")

</div>

> [@](#):
>
> I understand why keep some of the variables out of the cache (make sure to preserve cache locality for things we know we want to stay there)
> 
> But why would we want to disable L1 cache all together?

The most common case for wanting this is in applications with incoherent memory access patterns where the L1 cache doesn’t much help anyway. Memory fetches that go through L1 always result in 128-byte transactions, but accesses that skip L1 and access L2 directly can have smaller granularities, which can help reduce over-fetch in the case of scattered access.

–Cliff

---

<div class="post-metadata">

**Author:** ![AlexanderMalishev](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@AlexanderMalishev](https://forums.developer.nvidia.com/u/AlexanderMalishev)\
**Post date:** [August 28, 2010, 6:27am UTC](https://forums.developer.nvidia.com/t/bypassing-cache-in-fermi/18233/16 "2010-08-28T06:27:55Z")

</div>

Couple of complier intrinsic would solve this problem at the CUDA C level:

template T \_\_load(T \*address , LOAD\_OPTIONS options);  
template void \_\_store(T \*address , T value, STORE\_OPTIONS options);

---

<div class="post-metadata">

**Author:** ![AlexanderMalishev](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@AlexanderMalishev](https://forums.developer.nvidia.com/u/AlexanderMalishev)\
**Post date:** [August 28, 2010, 6:27am UTC](https://forums.developer.nvidia.com/t/bypassing-cache-in-fermi/18233/17 "2010-08-28T06:27:55Z")

</div>

Couple of complier intrinsic would solve this problem at the CUDA C level:

template T \_\_load(T \*address , LOAD\_OPTIONS options);  
template void \_\_store(T \*address , T value, STORE\_OPTIONS options);
