# Is cache access coalesced?

**URL:** <https://forums.developer.nvidia.com/t/is-cache-access-coalesced/24698>\
**Category:** CUDA Programming and Performance\
**Created:** [October 23, 2011, 10:37am UTC](https://forums.developer.nvidia.com/t/is-cache-access-coalesced/24698 "2011-10-23T10:37:11Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![mjmawson](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@mjmawson](https://forums.developer.nvidia.com/u/mjmawson)\
**Post date:** [October 23, 2011, 10:37am UTC](https://forums.developer.nvidia.com/t/is-cache-access-coalesced/24698/1 "2011-10-23T10:37:11Z")

</div>

Can accesses to L1 cache be coalesced in the same manner that accesses to global memory are? E.g. a 128 byte segment of global memory is accessed in a coalesced fashion, the segment is cached in L1. If I later access the same 128 byte segment will access to it be coalesced in L1, or carried out serially? I couldn’t find anything in the programming guide, only information about if the segment was in global memory.

---

<div class="post-metadata">

**Author:** ![Asteroid](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@Asteroid](https://forums.developer.nvidia.com/u/Asteroid)\
**Post date:** [October 26, 2011, 9:13pm UTC](https://forums.developer.nvidia.com/t/is-cache-access-coalesced/24698/2 "2011-10-26T21:13:34Z")

</div>

The Fermi Tuning Guide states: “The same on-chip memory is used for both L1 and shared memory, …”

Shared memory accesses do not need coalescing to have optimal performance, you only have to be careful about bank conflicts. Therefore I think you don’t have to worry about memory coalescing for L1 cache as well. I’m not sure if the accesses are actually coalesced or not.

---

<div class="post-metadata">

**Author:** ![bit\_mapper](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@bit\_mapper](https://forums.developer.nvidia.com/u/bit_mapper)\
**Post date:** [October 26, 2011, 10:03pm UTC](https://forums.developer.nvidia.com/t/is-cache-access-coalesced/24698/3 "2011-10-26T22:03:51Z")

</div>

I have the same question regarding access pattern of L1 cache with you. Further, what is the access pattern of L2 cache?

---

<div class="post-metadata">

**Author:** ![AllieW](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/alliew/32/12664_2.png) [@AllieW](https://forums.developer.nvidia.com/u/AllieW)\
**Post date:** [September 5, 2016, 4:40am UTC](https://forums.developer.nvidia.com/t/is-cache-access-coalesced/24698/4 "2016-09-05T04:40:19Z")

</div>

Hi, I have a similar question here! It is reasonable for shared memory to be un-coalesced.

But what is the case for L2 cache? Does L2 cache perform coalescing if I have the L1 cache disabled? So does L2 access perform similarly as main memory access?

Thank you!

---

<div class="post-metadata">

**Author:** ![Robert\_Crovella](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/robert_crovella/32/14043_2.png) [@Robert\_Crovella](https://forums.developer.nvidia.com/u/Robert_Crovella)\
**Post date:** [September 5, 2016, 9:43am UTC](https://forums.developer.nvidia.com/t/is-cache-access-coalesced/24698/5 "2016-09-05T09:43:25Z")

</div>

coalescing is not a function of whether a particular cache is enabled or not.

You can have proper coalescing even on CC 1.x devices that had no L1 and no L2 cache.

You can determine whether or not a read or write transaction will coalesce based _strictly_ on the addresses generated for that transaction by each thread within a warp.

You may want to study carefully a presentation such as this one:

[url][http://on-demand.gputechconf.com/gtc/2012/presentations/S0514-GTC2012-GPU-Performance-Analysis.pdf[/url]](http://on-demand.gputechconf.com/gtc/2012/presentations/S0514-GTC2012-GPU-Performance-Analysis.pdf%5B/url%5D)

coalesced (or uncoalesced) access is a characteristic of global memory transactions.

with respect to shared memory, the question is whether or not a particular transaction issued by a warp instruction will have bank conflicts. The rules for determining bank conflicts have some similarities to the rules for determining coalesced access but they are not the same rules.
