# Memory coalescing

**URL:** <https://forums.developer.nvidia.com/t/memory-coalescing/883>\
**Category:** CUDA Programming and Performance\
**Created:** [June 10, 2007, 5:25am UTC](https://forums.developer.nvidia.com/t/memory-coalescing/883 "2007-06-10T05:25:18Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![sashang](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@sashang](https://forums.developer.nvidia.com/u/sashang)\
**Post date:** [June 10, 2007, 5:25am UTC](https://forums.developer.nvidia.com/t/memory-coalescing/883/1 "2007-06-10T05:25:18Z")

</div>

Hi

I’d just like to confirm my understanding of memory coalescing. The access pattern should conform to BaseAddress + N where N is the number of the thread in the half-warp. The BaseAddress should be aligned to 16 bytes. You can test for this like this: address & 0xf == 0. An example of non-coalesced writes is in the CUDA SDK transpose project. Printing the in and out indices of the naive transpose shows the access pattern:

index\_out = 0 index\_in = 0  
index\_out = 16 index\_in = 1  
index\_out = 32 index\_in = 2  
index\_out = 48 index\_in = 3  
index\_out = 64 index\_in = 4  
index\_out = 80 index\_in = 5  
index\_out = 96 index\_in = 6  
index\_out = 112 index\_in = 7  
index\_out = 128 index\_in = 8  
etc…

The writes are obviously non-coalesced. Reads are coalesced because address locations are accessed sequentially per thread.

Does this make sense?
