# Is my bandwidth calculation right? bandwidth

**URL:** <https://forums.developer.nvidia.com/t/is-my-bandwidth-calculation-right-bandwidth/13214>\
**Category:** CUDA Programming and Performance\
**Created:** [November 12, 2009, 8:21am UTC](https://forums.developer.nvidia.com/t/is-my-bandwidth-calculation-right-bandwidth/13214 "2009-11-12T08:21:20Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![GoGdizzY](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@GoGdizzY](https://forums.developer.nvidia.com/u/GoGdizzY)\
**Post date:** [November 12, 2009, 8:21am UTC](https://forums.developer.nvidia.com/t/is-my-bandwidth-calculation-right-bandwidth/13214/1 "2009-11-12T08:21:20Z")

</div>

In BestPracticeGuide, it says

Effective bandwidth = (( Br + Bw ) / 10^9) / time

Theoretical Bandwidth = ( clockRate \* 10^6 \* (bitwidth/8) \* 2 ) / 10^9

so my GTX260 216sp 's theoretical bandwidth is

( 1175 \* 10^6 \* (448/8) \* 2) / 10^9 = 131.6 GB/s

In practice, my effective bandwidth is only 1.832 GB/s, is that too small?

[codebox] #define z\_uint8 unsigned char

#define z\_float32 float

#define z\_int32 int

#define N 1024

#define R 1.23456789

// global var

z\_uint8 G\_Input[N\*N];

z\_uint8 G\_Output[int(N\*R)_int(N_R)];

z\_float32 G\_Input2[N\*N];

z\_float32 G\_Output2[int(N\*R)_int(N_R)];

…

…

{

cudaEvent\_t start, stop;

float time;

cudaEventCreate(&start);

cudaEventCreate(&stop);

cudaEventRecord( start, 0 );

z\_int32 size1 = N_N, size2 = int(N_R)_int(N_R), size3 = int(N\*R);

z\_float32\* d\_dest;

z\_float32\* d\_src;

cudaMalloc( (void\*\*)&d\_src, size1 \* sizeof(z\_float32) );

cudaMalloc( (void\*\*)&d\_dest, size2 \* sizeof(z\_float32) );

for(z\_int32 i = 0; i \< size1; i++) G\_Input2[i] = z\_float32(G\_Input[i]);

z\_int32 iter = 10;

for(int i =0; i\<iter;i++)

{

cudaMemcpy( d\_src, G\_Input2, size1 \* sizeof(z\_float32), cudaMemcpyHostToDevice);

cudaMemcpy( G\_Output2, d\_dest, size2 \* sizeof(z\_float32), cudaMemcpyDeviceToHost);

}

for(z\_int32 i = 0; i \< size2; i++) G\_Output[i] = z\_uint8(G\_Output2[i]);

cudaFree( d\_dest );

cudaFree( d\_src );

cudaEventRecord( stop, 0 );

cudaEventSynchronize( stop );

cudaEventElapsedTime( &time, start, stop );

cudaEventDestroy( start );

cudaEventDestroy( stop );

printf(“Gpu time: %f miliseconds. \n”, time );

printf(“Gpu bandwidth: %f Gflops. \n”, (size1+size2)_sizeof(z\_float32)/1e6/time_iter );

}[/codebox]

---

<div class="post-metadata">

**Author:** ![GoGdizzY](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@GoGdizzY](https://forums.developer.nvidia.com/u/GoGdizzY)\
**Post date:** [November 12, 2009, 9:45am UTC](https://forums.developer.nvidia.com/t/is-my-bandwidth-calculation-right-bandwidth/13214/2 "2009-11-12T09:45:39Z")

</div>

i know that the first Bandwidth is between device memory and GPU, and the second Bandwidth is between  
Host and Device memory.

so my peak rate is 250 MB/s \* 16 = 4 GB/s (PCIE1.0 x16)

Now another question, is the GTX260 using PCIE1.0 or 2.0?

---

<div class="post-metadata">

**Author:** ![LSChien](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@LSChien](https://forums.developer.nvidia.com/u/LSChien)\
**Post date:** [November 12, 2009, 2:59pm UTC](https://forums.developer.nvidia.com/t/is-my-bandwidth-calculation-right-bandwidth/13214/3 "2009-11-12T14:59:44Z")

</div>

no, Theoretical Bandwidth = ( clockRate \* 10^6 \* (bitwidth/8) \* 2 ) / 10^9

means transfer rate among device memory.

However your code does not measure bandwidth of device memory, you measure

1. data transfer between host memory

```auto
for(z_int32 i = 0; i < size1; i++) G_Input2[i] = z_float32(G_Input[i]);

...

for(z_int32 i = 0; i < size2; i++) G_Output[i] = z_uint8(G_Output2[i]);

```

this depends on FSB, and how many cores you use. you use only one core, so

bandwidth is about 2GB/s

1. transfer between host memory and device memory

```auto
cudaMemcpy( d_src, G_Input2, size1 * sizeof(z_float32), cudaMemcpyHostToDevice);

cudaMemcpy( G_Output2, d_dest, size2 * sizeof(z_float32), cudaMemcpyDeviceToHost);

```

this depends on PCI express, roughly speaking, bandwidth is 1.7~2.5 GB/s for non-pinned memory

in my machine (ASUS P5Q PRO).

you must write a kernel function to measure bandwidth of device memory, for example, data copy

---

<div class="post-metadata">

**Author:** ![GoGdizzY](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@GoGdizzY](https://forums.developer.nvidia.com/u/GoGdizzY)\
**Post date:** [November 13, 2009, 7:16am UTC](https://forums.developer.nvidia.com/t/is-my-bandwidth-calculation-right-bandwidth/13214/4 "2009-11-13T07:16:56Z")

</div>

> [@](#):
>
> no, Theoretical Bandwidth = ( clockRate \* 10^6 \* (bitwidth/8) \* 2 ) / 10^9
> 
> means transfer rate among device memory.
> 
> However your code does not measure bandwidth of device memory, you measure
> 
> 1. data transfer between host memory
> 
> ```auto
> for(z_int32 i = 0; i < size1; i++) G_Input2[i] = z_float32(G_Input[i]);
> 
> ...
> 
> for(z_int32 i = 0; i < size2; i++) G_Output[i] = z_uint8(G_Output2[i]);
> 
> ```
> 
> this depends on FSB, and how many cores you use. you use only one core, so
> 
> bandwidth is about 2GB/s
> 
> 1. transfer between host memory and device memory
> 
> ```auto
> cudaMemcpy( d_src, G_Input2, size1 * sizeof(z_float32), cudaMemcpyHostToDevice);
> 
> cudaMemcpy( G_Output2, d_dest, size2 * sizeof(z_float32), cudaMemcpyDeviceToHost);
> 
> ```
> 
> this depends on PCI express, roughly speaking, bandwidth is 1.7~2.5 GB/s for non-pinned memory
> 
> in my machine (ASUS P5Q PRO).
> 
> you must write a kernel function to measure bandwidth of device memory, for example, data copy

Thank you very much!!

In fact, i realize it myself later. :)
