# How portable are compiled binaries?

**URL:** <https://forums.developer.nvidia.com/t/how-portable-are-compiled-binaries/7309>\
**Category:** CUDA Programming and Performance\
**Created:** [January 19, 2009, 1:49pm UTC](https://forums.developer.nvidia.com/t/how-portable-are-compiled-binaries/7309 "2009-01-19T13:49:44Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![\_Big\_Mac](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@\_Big\_Mac](https://forums.developer.nvidia.com/u/_Big_Mac)\
**Post date:** [January 19, 2009, 1:49pm UTC](https://forums.developer.nvidia.com/t/how-portable-are-compiled-binaries/7309/1 "2009-01-19T13:49:44Z")

</div>

Can I write portable code that will use wither arch 1.1 or 1.3 depending on device’s compute capability using Runtime API?

Let’s assume I write a kernel that doesn’t use double precision and vote intrinsics, ie. that can be compiled to architecture 1.1 or even 1.0, but I’d like it to use the extra registers and better coalescing when ran on a 1.3 device. Does the Runtime API support this? Can the compiler automagically build two kernels, link both (on the expense of having a slightly bigger .exe file) and have the program decide at runtime which to launch?

---

<div class="post-metadata">

**Author:** ![jack](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/jack/32/12069_2.png) [@jack](https://forums.developer.nvidia.com/u/jack)\
**Post date:** [January 19, 2009, 2:44pm UTC](https://forums.developer.nvidia.com/t/how-portable-are-compiled-binaries/7309/2 "2009-01-19T14:44:16Z")

</div>

> [@](#):
>
> Can I write portable code that will use wither arch 1.1 or 1.3 depending on device’s compute capability using Runtime API?
> 
> Let’s assume I write a kernel that doesn’t use double precision and vote intrinsics, ie. that can be compiled to architecture 1.1 or even 1.0, but I’d like it to use the extra registers and better coalescing when ran on a 1.3 device. Does the Runtime API support this? Can the compiler automagically build two kernels, link both (on the expense of having a slightly bigger .exe file) and have the program decide at runtime which to launch?

I remember reading something in the driver API section once about ‘fat binaries’ – the idea was to be able to pre-compile your kernels for different compute levels (1.0, 1.1, 1.3, etc.) and put them all in one file (the ‘fat binary’). When the host program was run, it would automatically choose the most compatible kernel from the binary and run it.

I don’t know if nvidia has released any tools or whatever to use this feature as of the release of CUDA 2.1 (it said so in the programming manual), but it seems like a good feature to have, so maybe one day they will get around to it. Even if they didn’t release anything official yet, I wouldn’t think it would take someone more than an hour or so to make a little “packaging” program to import .cubins and package them all into the fat binary, and they could just post it up as one of the ‘tools’ available from the forum.

---

<div class="post-metadata">

**Author:** ![AndreiB](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/andreib/32/9727_2.png) [@AndreiB](https://forums.developer.nvidia.com/u/AndreiB)\
**Post date:** [January 19, 2009, 3:03pm UTC](https://forums.developer.nvidia.com/t/how-portable-are-compiled-binaries/7309/3 "2009-01-19T15:03:05Z")

</div>

AFAIK you don’t have to compile for sm\_13 to use softer coalescing rules and/or extra registers. This is done on hardware level, so even code compiled for sm\_10 will benefit from this hardware changes.

---

<div class="post-metadata">

**Author:** ![\_Big\_Mac](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@\_Big\_Mac](https://forums.developer.nvidia.com/u/_Big_Mac)\
**Post date:** [January 19, 2009, 3:33pm UTC](https://forums.developer.nvidia.com/t/how-portable-are-compiled-binaries/7309/4 "2009-01-19T15:33:15Z")

</div>

> [@](#):
>
> AFAIK you don’t have to compile for sm\_13 to use softer coalescing rules and/or **extra registers**. This is done on hardware level, so even code compiled for sm\_10 will benefit from this hardware changes.

Does that mean that superfluous registers spill to lmem at runtime and not at compilation? I thought you could judge how many registers have spilled by reading the .cubin.

True about coalescing though, I think I remember reading someone’s remark that the “naive” and “optimized” kernels from transpose example run equally fast on his GT 200, meaning coalescing is automatically optimized despite the example being compiled to 1.0.

---

<div class="post-metadata">

**Author:** ![E.D\_Riedijk](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@E.D\_Riedijk](https://forums.developer.nvidia.com/u/E.D_Riedijk)\
**Post date:** [January 19, 2009, 4:14pm UTC](https://forums.developer.nvidia.com/t/how-portable-are-compiled-binaries/7309/5 "2009-01-19T16:14:41Z")

</div>

> [@](#):
>
> Does that mean that superfluous registers spill to lmem at runtime and not at compilation? I thought you could judge how many registers have spilled by reading the .cubin.
> 
> True about coalescing though, I think I remember reading someone’s remark that the “naive” and “optimized” kernels from transpose example run equally fast on his GT 200, meaning coalescing is automatically optimized despite the example being compiled to 1.0.

Spilling to lmem is done at ptxas compile time, so if you are spilling to local memory it is better to recompile with the sm\_13 switch.

---

<div class="post-metadata">

**Author:** ![\_Big\_Mac](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@\_Big\_Mac](https://forums.developer.nvidia.com/u/_Big_Mac)\
**Post date:** [January 19, 2009, 4:47pm UTC](https://forums.developer.nvidia.com/t/how-portable-are-compiled-binaries/7309/6 "2009-01-19T16:47:06Z")

</div>

> [@](#):
>
> Spilling to lmem is done at ptxas compile time, so if you are spilling to local memory it is better to recompile with the sm\_13 switch.

Yes, but if I compiled with sm\_13 (16k registers) and then ran on an sm\_11 GPU, the binary would try to use non-existent registers. They wouldn’t just spill (if spilling is done @ compiling ptx) so the program would simply fail to run.

If assigning registers is done at compile time, using GT 200’s extra registers is not automatic at runtime for 1.1 builds (in contrast to coalescing). You’d need two compilations if you wanted to use extra registers and yet make the program runnable under old archs.

---

<div class="post-metadata">

**Author:** ![AndreiB](https://sea2.discourse-cdn.com/nvidia/user_avatar/forums.developer.nvidia.com/andreib/32/9727_2.png) [@AndreiB](https://forums.developer.nvidia.com/u/AndreiB)\
**Post date:** [January 19, 2009, 6:33pm UTC](https://forums.developer.nvidia.com/t/how-portable-are-compiled-binaries/7309/7 "2009-01-19T18:33:45Z")

</div>

If you compile for sm\_13 it will not run on anything earlier (or at least not supposed to run).

Physical register allocation is done at ptxas level. More registers per MP means you can have more blocks per MP for better occupancy.

---

<div class="post-metadata">

**Author:** ![E.D\_Riedijk](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@E.D\_Riedijk](https://forums.developer.nvidia.com/u/E.D_Riedijk)\
**Post date:** [January 20, 2009, 3:58am UTC](https://forums.developer.nvidia.com/t/how-portable-are-compiled-binaries/7309/8 "2009-01-20T03:58:35Z")

</div>

> [@](#):
>
> Yes, but if I compiled with sm\_13 (16k registers) and then ran on an sm\_11 GPU, the binary would try to use non-existent registers. They wouldn’t just spill (if spilling is done @ compiling ptx) so the program would simply fail to run.
> 
> If assigning registers is done at compile time, using GT 200’s extra registers is not automatic at runtime for 1.1 builds (in contrast to coalescing). You’d need two compilations if you wanted to use extra registers and yet make the program runnable under old archs.

but, but non-sm\_13 hardware is sooo 2008 ;)
