# clCreateKernelsInProgram returns CL\_INVALID\_KERNEL\_DEFINITION

**URL:** <https://forums.developer.nvidia.com/t/clcreatekernelsinprogram-returns-cl-invalid-kernel-definition/17179>\
**Category:** CUDA Programming and Performance\
**Created:** [June 15, 2010, 4:26am UTC](https://forums.developer.nvidia.com/t/clcreatekernelsinprogram-returns-cl-invalid-kernel-definition/17179 "2010-06-15T04:26:29Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![stanr](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@stanr](https://forums.developer.nvidia.com/u/stanr)\
**Post date:** [June 15, 2010, 4:26am UTC](https://forums.developer.nvidia.com/t/clcreatekernelsinprogram-returns-cl-invalid-kernel-definition/17179/1 "2010-06-15T04:26:29Z")

</div>

Hi,

So I’m using CUDA 3 SDK, and in an otherwise kosher situation the call to “cl\_int status = clCreateKernelsInProgram(program, 0, NULL, &numKernels),” status is -47 (CL\_INVALID\_KERNEL\_DEFINITION).  
The program/kernel in question is reduce0 from the SDK examples, with “#define T float” on top; it compiles with no problems.

Note that if I used clCreateKernel(program, “reduce0”, &status) instead of the other call, the kernel is created correctly.

The OpenCL doc does not even specify that clCreateKernelsInProgram can return this error code.

Any thoughts on what might be happening and thoughts on how to debug this problem would be appreciated.

---

<div class="post-metadata">

**Author:** ![stanr](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@stanr](https://forums.developer.nvidia.com/u/stanr)\
**Post date:** [June 16, 2010, 6:40pm UTC](https://forums.developer.nvidia.com/t/clcreatekernelsinprogram-returns-cl-invalid-kernel-definition/17179/2 "2010-06-16T18:40:54Z")

</div>

I figured out the problem, and will share it in case someone else encounters it.

The issue was that the program was within a context created for two GPU devices using clCreateContextFromType, but compiled for only one of the devices using clBuildProgram. That later resulted in the weird error when using clCreateKernelsInProgram.

---

<div class="post-metadata">

**Author:** ![sobokhan](https://developer.download.nvidia.com/images/forums/profile-default-devtalk-84.png) [@sobokhan](https://forums.developer.nvidia.com/u/sobokhan)\
**Post date:** [July 5, 2010, 4:27pm UTC](https://forums.developer.nvidia.com/t/clcreatekernelsinprogram-returns-cl-invalid-kernel-definition/17179/3 "2010-07-05T16:27:54Z")

</div>

> [@](#):
>
> I figured out the problem, and will share it in case someone else encounters it.
> 
> The issue was that the program was within a context created for two GPU devices using clCreateContextFromType, but compiled for only one of the devices using clBuildProgram. That later resulted in the weird error when using clCreateKernelsInProgram.

I had the same issue. It’s more specific than just this is indicating. When N devices are in a context (where N \> 0) there needs to be N cl\_device\_id, N cl\_command\_queue, at least 1 program, and N\*K cl\_kernel handles (where K is the number of kernels in your .cl file). Each Kernel is compiled for a specific device. In my instance I had only K cl\_kernel handles. When I build for 1 device, it works fine, when I build against 2 devices it fails with CL\_INVALID\_KERNEL\_DEFINITION. This was happening because clGetDeviceIDs was returning 2 devices when I asked only for 1. I consider this a _BUG_ in the NVidia OpenCL code, (since this can result in a buffer overrun).

This should be able to query the maximum number of devices in the context:

```auto
// NULL platform means default implementation

// 0 size indicates no array

// NULL is acceptable as pointer

clGetDeviceIDs(NULL, CL_DEVICE_TYPE_ALL, 0, NULL, &numDevices);

```

However when I’m actively trying to get the requested number of device handles, _do not overfill_ my buffer!

```auto
err = clGetDeviceIDs(pEnv->platform, dev_types, pEnv->numDevices, pEnv->devices, &pEnv->numDevices);

```

:angry: This code causes a overrun!

This side effect flushed out the fact that my code only really supports 1 device since I don’t have N\*K cl\_kernel handles.
