CUDA separable compilation + shared libraries -> "Invalid function" error

Would it be fixed? It seems that initialization code should be arranged as singleton, that is not the case currently.