In cutensor’s 2.0 API, there are various objects that need to be directly managed with Create/Destroy mechanisms.
In the documentation, it states that Create routines are non-blocking, but Destroy routines are blocking.
Are the various Execute routines also blocking? It is not clear from the documentation.
Do I need to insert synchronization calls (e.g. cudaDeviceSynchronize or cudaStreamSynchronize) between the Execute operation (e.g. cutensorContract) and the associated Destroy routines ( cutensorDestroyOperationDescriptor, cutensorDestroyPlan, cutensorDestroyTensorDescriptor)?
In practice, it seems like everything goes well (I get correct results) if I don’t explicitly synchronize between any cutensor calls. I do see a significant performance difference when I want to launch many calls to cutensor operations in succession, depending on whether I explicitly synchronize or not.
In our implementation, a call to a cutensor operation packages all Create/Execute/Destroy calls into a single high-level routine. When I call the high-level routine many times without explicitly synchronizing between calls (only at the very end of the set of calls), I get a significant speedup compared to the case when I synchronize inside the high-level routine. This can be expected since the GPU will wait for all the operations to complete inside the routine, which can cause some overhead. But depending on the way the Destroy routines block, this should be happening anyway.
I think a more detail explanation of the blocking mechanisms within cutensor is needed.
Thanks!