I’m trying to use tensor maps with swizzling, but any swizzle size that is smaller than the innermost tile size (i.e. 32B swizzle with a 64B tile) returns a CUDA_ERROR_INVALID_VALUE. I can’t find any requirement for this in either the PTX ISA or the cuTensorMapEncodeTiled documentation.
- Is this intentional?
- If so, is this invariant documented anywhere?
As an example of a failing configuration:
tensorDataType: CU_TENSOR_MAP_DATA_TYPE_FLOAT16
tensorRank: 3
globalDim: [6144, 6144, 1]
globalStrides: [6144, 37748736]
boxDim: [128, 32, 1]
elementStrides: [1, 1, 1]
interleave: CU_TENSOR_MAP_INTERLEAVE_NONE
swizzle: CU_TENSOR_MAP_SWIZZLE_128B
l2Promotion: CU_TENSOR_MAP_L2_PROMOTION_NONE
oobFill: CU_TENSOR_MAP_FLOAT_OOB_FILL_NONE