Does wmma::load_matrix_sync() or wmma::mma_sync() performance benefit from row-major or col-major layouts?

Prior to performing any matrix multiplication using tensor cores and wmma, a part of the matrix, stored in the device memory, must be loaded into a fragment using wmma::load_matrix_sync(). You must specify whether the memory is row-major or col-major for proper results. I am curious if there is any memory coalescing benefit if the A fragment is row-major and the B fragment is col-major? My gut tells me that load_matrix_sync() is designed to utilize memory coalescing regardless of your layout, and it relies on you to tell it the layout so your results are correct.

wmma are high-level convenience functions. They are mapped to mma functions.

For loading matrices from memory into registers, and more specifically for the step of loading from shared memory, the PTX instruction ldmatrix is used.

It has a normal and a .trans variant. I would expect both to have the same performance.

So the copies from global device memory can be just done 1:1 and in a coalesced fashion regardless of the layout.