CUDA tensor core register mapping?

I’m not sure what you mean by “down to the mma instruction itself”. You may wish to read the relevant sections of the PTX doc. Or study an example. You can find examples on these forums. Here is one.