Thread block clustering in Blackwell GPUs

I think that is due to differing levels of description for GB202 vs. the others. The description of the GB202 includes a die “layout” graphic as well as the note below Figure 3 that I think you are referring to. The descriptions of the other dies (GB203, etc.) in the appendices starting on page 49 don’t go to this level of depth. If they did, I think they would have similar notes about FP64 cores and TC hardware for “binary compatibility”.

It’s highly unlikely that the RTX 5090 GPU would have a SM with a maximum of 100 KB L1/shared, whereas the lesser RTX 50 series GPUs would have a SM with >200KB L1/shared, to pick one example, which would be an implication if the lesser GPUs were cc10.0.

The FP64 core count of the GB202 is mentioned below the block diagram of the GB202 in the whitepaper. There is no block diagram of any other die (neither GB203 nor GB205) shown in the whitepaper. This is probably the reason, they do not mention the FP64 core count on those dies, not because they have none or have 64.

BTW the programming guide has likely an error in chapter 16.10.1 listing the number of INT32 cores for the 12.0 with 64 instead of the doubled 128, probably a copy&paste error from 10.0.

In the whitepaper, you can calculate the L1 cache size per SM from the tables. It is 128 KB for all consumer blackwells. The programming guide says 256 KB for 10.0 and 128 KB for 12.0.