It is also discussed here: GB10 really does hit ~1 PFLOP NVFP4 (2:4 sparse) — measured, with an open-source tool to reproduce it A little bit.
Like you said, most of our runs are limited by LPDDR5X bandwidth rather than compute, hence why down-clocking it’s barely affecting performance. Also model-dependant, there are models that I get a 10% net loss in Tok/s at 2000mhz (i.e.: Qwen3.5-a3b-nvfp4) and others that I get margin of error results of +/- 1 tok/s (i.e.: Qwen3.5-122b-a10b-hybrid).
Also, if you see most GPUs and CPUs behavior when plotting Clock vs TDP, you’ll see there’s always a point of diminishing returns. I think the ceiling for that point on the GB10 is 2150Mhz (at least on mine), with the most optimal being 2000Mhz. Everything above that it’s like “overclocking” which is what probably nVidia did in order to hit that “Petaflop NVFP4 performance” datapoint.