DeepSeek v4 Flash (Aiden Recipe from Reddit) - 1M token session operational, Cuda 12.1 tailored for DGX Spark GB10

There was some optimizations to shard more parts of the model to reduce overhead.

I’ll work more on this once 4.1 comes out