There was some optimizations to shard more parts of the model to reduce overhead.
I’ll work more on this once 4.1 comes out