I agree that this model is not usable with this quantizattion. I went back to Intel’s original Qwen 397B int4 AutoRound, and even that is having some issues. Maybe it is about how I am serving it as well.
Happy to attempt it if it can actually be done in 12hrs. Will your instructions above be sufficient Question for you: does your tool have all these MoE “exceptions” baked in where sensible parts of the MOE architecture are left in higher precision? What about MTP layers?
MTP layers are preserved intact. It has been tested to Qwen 3.5 and 3.6 so I expect Ornith 35b to quantise but haven’t tested it yet. I have fibre to the premises being installed next week, so in future should be able to upload quantised models successfully – so don’t expect anything form me until then.