I plan to post the inference commands and benchmark results for the 397B model here.
I’ve already completed a pure RTN-based W4A16G128 quantization using Intel AutoRound and am currently testing it. If everything goes well, I also plan to upload the quantized model to Hugging Face.
If you’re interested in the 397B model, feel free to follow this thread for updates!