Full transparency - I was running 2bit v4 flash and was happy, same here - I’m happy. Its not perfect but I also don’t have tens of thousands of dollars to spend on a cluster than could run remotely close to a higher quant. I’m okay with the quality loss under the circumstances
I guess the question is, have you tried comparing running your work load on qwen3.8 27b or flash next on a single spark and seeing what handles your tasks with less issue?
Don’t kid yourself. What’s the point of all that speed under this quantization setup? Speed only holds value when it’s built on top of model capability.
I was able to successfully deploy it with changes suggested in PRs 3 and 4. But the max speed is 50 tok/s before PRs (it crashes after ~10ktokens) 40 with PRs applied (works reliably). Not bad still, but please dont advertise 60 tok/s.