Different focus!! FujitsuPolycom have been a huge driver and have basically “solved” che switchless setup for my sparks, but that’s not where I have been focusing! My repo is focusing only on GLM 5.3 Flash, my goal was to make it FAST without using any quantised version of the model. I wanted fast, full precision, upstream FP8 - and I have optimised a lot of the numbers for performance! For a local upstream model of GLM 5.3 Flash, I think it now runs quite fast!
Sorry for the X’ish style punchcard, but the numbers here are very nice for an FP8 and I think it will answer your question: