I added results for DFlash and no thinking.
DFlash improved the bfcl speed for the dense model, but oddly took almost the same time for AgentBench (and scored a little worse, but that’s probably just normal variance since it’s supposed to be lossless). I think I need to start capturing the logs to review to understand this (or run more than 3 epochs to reduce the reduce the chance of random variance).
Disabling thinking had a big impact on times too, but also somehow scored higher on bfcl (again, I suspect some randomness here).
| name | AgentBench | bfcl |
|---|---|---|
| Qwen3.6 27B | 59.3% 2h 41m |
77.3% 1h 13m |
| Qwen3.6 27B speculative-config=dflash(15) |
58.0% 2h 42m |
77.3% 48m 58s |
| Qwen3.6 27B FP8 | 58.7% 1h 44m |
75.3% 37m 26s |
| Qwen3.6 27B enable_thinking=False |
56.0% 1h 40m |
78.0% 11m 52s |
| Qwen3.6 35B-A3B FP8 | 55.3% 2h 9m |
78.0% 17m 3s |
| Qwen3.6 35B-A3B | 52.7% 2h 34m |
78.0% 25m 5s |
| Qwen3.6 35B-A3B NVFP4 | 52.7% 2h 0m |
77.3% 18m 32s |
| Gemma4 31B | 45.3% 2h 4m |
77.3% 19m 49s |
| Qwen3 Coder Next FP8 | 46.0% 32m 49s |
|
| Gemma4 26B-A4B | 44.0% 2h 16m |
I think I’ll try to find a couple more good benchmarks to add, and then just run larger samples from them and more epochs (and just have to deal with them taking a long time) to try to get less variable numbers.