Collecting eval results for Spark-sized quants of models

I added results for DFlash and no thinking.

DFlash improved the bfcl speed for the dense model, but oddly took almost the same time for AgentBench (and scored a little worse, but that’s probably just normal variance since it’s supposed to be lossless). I think I need to start capturing the logs to review to understand this (or run more than 3 epochs to reduce the reduce the chance of random variance).

Disabling thinking had a big impact on times too, but also somehow scored higher on bfcl (again, I suspect some randomness here).

name AgentBench bfcl
Qwen3.6 27B 59.3%
2h 41m
77.3%
1h 13m
Qwen3.6 27B
speculative-config=dflash(15)
58.0%
2h 42m
77.3%
48m 58s
Qwen3.6 27B FP8 58.7%
1h 44m
75.3%
37m 26s
Qwen3.6 27B
enable_thinking=False
56.0%
1h 40m
78.0%
11m 52s
Qwen3.6 35B-A3B FP8 55.3%
2h 9m
78.0%
17m 3s
Qwen3.6 35B-A3B 52.7%
2h 34m
78.0%
25m 5s
Qwen3.6 35B-A3B NVFP4 52.7%
2h 0m
77.3%
18m 32s
Gemma4 31B 45.3%
2h 4m
77.3%
19m 49s
Qwen3 Coder Next FP8 46.0%
32m 49s
Gemma4 26B-A4B 44.0%
2h 16m

I think I’ll try to find a couple more good benchmarks to add, and then just run larger samples from them and more epochs (and just have to deal with them taking a long time) to try to get less variable numbers.