Meta releases Muse Glimmer, a new 30B open model that runs on 18GB RAM

Meta releases Muse Glimmer, a new 30B open model that runs on 18GB RAM.

Is it worth trying it on Spark?

The benchmark numbers look very good, definitely worth trying. I wanted to run some evals on it, but I canโ€™t get it to work. Thereโ€™s a vllm container for it, but it fails to start complaining that DFlashMuseGlimmerAssistantModel doesnโ€™t exist (Iโ€™m trying to use DFlash Spec Decode).

If anyone gets it working (with spec decode), please do post back. If itโ€™s not working by the end of the day, I might kick off the evals without spec decode to see if that works and at least get some slow numbers.

30b dense, no mtp/dspark yet.
only one mlx quant yet on hf and reports 17 t/s on m5 max 64gb, which is like 480 gb/s, so on spark expect to be 12 - 14 t/s.

Very early days. Give it few weeks but qwen 3.8 27b releases later this week so likely attention will go towards it

There is original ExecuTorch version for both Cuda and Metal, has DFlash, great numbers. But I have no idea yet how to run it on either platform. Will dispatch agent to investigate.

Muse-Glimmer-30B-GGUF:UD-Q6_K_XL + dflash

๐Ÿ”ง Tool-Call Benchmark
  Server: http://127.0.0.1:8000
  Querying http://127.0.0.1:8000/v1/models โ€ฆ โœ“ unsloth/Muse-Glimmer-30B-GGUF:UD-Q6_K_XL

  โœ“ Warm-up complete (518 ms)
  ๐Ÿ” Engine: llama.cpp b10354-d2f83055d

โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ โšก llama-benchy Throughput Benchmark โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ unsloth/Muse-Glimmer-30B-GGUF:UD-Q6_K_XL                                                                                                                 โ”‚
โ”‚ pp=[2048]  tg=[128]  depth=[0, 4096, 8192]  concurrency=[1]  runs=3  latency=generation                                                                  โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ

  โœ“ Complete โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ” 9/9 0:01:58

  llama-benchy 0.4.0
  Estimated latency: 323.9 ms

                                                                    llama-benchy Results                                                                    
โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”ณโ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”ณโ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”ณโ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”ณโ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”ณโ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”ณโ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”“
โ”ƒ Test                                 โ”ƒ     c     โ”ƒ            pp t/s โ”ƒ            tg t/s โ”ƒ           TTFT (ms) โ”ƒ         Total (ms) โ”ƒ             Tokens โ”ƒ
โ”กโ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ•‡โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ•‡โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ•‡โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ•‡โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ•‡โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ•‡โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”ฉ
โ”‚ pp2048 tg128 @ d0                    โ”‚    c1     โ”‚               673 โ”‚              44.6 โ”‚               3,408 โ”‚              5,758 โ”‚           2048+128 โ”‚
โ”‚ pp2048 tg128 @ d4096                 โ”‚    c1     โ”‚               685 โ”‚              35.0 โ”‚               8,719 โ”‚             11,906 โ”‚           2048+128 โ”‚
โ”‚ pp2048 tg128 @ d8192                 โ”‚    c1     โ”‚               682 โ”‚              26.7 โ”‚              14,174 โ”‚             18,492 โ”‚           2048+128 โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

  โ„น Metrics sourced from llama-benchy โ€” see https://github.com/eugr/llama-benchy for methodology.


โ•ญโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ ๐Ÿ”ฎ Speculative Decoding Benchmark โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฎ
โ”‚ unsloth/Muse-Glimmer-30B-GGUF:UD-Q6_K_XL                                                                                                                 โ”‚
โ”‚ tg=128  depth=[0, 4096, 8192]  prompts=['filler', 'code', 'structured']  method=auto                                                                     โ”‚
โ•ฐโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ•ฏ

  โœ“     filler @ d0  14.8 eff t/s  14.7 stream t/s  ฮฑ=38.3%  waste=62%
  โœ“       code @ d0  29.2 eff t/s  28.9 stream t/s  ฮฑ=59.3%  waste=41%
  โœ“ structured @ d0  30.5 eff t/s  30.3 stream t/s  ฮฑ=70.8%  waste=29%
  โœ“     filler @ d4096  11.2 eff t/s  11.1 stream t/s  ฮฑ=84.0%  waste=16%
  โœ“       code @ d4096  27.1 eff t/s  26.9 stream t/s  ฮฑ=54.2%  waste=46%
  โœ“ structured @ d4096  33.1 eff t/s  32.8 stream t/s  ฮฑ=70.8%  waste=29%
  โœ“     filler @ d8192  6.1 eff t/s  6.1 stream t/s  ฮฑ=25.3%  waste=75%
  โœ“       code @ d8192  27.1 eff t/s  26.9 stream t/s  ฮฑ=54.2%  waste=46%
  โœ“ structured @ d8192  33.1 eff t/s  32.9 stream t/s  ฮฑ=70.8%  waste=29%

                                  Speculative Decoding Results                                  
โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”ณโ”โ”โ”โ”โ”โ”โ”โ”ณโ”โ”โ”โ”โ”โ”โ”โ”โ”โ”ณโ”โ”โ”โ”โ”โ”โ”โ”โ”ณโ”โ”โ”โ”โ”โ”โ”โ”ณโ”โ”โ”โ”โ”โ”โ”โ”ณโ”โ”โ”โ”โ”โ”ณโ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”ณโ”โ”โ”โ”โ”โ”โ”โ”โ”โ”ณโ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”“
โ”ƒ Prompt     โ”ƒ Depth โ”ƒ Eff t/s โ”ƒ    ฮฑ % โ”ƒ Waste โ”ƒ ฯ„ len โ”ƒ Win โ”ƒ Draft t/s โ”ƒ TTFT ms โ”ƒ Total ms โ”ƒ
โ”กโ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ•‡โ”โ”โ”โ”โ”โ”โ”โ•‡โ”โ”โ”โ”โ”โ”โ”โ”โ”โ•‡โ”โ”โ”โ”โ”โ”โ”โ”โ•‡โ”โ”โ”โ”โ”โ”โ”โ•‡โ”โ”โ”โ”โ”โ”โ”โ•‡โ”โ”โ”โ”โ”โ•‡โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ•‡โ”โ”โ”โ”โ”โ”โ”โ”โ”โ•‡โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”ฉ
โ”‚ filler     โ”‚     0 โ”‚    14.8 โ”‚  38.3% โ”‚   62% โ”‚     โ€” โ”‚   โ€” โ”‚      26.7 โ”‚      34 โ”‚    8,664 โ”‚
โ”‚ code       โ”‚     0 โ”‚    29.2 โ”‚  59.3% โ”‚   41% โ”‚     โ€” โ”‚   โ€” โ”‚      38.1 โ”‚      86 โ”‚    4,475 โ”‚
โ”‚ structured โ”‚     0 โ”‚    30.5 โ”‚  70.8% โ”‚   29% โ”‚     โ€” โ”‚   โ€” โ”‚      34.4 โ”‚      13 โ”‚    4,203 โ”‚
โ”‚ filler     โ”‚    4K โ”‚    11.2 โ”‚  84.0% โ”‚   16% โ”‚     โ€” โ”‚   โ€” โ”‚      10.9 โ”‚      56 โ”‚   11,485 โ”‚
โ”‚ code       โ”‚    4K โ”‚    27.1 โ”‚  54.2% โ”‚   46% โ”‚     โ€” โ”‚   โ€” โ”‚      37.4 โ”‚     146 โ”‚    4,873 โ”‚
โ”‚ structured โ”‚    4K โ”‚    33.1 โ”‚  70.8% โ”‚   29% โ”‚     โ€” โ”‚   โ€” โ”‚      37.2 โ”‚      22 โ”‚    3,891 โ”‚
โ”‚ filler     โ”‚    8K โ”‚     6.1 โ”‚  25.3% โ”‚   75% โ”‚     โ€” โ”‚   โ€” โ”‚      14.4 โ”‚     124 โ”‚   21,007 โ”‚
โ”‚ code       โ”‚    8K โ”‚    27.1 โ”‚  54.2% โ”‚   46% โ”‚     โ€” โ”‚   โ€” โ”‚      37.4 โ”‚     160 โ”‚    4,890 โ”‚
โ”‚ structured โ”‚    8K โ”‚    33.1 โ”‚  70.8% โ”‚   29% โ”‚     โ€” โ”‚   โ€” โ”‚      37.3 โ”‚      23 โ”‚    3,887 โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

  Highest acceptance: filler (84.0%)  Lowest: filler (25.3%)

thank you, currently downloading it and will try to get it up. Howโ€™s the quality?

Since DFlash is broken, Iโ€™m running it without and itโ€™s very slow.

The first benchmark to complete is bcfl, and it scored just 10% (compared to Gemma4โ€™s 76%). I donโ€™t know if itโ€™s just too slow and timing out or if something else is broken. Iโ€™ll leave it going overnight, but probably will then just wait for the DFlash stuff to be fixed before re-running.

DFlash works fine. Iโ€™m not sure what you mean about it being broken. Iโ€™m using the same config as I posted about for the RTX 3090, essentially, just with some tweaks to use unified KV and allocate more KV space for multiple slots: Reddit

If youโ€™re getting that drastic of a score difference and you canโ€™t run DFlash, then I think something is broken with the setup youโ€™re using.

Iโ€™m seeing 13 tok/s without DFlash, 25 tok/s for prose with DFlash, and 35 tok/s for code with DFlash.

Prefill speeds are about 1000 tok/s without DFlash, and 725 tok/s with DFlash, so it currently hurts that quite a bit.

The model seems solid in my limiting testing. Not revolutionary, and probably better suited for an RTX 3090, where it runs substantially faster. Unlike Gemma 4, Glimmer is actually willing to call tools, which I appreciate.

EDIT: I guess maybe vLLM DFlash support is broken? llama.cpp support seems to be working fine. If vLLM is generally broken for this model, maybe itโ€™s not the best place to run the evals right now.

Needed to patch spark-vllm but it works with DFlash, tried the full base BF16 and the NVFP4 from Preyazz/Muse-Glimmer-30B-NVFP4, threw random stuff at it and got

BF16  + DFlash :  7.72 tok/s aggregate  (8.19 mean/turn)
NVFP4 + DFlash : 18.65 tok/s aggregate (19.66 mean/turn)   -> 2.42x

Caveats:

  • reasoning_content - Keep max_tokens โ‰ฅ 1500 or a mid-CoT cutoff blanks both fields
  • Had to add a few cherry picked patches to the spark-vllm-docker from this PR to fix DFlash

And this is in the recipe:

defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 1
  gpu_memory_utilization: 0.45    # ~54.7 GiB -> ~11.8x full-131K concurrency
  max_model_len: 131072
  max_num_seqs: 16                # DFlash-safe house rule (<=32; 256 crashes vLLM under DFlash)

command: |
  vllm serve Preyazz/Muse-Glimmer-30B-NVFP4 \
    --host {host} \
    --port {port} \
    --served-model-name muse-glimmer-nvfp4 \
    --tensor-parallel-size {tensor_parallel} \
    --gpu-memory-utilization {gpu_memory_utilization} \
    --max-model-len {max_model_len} \
    --max-num-seqs {max_num_seqs} \
    --max-num-batched-tokens 8192 \
    --enable-prefix-caching \
    --enable-chunked-prefill \
    --enable-auto-tool-choice \
    --tool-call-parser muse_glimmer \
    --reasoning-parser muse_glimmer \
    --speculative-config '{{"model":"meta-models/Muse-Glimmer-30B-assistant","num_speculative_tokens":15,"method":"dflash"}}' \
    --override-generation-config '{{"temperature":1.0,"top_p":0.95,"top_k":64}}'

Yep, it is just vllm thatโ€™s broken right now. I might try something else tomorrow evening if this isnโ€™t fixed.

I donโ€™t know if itโ€™s the cause of the poor benchmark score, but I noticed that the temperature/top_p/top_k values they give donโ€™t seem to be in the json config file, so Iโ€™m wondering if it got vllm defaults instead. I wasnโ€™t able to find how to tell what vllm is running with, so Iโ€™m restarting them with the explicit recommended values to run again overnight to see if it makes a difference.

I got a metal quant working, also poor results, not inference errors, just model (or quant). Gave it coding task in Opencode, first call it made failed, called read.read() instead of read(), despite tool list injected. Corrected of course. But looks sloppy

Seems like at least some of the issues are that bcfl expects multiple tool calls but it only calls the first one. All the other models handled this fine, so if this isnโ€™t the model, I wonder if itโ€™s related to tool call parsing or something.

Will retry using sglang and see if that does any better.

sglang seems to do the same on the first few Iโ€™ve checked, but Iโ€™ll let it finish to see if the overall score is any better (at least itโ€™s way faster because DFlash is working ๐Ÿ™‚).

Edit: It finished bfcl at 12%. I donโ€™t know all of this is down to failing to call multiple tools, but itโ€™s not a good start. Iโ€™ll let it finish IfEvalCode and BigCodeBench and see how those compare.

Interesting result. I ran a controlled A/B test on a DGX Spark (GB10) with Unsloth Muse Glimmer 30B UD-Q6_K_XL, 8K context, parallel=1, CUDA 13/sm121a, and llama.cpp built from the Muse parser PR #26849 head.

With sufficiently generous reasoning budgets, both plain Q6 and Q6+DFlash passed:

  • 20/20 forced single-tool calls
  • 5/5 tool-error recovery cases
  • 5/5 JSON cases
  • 5/5 bounded coding cases

DFlash preserved those scores while reducing the complete 40-case run from 858.5s to 210.4s โ€” a 4.08ร— end-to-end speedup. Tool decoding increased from 8.18 to 34.81 tok/s, with 35.9% of proposed draft tokens accepted. System-used unified memory was about 31 GiB, with no swap, OOM or crash.

However, a smaller token-budget profile collapsed to only 8/20 tool calls with DFlash. This model appears very sensitive to reasoning truncation and parser/template correctness.

Important caveat: my 20 tool tests force exactly one tool call per request. Therefore, they do not contradict the BFCL result showing failures on multiple simultaneous tool calls. Muse may be perfectly capable of serializing one forced call while still failing multi-tool planning or multi-call serialization. Exact-format French instruction following was also only 3/5, so I would not call it generally reliable yet.

My current hypothesis is that three separate things need to be evaluated:

  1. the correct Muse reasoning/tool parser and chat template;
  2. enough completion budget to avoid cutting the reasoning phase;
  3. genuine multi-tool planning and serialization.

Since both vLLM and SGLang appear to reproduce the multi-call issue, point 3 may indeed be a model limitation rather than only a runtime bug. A matched single-tool vs multi-tool test on the same parser would help isolate it. For now, my verdict is: very promising for local, adult, on-demand batch/coding work, especially with DFlash, but not production-ready tool calling yet.

I can concur, that Apple/Metal quants have same issue - poor mutli-tool call performance. I observed it in opencode with a single tool call too - it called supplied read() tool as read.read() - I hardly can attribute it to parser quality, rather model/quant property. Tested 4, 6, and 8 bit quants - all performed with generally same poor tool call quality (77/100 on 2.0.1. old-style tool eval bench where good numbers go 90+). Plus its very slow.

I would not waste time in my humble opinion, Qwen 3.8 27b is to drop in 24 hours and we can forget about Zukโ€™s generosity like a fever dream

Is there a token budget or a timeout in the cases? Did the failed cases return empty result?

FWIW Iโ€™m running the official bf16 model and seeing the issues I noted above (with both vllm and sglangโ€™s custom containers for this model).

I donโ€™t have full logs (seems like Inspect AI doesnโ€™t keep them when outputting json?), but in the screenshots I posted above, the task was ended when the model made a single tool call and considered a fail, as it was expected to emit multiple. I donโ€™t know what portion of the failures are a result of this, but since this benchmark is to test tool calling, I suspect that the low score is a result of bad tool calling and not anything like a timeout.

(I donโ€™t know if thereโ€™s a token budget, but Iโ€™m running these exactly like I run for ever other model - the commands are ll in my GH repo).

When I get some more time (and itโ€™s finished the others that are still running - although they seem to be going very slowly, even with DFlashโ€ฆ I am expecting bad scores) I will re-run the bfcl ones with the native output format so I can go through some failures in more detail.

(I wouldnโ€™t rule out something be wrong on InspectAIโ€™s side here, Iโ€™ve had many issues with that.. however since Iโ€™m running the exact same thing that other models have scores significantly higher in, I donโ€™t think itโ€™s to blame this time)

it is interesting though that I see many people running it on graphic cards not to complain. Reddit

Is it we are hit by some specific quantization on Sparks?

I am not sure how many of these actually use it for any practical purpose, not just trying to squeeze maximum throughput. See if any YT influencers show any practical use besides one-shot โ€œbuild a website/gameโ€ tests.

I did run a practical, repeatable test using DragonScale. It did built it, but with defects, the task wasnโ€™t finished competely - it stopped after doing a test, never did a summary and task progress, like every other model did. Some tool calls. But git done properly, levels 0 of the test game works (but then its stuck). So itโ€™s not useless, just not among the best for its class. For comparison Qwen 3.6 35B 8 bit did it much better, and 4x faster. End to end. Exactly same starting points.

Iโ€™m using the full bf16, so I donโ€™t think itโ€™s in any way related to the Spark. I just think these benchmarks are designed to thoroughly test the model (bfcl is specifically a test for tool calling, probably including many scenarios that are not often hit) and maybe these edge cases arenโ€™t covered. The issues might not be the model, but could be issues with the implementation in the inference engines, or issues with the tool parsers, chat templates, etc.

Gemma4 had a lot of issues at launch that were solved with a chat template update. I suspect as people try it out more and are able to file good bug reports, things will get much better - I donโ€™t think the results Iโ€™m seeing today are all the model is capable of.

I do wish when a new model came out, it came with a set of reproducible benchmarks - surely now that software development is solved, we can have a unified benchmark harness and have all models tested like-for-like? It would also be a useful reference for different inference engines to test against to ensure theyโ€™re getting similar results to each other.