Meta releases Muse Glimmer, a new 30B open model that runs on 18GB RAM.
Is it worth trying it on Spark?
Meta releases Muse Glimmer, a new 30B open model that runs on 18GB RAM.
Is it worth trying it on Spark?
The benchmark numbers look very good, definitely worth trying. I wanted to run some evals on it, but I canโt get it to work. Thereโs a vllm container for it, but it fails to start complaining that DFlashMuseGlimmerAssistantModel doesnโt exist (Iโm trying to use DFlash Spec Decode).
If anyone gets it working (with spec decode), please do post back. If itโs not working by the end of the day, I might kick off the evals without spec decode to see if that works and at least get some slow numbers.
30b dense, no mtp/dspark yet.
only one mlx quant yet on hf and reports 17 t/s on m5 max 64gb, which is like 480 gb/s, so on spark expect to be 12 - 14 t/s.
Very early days. Give it few weeks but qwen 3.8 27b releases later this week so likely attention will go towards it
There is original ExecuTorch version for both Cuda and Metal, has DFlash, great numbers. But I have no idea yet how to run it on either platform. Will dispatch agent to investigate.
Muse-Glimmer-30B-GGUF:UD-Q6_K_XL + dflash
๐ง Tool-Call Benchmark
Server: http://127.0.0.1:8000
Querying http://127.0.0.1:8000/v1/models โฆ โ unsloth/Muse-Glimmer-30B-GGUF:UD-Q6_K_XL
โ Warm-up complete (518 ms)
๐ Engine: llama.cpp b10354-d2f83055d
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โก llama-benchy Throughput Benchmark โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ unsloth/Muse-Glimmer-30B-GGUF:UD-Q6_K_XL โ
โ pp=[2048] tg=[128] depth=[0, 4096, 8192] concurrency=[1] runs=3 latency=generation โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
โ Complete โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ 9/9 0:01:58
llama-benchy 0.4.0
Estimated latency: 323.9 ms
llama-benchy Results
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโณโโโโโโโโโโโโณโโโโโโโโโโโโโโโโโโโโณโโโโโโโโโโโโโโโโโโโโณโโโโโโโโโโโโโโโโโโโโโโณโโโโโโโโโโโโโโโโโโโโโณโโโโโโโโโโโโโโโโโโโโโ
โ Test โ c โ pp t/s โ tg t/s โ TTFT (ms) โ Total (ms) โ Tokens โ
โกโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฉ
โ pp2048 tg128 @ d0 โ c1 โ 673 โ 44.6 โ 3,408 โ 5,758 โ 2048+128 โ
โ pp2048 tg128 @ d4096 โ c1 โ 685 โ 35.0 โ 8,719 โ 11,906 โ 2048+128 โ
โ pp2048 tg128 @ d8192 โ c1 โ 682 โ 26.7 โ 14,174 โ 18,492 โ 2048+128 โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโ
โน Metrics sourced from llama-benchy โ see https://github.com/eugr/llama-benchy for methodology.
โญโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ ๐ฎ Speculative Decoding Benchmark โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฎ
โ unsloth/Muse-Glimmer-30B-GGUF:UD-Q6_K_XL โ
โ tg=128 depth=[0, 4096, 8192] prompts=['filler', 'code', 'structured'] method=auto โ
โฐโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฏ
โ filler @ d0 14.8 eff t/s 14.7 stream t/s ฮฑ=38.3% waste=62%
โ code @ d0 29.2 eff t/s 28.9 stream t/s ฮฑ=59.3% waste=41%
โ structured @ d0 30.5 eff t/s 30.3 stream t/s ฮฑ=70.8% waste=29%
โ filler @ d4096 11.2 eff t/s 11.1 stream t/s ฮฑ=84.0% waste=16%
โ code @ d4096 27.1 eff t/s 26.9 stream t/s ฮฑ=54.2% waste=46%
โ structured @ d4096 33.1 eff t/s 32.8 stream t/s ฮฑ=70.8% waste=29%
โ filler @ d8192 6.1 eff t/s 6.1 stream t/s ฮฑ=25.3% waste=75%
โ code @ d8192 27.1 eff t/s 26.9 stream t/s ฮฑ=54.2% waste=46%
โ structured @ d8192 33.1 eff t/s 32.9 stream t/s ฮฑ=70.8% waste=29%
Speculative Decoding Results
โโโโโโโโโโโโโโณโโโโโโโโณโโโโโโโโโโณโโโโโโโโโณโโโโโโโโณโโโโโโโโณโโโโโโณโโโโโโโโโโโโณโโโโโโโโโโณโโโโโโโโโโโ
โ Prompt โ Depth โ Eff t/s โ ฮฑ % โ Waste โ ฯ len โ Win โ Draft t/s โ TTFT ms โ Total ms โ
โกโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโฉ
โ filler โ 0 โ 14.8 โ 38.3% โ 62% โ โ โ โ โ 26.7 โ 34 โ 8,664 โ
โ code โ 0 โ 29.2 โ 59.3% โ 41% โ โ โ โ โ 38.1 โ 86 โ 4,475 โ
โ structured โ 0 โ 30.5 โ 70.8% โ 29% โ โ โ โ โ 34.4 โ 13 โ 4,203 โ
โ filler โ 4K โ 11.2 โ 84.0% โ 16% โ โ โ โ โ 10.9 โ 56 โ 11,485 โ
โ code โ 4K โ 27.1 โ 54.2% โ 46% โ โ โ โ โ 37.4 โ 146 โ 4,873 โ
โ structured โ 4K โ 33.1 โ 70.8% โ 29% โ โ โ โ โ 37.2 โ 22 โ 3,891 โ
โ filler โ 8K โ 6.1 โ 25.3% โ 75% โ โ โ โ โ 14.4 โ 124 โ 21,007 โ
โ code โ 8K โ 27.1 โ 54.2% โ 46% โ โ โ โ โ 37.4 โ 160 โ 4,890 โ
โ structured โ 8K โ 33.1 โ 70.8% โ 29% โ โ โ โ โ 37.3 โ 23 โ 3,887 โ
โโโโโโโโโโโโโโดโโโโโโโโดโโโโโโโโโโดโโโโโโโโโดโโโโโโโโดโโโโโโโโดโโโโโโดโโโโโโโโโโโโดโโโโโโโโโโดโโโโโโโโโโโ
Highest acceptance: filler (84.0%) Lowest: filler (25.3%)
thank you, currently downloading it and will try to get it up. Howโs the quality?
Since DFlash is broken, Iโm running it without and itโs very slow.
The first benchmark to complete is bcfl, and it scored just 10% (compared to Gemma4โs 76%). I donโt know if itโs just too slow and timing out or if something else is broken. Iโll leave it going overnight, but probably will then just wait for the DFlash stuff to be fixed before re-running.
DFlash works fine. Iโm not sure what you mean about it being broken. Iโm using the same config as I posted about for the RTX 3090, essentially, just with some tweaks to use unified KV and allocate more KV space for multiple slots: Reddit
If youโre getting that drastic of a score difference and you canโt run DFlash, then I think something is broken with the setup youโre using.
Iโm seeing 13 tok/s without DFlash, 25 tok/s for prose with DFlash, and 35 tok/s for code with DFlash.
Prefill speeds are about 1000 tok/s without DFlash, and 725 tok/s with DFlash, so it currently hurts that quite a bit.
The model seems solid in my limiting testing. Not revolutionary, and probably better suited for an RTX 3090, where it runs substantially faster. Unlike Gemma 4, Glimmer is actually willing to call tools, which I appreciate.
EDIT: I guess maybe vLLM DFlash support is broken? llama.cpp support seems to be working fine. If vLLM is generally broken for this model, maybe itโs not the best place to run the evals right now.
Needed to patch spark-vllm but it works with DFlash, tried the full base BF16 and the NVFP4 from Preyazz/Muse-Glimmer-30B-NVFP4, threw random stuff at it and got
BF16 + DFlash : 7.72 tok/s aggregate (8.19 mean/turn)
NVFP4 + DFlash : 18.65 tok/s aggregate (19.66 mean/turn) -> 2.42x
Caveats:
reasoning_content - Keep max_tokens โฅ 1500 or a mid-CoT cutoff blanks both fieldsAnd this is in the recipe:
defaults:
port: 8000
host: 0.0.0.0
tensor_parallel: 1
gpu_memory_utilization: 0.45 # ~54.7 GiB -> ~11.8x full-131K concurrency
max_model_len: 131072
max_num_seqs: 16 # DFlash-safe house rule (<=32; 256 crashes vLLM under DFlash)
command: |
vllm serve Preyazz/Muse-Glimmer-30B-NVFP4 \
--host {host} \
--port {port} \
--served-model-name muse-glimmer-nvfp4 \
--tensor-parallel-size {tensor_parallel} \
--gpu-memory-utilization {gpu_memory_utilization} \
--max-model-len {max_model_len} \
--max-num-seqs {max_num_seqs} \
--max-num-batched-tokens 8192 \
--enable-prefix-caching \
--enable-chunked-prefill \
--enable-auto-tool-choice \
--tool-call-parser muse_glimmer \
--reasoning-parser muse_glimmer \
--speculative-config '{{"model":"meta-models/Muse-Glimmer-30B-assistant","num_speculative_tokens":15,"method":"dflash"}}' \
--override-generation-config '{{"temperature":1.0,"top_p":0.95,"top_k":64}}'
Yep, it is just vllm thatโs broken right now. I might try something else tomorrow evening if this isnโt fixed.
I donโt know if itโs the cause of the poor benchmark score, but I noticed that the temperature/top_p/top_k values they give donโt seem to be in the json config file, so Iโm wondering if it got vllm defaults instead. I wasnโt able to find how to tell what vllm is running with, so Iโm restarting them with the explicit recommended values to run again overnight to see if it makes a difference.
I got a metal quant working, also poor results, not inference errors, just model (or quant). Gave it coding task in Opencode, first call it made failed, called read.read() instead of read(), despite tool list injected. Corrected of course. But looks sloppy
Seems like at least some of the issues are that bcfl expects multiple tool calls but it only calls the first one. All the other models handled this fine, so if this isnโt the model, I wonder if itโs related to tool call parsing or something.
Will retry using sglang and see if that does any better.
sglang seems to do the same on the first few Iโve checked, but Iโll let it finish to see if the overall score is any better (at least itโs way faster because DFlash is working ๐).
Edit: It finished bfcl at 12%. I donโt know all of this is down to failing to call multiple tools, but itโs not a good start. Iโll let it finish IfEvalCode and BigCodeBench and see how those compare.
Interesting result. I ran a controlled A/B test on a DGX Spark (GB10) with Unsloth Muse Glimmer 30B UD-Q6_K_XL, 8K context, parallel=1, CUDA 13/sm121a, and llama.cpp built from the Muse parser PR #26849 head.
With sufficiently generous reasoning budgets, both plain Q6 and Q6+DFlash passed:
DFlash preserved those scores while reducing the complete 40-case run from 858.5s to 210.4s โ a 4.08ร end-to-end speedup. Tool decoding increased from 8.18 to 34.81 tok/s, with 35.9% of proposed draft tokens accepted. System-used unified memory was about 31 GiB, with no swap, OOM or crash.
However, a smaller token-budget profile collapsed to only 8/20 tool calls with DFlash. This model appears very sensitive to reasoning truncation and parser/template correctness.
Important caveat: my 20 tool tests force exactly one tool call per request. Therefore, they do not contradict the BFCL result showing failures on multiple simultaneous tool calls. Muse may be perfectly capable of serializing one forced call while still failing multi-tool planning or multi-call serialization. Exact-format French instruction following was also only 3/5, so I would not call it generally reliable yet.
My current hypothesis is that three separate things need to be evaluated:
Since both vLLM and SGLang appear to reproduce the multi-call issue, point 3 may indeed be a model limitation rather than only a runtime bug. A matched single-tool vs multi-tool test on the same parser would help isolate it. For now, my verdict is: very promising for local, adult, on-demand batch/coding work, especially with DFlash, but not production-ready tool calling yet.
I can concur, that Apple/Metal quants have same issue - poor mutli-tool call performance. I observed it in opencode with a single tool call too - it called supplied read() tool as read.read() - I hardly can attribute it to parser quality, rather model/quant property. Tested 4, 6, and 8 bit quants - all performed with generally same poor tool call quality (77/100 on 2.0.1. old-style tool eval bench where good numbers go 90+). Plus its very slow.
I would not waste time in my humble opinion, Qwen 3.8 27b is to drop in 24 hours and we can forget about Zukโs generosity like a fever dream
Is there a token budget or a timeout in the cases? Did the failed cases return empty result?
FWIW Iโm running the official bf16 model and seeing the issues I noted above (with both vllm and sglangโs custom containers for this model).
I donโt have full logs (seems like Inspect AI doesnโt keep them when outputting json?), but in the screenshots I posted above, the task was ended when the model made a single tool call and considered a fail, as it was expected to emit multiple. I donโt know what portion of the failures are a result of this, but since this benchmark is to test tool calling, I suspect that the low score is a result of bad tool calling and not anything like a timeout.
(I donโt know if thereโs a token budget, but Iโm running these exactly like I run for ever other model - the commands are ll in my GH repo).
When I get some more time (and itโs finished the others that are still running - although they seem to be going very slowly, even with DFlashโฆ I am expecting bad scores) I will re-run the bfcl ones with the native output format so I can go through some failures in more detail.
(I wouldnโt rule out something be wrong on InspectAIโs side here, Iโve had many issues with that.. however since Iโm running the exact same thing that other models have scores significantly higher in, I donโt think itโs to blame this time)
it is interesting though that I see many people running it on graphic cards not to complain. Reddit
Is it we are hit by some specific quantization on Sparks?
I am not sure how many of these actually use it for any practical purpose, not just trying to squeeze maximum throughput. See if any YT influencers show any practical use besides one-shot โbuild a website/gameโ tests.
I did run a practical, repeatable test using DragonScale. It did built it, but with defects, the task wasnโt finished competely - it stopped after doing a test, never did a summary and task progress, like every other model did. Some tool calls. But git done properly, levels 0 of the test game works (but then its stuck). So itโs not useless, just not among the best for its class. For comparison Qwen 3.6 35B 8 bit did it much better, and 4x faster. End to end. Exactly same starting points.
Iโm using the full bf16, so I donโt think itโs in any way related to the Spark. I just think these benchmarks are designed to thoroughly test the model (bfcl is specifically a test for tool calling, probably including many scenarios that are not often hit) and maybe these edge cases arenโt covered. The issues might not be the model, but could be issues with the implementation in the inference engines, or issues with the tool parsers, chat templates, etc.
Gemma4 had a lot of issues at launch that were solved with a chat template update. I suspect as people try it out more and are able to file good bug reports, things will get much better - I donโt think the results Iโm seeing today are all the model is capable of.
I do wish when a new model came out, it came with a set of reproducible benchmarks - surely now that software development is solved, we can have a unified benchmark harness and have all models tested like-for-like? It would also be a useful reference for different inference engines to test against to ensure theyโre getting similar results to each other.