I built OpenGauntlet.com while developing a locally hosted voice-AI system on a single DGX Spark (GB10). I directly benchmarked 31 conversational LLMs on the Spark, measuring quality, time to first token, throughput at multiple context sizes, and voice-conversation suitability.
The repository also includes the larger TTS and STT research catalogs I assembled during development. Those directories are sourced comparison research—not a claim that every speech system was benchmarked on the Spark.
I’m publishing the project because the GB10-specific results, aarch64 compatibility notes, serving-backend findings, and reproducible methodology may help other Spark owners.
Source and results: GitHub - D3velop-llc/open-gauntlet-leaderboard: OpenGauntlet — independent conversational-EQ rankings and sourced voice-AI research for locally hosted models · GitHub
I’d appreciate corrections, comparable Spark results, and suggestions for additional models to test.