|
GLM-5.2-Int4-Int8 on 8× GB10: ~1,200 t/s prefill, 33–54 t/s avg decode (generic - coding/structured)
|
|
27
|
1800
|
July 30, 2026
|
|
Laguna S 2.1 Config & Benchmarks
|
|
58
|
7203
|
July 29, 2026
|
|
Request to enable Public API Endpoints for my account — 404 Function not found
|
|
0
|
32
|
July 29, 2026
|
|
Nemoguard-8b-topic-control NIM returns HTTP 500 (TensorRT-LLM CUDA illegal memory access) — reproducible
|
|
0
|
41
|
July 28, 2026
|
|
Request to enable "Public API Endpoints" permission for my personal organization
|
|
0
|
20
|
July 28, 2026
|
|
Manual Account Verification Request
|
|
0
|
26
|
July 26, 2026
|
|
SIPA OS fine tuning experiment
|
|
0
|
28
|
July 26, 2026
|
|
Models That Don't Work
|
|
0
|
124
|
July 25, 2026
|
|
Running Speech-to-Speech with Qwen3-TTS on NVIDIA GB10 (DGX Spark): Bypassing GGML CUDA Crashes
|
|
2
|
256
|
July 24, 2026
|
|
Solar Open2 250B (NVFP4) on a 2× GB10 pair — FP8 KV cache A/B on a hybrid linear-attention MoE
|
|
0
|
250
|
July 24, 2026
|
|
Rate Limit Increase Request: Benchmarking Novel Grounding Framework on Small Language Models
|
|
2
|
76
|
July 24, 2026
|
|
Request to enable "Public API Endpoints" for my account — moonshotai/kimi-k2.6 returns 404 Function not found
|
|
0
|
116
|
July 24, 2026
|
|
Function not found for account — kimi-k2.6 and minimax-m3 return 404 on Free Endpoints
|
|
0
|
86
|
July 24, 2026
|
|
Public API Endpoints scope missing on personal org — Llama/Gemma work, Kimi/DeepSeek/Qwen/Nemotron all 404
|
|
0
|
50
|
July 24, 2026
|
|
Request to enable "Public API Endpoints" permission for my personal organization
|
|
0
|
25
|
July 24, 2026
|
|
Account Activation Request - No API Key Permissions
|
|
0
|
24
|
July 24, 2026
|
|
Eugr joins NVIDIA Spark Team!
|
|
109
|
4892
|
July 23, 2026
|
|
Request to enable "Public API Endpoints" permission — personal org (Belgium)
|
|
0
|
52
|
July 22, 2026
|
|
Nemotron 3 Super (nemotron-3-super-120b-a12b) returning 429 continuously for 3+ days, new API key didn't help
|
|
0
|
60
|
July 21, 2026
|
|
MiniMax-M3-AWQ on 4× GB10, fp8 KV, 262k context, adaptive reasoning, ~30 tok/s
|
|
17
|
1235
|
July 21, 2026
|
|
Three times ( VoiceClone | VoiceDesign | CustomVoice ) - Faster-Qwen3-TTS for NVIDIA DGX Spark (GB10)
|
|
56
|
2733
|
July 21, 2026
|
|
Request to enable "Public API Endpoints" for my account - kimi-k2.6 returns 404 Function not found
|
|
0
|
120
|
July 21, 2026
|
|
Jetson Agent Skills : AI-Assisted Workflows for Device & BSP Customization
|
|
1
|
236
|
July 21, 2026
|
|
Credit & rate-limit increase request (1,000 → 5,000 credits, 40 → 200 RPM)
|
|
1
|
53
|
July 20, 2026
|
|
sparkDash - Multi-unit monitoring dashboard for NVIDIA DGX Spark
|
|
0
|
322
|
July 20, 2026
|
|
Three node Spark clusters (without a switch) are now supported in spark-vllm-docker and sparkrun!
|
|
15
|
2777
|
July 19, 2026
|
|
DeepSeekv4 Flash for 1x Spark [REAP25] [PrismaAURA]
|
|
3
|
822
|
July 19, 2026
|
|
How three.ws Translates a Web App into 100+ Languages with NVIDIA NIM: an LLM-Powered i18n Pipeline
|
|
0
|
65
|
July 19, 2026
|
|
Request for NVIDIA Build API Rate Limit Increase (40 RPM → 200 RPM)
|
|
0
|
30
|
July 18, 2026
|
|
cuBLASLt sm_120 (Blackwell): TF32 split-K nvjet kernel raises "Warp Barrier Arrival Mismatch" — intermittent illegal access / GPU hang (RTX 5090)
|
|
2
|
174
|
July 17, 2026
|