|
Running Nemotron 3 Ultra on 4 DGX Sparks!
|
|
12
|
647
|
June 29, 2026
|
|
Add missing Nemotron model to Inference Endpoints
|
|
0
|
41
|
June 29, 2026
|
|
Open-source recipe + scaffold: training a DSpark-class speculative-decoding draft for Nemotron
|
|
3
|
234
|
June 29, 2026
|
|
Claude Code + VLLM on nvcr.io/nvidia/vllm
|
|
1
|
217
|
June 28, 2026
|
|
How to run local Docker image of Nemotron 3 Nano 4B
|
|
2
|
61
|
June 27, 2026
|
|
Deliberations on 4-sparks cluster advantages
|
|
42
|
1944
|
June 26, 2026
|
|
Creating the NVIDIA Nemotron 3 Ultra NVFP4 Checkpoint with NVIDIA Model Optimizer
|
|
0
|
23
|
June 26, 2026
|
|
Request to enable "Public API Endpoints" permission – 403 on all models
|
|
0
|
50
|
June 26, 2026
|
|
A Spark to beat M5 Ultra and a MegaSpark to beat 2x Rubin PRO 6000!
|
|
50
|
2741
|
June 25, 2026
|
|
Now GA: NVIDIA VSS Blueprint Version 3
|
|
0
|
67
|
June 25, 2026
|
|
Atlas: Open-source inference engine for DGX Spark <2minute cold start, 100+ tok/s on Qwen3.6-35B-FP8, 13+ supported models
|
|
100
|
5879
|
June 24, 2026
|
|
Nemotron OCR v2 1.4.0 ARM64 image contains x86_64 extension on DGX Spark
|
|
0
|
30
|
June 24, 2026
|
|
Tauergon agent harness optimized for local llm
|
|
0
|
131
|
June 23, 2026
|
|
Building Local + Hybrid LLMs on DGX Spark That Outperform Top Cloud Models
|
|
25
|
7355
|
June 23, 2026
|
|
Cannot verify phone number from Pakistan (+92) - Request manual account verification for API access
|
|
0
|
13
|
June 21, 2026
|
|
Vulkan as alternative backend for llama.cpp
|
|
6
|
1612
|
June 21, 2026
|
|
Genesis — open-source multi-agent scientific discovery engine in Go (NIM + Earth-2 + BioNeMo)
|
|
0
|
24
|
June 20, 2026
|
|
Request for NVIDIA NIM API Rate Limit Increase (40 → 200 RPM)
|
|
1
|
77
|
June 18, 2026
|
|
Nemotron-3-Super-120B-A12B-NVFP4 + MTP on 4× DGX Spark via SGLang (TP=4, RoCE) - MTP actually pays off: 1.70× single-stream, accept-len ≈ 2.7
|
|
5
|
283
|
June 18, 2026
|
|
8x DGX Spark Cluster Build Report: CRS812 + 400DD→4x100G Breakouts, Nemotron 3 Ultra at TP=8
|
|
4
|
538
|
June 17, 2026
|
|
Rate limit increase request — 40 RPM to 200 RPM for Hermes Agent
|
|
1
|
115
|
June 17, 2026
|
|
ASUS Ascent GX10 — Public API Endpoints permission missing from NGC Personal Key — NIM containers returning 403
|
|
3
|
202
|
June 15, 2026
|
|
"Qwen3.6-35B-A3B-NVFP4 hangs after attention backend selection across 3 vLLM images, including NVIDIA's own official recipe
|
|
2
|
401
|
June 14, 2026
|
|
Nemotron 3 Super & Ultra Models leaking metadata and chatting in longform content
|
|
0
|
62
|
June 14, 2026
|
|
Request for NVIDIA NIM API Rate Limit Increase (40 → 200 RPM) – Student Learning & Agentic Coding Workflow
|
|
0
|
43
|
June 12, 2026
|
|
DGX Manager — an open-source control plane for your DGX Spark cluster (looking for testers & feedback)
|
|
0
|
307
|
June 12, 2026
|
|
Pushing GB10 to the Limit: Qwen3 235B MoE + Concurrent Best-of-4 + Persistent Agent Layer. Architecture check & Optimization tips?
|
|
0
|
211
|
June 12, 2026
|
|
Request for NVIDIA NIM API Rate Limit Increase (40 → 200 RPM)
|
|
0
|
27
|
June 11, 2026
|
|
NVIDIA NIM API Rate Limit Increase Request (40 → 200 RPM) – Agentic Coding Workflows
|
|
0
|
44
|
June 11, 2026
|
|
Open-webui and utilizing Nemotron VL Embed 1B
|
|
0
|
29
|
June 10, 2026
|