|
Open WebUI needs sign in
|
|
2
|
22
|
October 2, 2026
|
|
Reducing LLM optimizer VRAM by >50% using frequency-domain gradients (cuFFT)
|
|
0
|
13
|
September 29, 2026
|
|
Running GLM-5.3-Flash-NVFP4 on 4xRTX PRO 6000 Blackwell
|
|
2
|
202
|
September 23, 2026
|
|
Three times ( VoiceClone | VoiceDesign | CustomVoice ) - Faster-Qwen3-TTS for NVIDIA DGX Spark (GB10)
|
|
59
|
3715
|
September 21, 2026
|
|
[THOR] Cannot run LLM - System Reboots
|
|
52
|
924
|
September 21, 2026
|
|
Your DGX Spark as a production AI server: one command, a cockpit, and an agent in the browser
|
|
2
|
471
|
September 20, 2026
|
|
How do you think models are going to progress from here out, VRAM usage and efficiency wise? Does it change your spark configuration choices?
|
|
27
|
883
|
September 18, 2026
|
|
Qwen3.8-27B at 34–38 tok/s on DGX Spark — open-source one-command setup (SGLang + NVFP4 + DSpark)
|
|
106
|
20298
|
September 16, 2026
|
|
The T5000 module shuts down when running large models with Thor and the carrier board
|
|
10
|
239
|
September 15, 2026
|
|
From Conversation to Combat: A Unified 3D AI Agent with LLM, CUDA and OpenGL on Jetson Orin Nano
|
|
2
|
132
|
September 7, 2026
|
|
Qwen3.8-27B benchmarking on one DGX Spark: DFlash2 beat vLLM+MTP, and greedy beat the "thinking" sampler
|
|
24
|
5984
|
September 7, 2026
|
|
Motif-3 315B core on one DGX Spark: 83.56 GiB, 316.7 pp / 16.5 tg, reproducible build
|
|
1
|
261
|
September 5, 2026
|
|
FP8 Qwen3.8-Flash-Next on 2x DGX Spark via SGLang: 37-40 tok/s
|
|
3
|
629
|
September 5, 2026
|
|
Local Qwen-Controlled 3D Agent with CUDA Action Loop on Jetson Orin Nano
|
|
4
|
120
|
September 3, 2026
|
|
How to run latest qwen models e.g. 3.6, 3.8 models on jetpack 6.2
|
|
1
|
117
|
September 2, 2026
|
|
GLM-5.3-Flash on 4× DGX Spark: ~30–43 tok/s, 1M context, uncensored, multimodal, cuda graphs on,
|
|
2
|
1125
|
September 1, 2026
|
|
What Matters Most for Long-Form AI Writing: Context Window, Memory, or Prompt Structure?
|
|
1
|
124
|
August 28, 2026
|
|
Qwen3.8-27B on DGX Spark using vllm: NVFP4 vs FP8 performance
|
|
8
|
5300
|
August 24, 2026
|
|
Nemotron-3-Omni with audio + video in llama.cpp on DGX Spark (GGUF, prebuilt binaries)
|
|
3
|
304
|
August 24, 2026
|
|
What are good practices for building reliable AI workflows that interact with external APIs?
|
|
2
|
99
|
August 18, 2026
|
|
Qwen3.8-27B at 256K on a 24 GB Blackwell target GPU: iMatrix NVFP4 + MTP, 55.4 tok/s
|
|
0
|
546
|
August 18, 2026
|
|
Advice requested: voice assistant on Jetson Orin Nano within a 5.4 GB memory ceiling
|
|
1
|
134
|
August 17, 2026
|
|
Running Kimi K3 on 24 DGX Spark Systems — TP24 vs TP8/PP3 Advice
|
|
39
|
3954
|
August 14, 2026
|
|
Deepseek 4 0731 Single Spark - 40t/s decode - 131k ctx - Dwarfstar CUDA custom engine
|
|
3
|
670
|
August 5, 2026
|
|
Compute bottleneck evaluating multiple trajectories through a 7B VLM (VLA Architecture Design)
|
|
4
|
172
|
August 4, 2026
|
|
One-Forward-Pass Readout Engine for TensorRT-LLM—1,014 tokens in 92.3 ms
|
|
0
|
181
|
August 3, 2026
|
|
Any opinions on OrionLLM/GRM-3.2-Sky? (70GB bf16)
|
|
4
|
458
|
August 1, 2026
|
|
Nightly QLoRA on a DGX Spark: fine-tuning Qwen3.6-35B-A3B on my own coding-agent logs (hobby project)
|
|
6
|
834
|
July 28, 2026
|
|
Benchmarking deterministic JSON execution on AGX Thor (Zero Hallucination Architecture)
|
|
1
|
91
|
July 27, 2026
|
|
Eliminating stochastic risk in multi-agent pipelines (State-Locked AES-GCM Architecture)
|
|
0
|
67
|
July 24, 2026
|