|
RTX 4070: Linux Vulkan atomic add throughput (u32, random access) degrades above 1MiB buffer size
|
|
2
|
72
|
September 27, 2026
|
|
DGX Spark GB10: hard freeze under sustained load (RCU stall on CPU 11), watchdog + kdump both fail — working pstore-only crash capture recipe + eviden
|
|
1
|
162
|
September 24, 2026
|
|
GPU ↔ FPGA over PCIe: asynchronous P2P / GPUDirect RDMA
|
|
0
|
26
|
September 24, 2026
|
|
Deep-fold: NF4 GEMV and CUDA graphs on RTX 3080 - feedback on performance and portability
|
|
0
|
41
|
September 18, 2026
|
|
cudaMallocManaged is 13× slower than cudaMalloc on Windows — here's what I found
|
|
1
|
70
|
September 18, 2026
|
|
Qwen3.8-Flash-Next Large Context?
|
|
6
|
781
|
September 17, 2026
|
|
RTX 4060 QLoRA training: 38% higher throughput but 7°C higher temperature what measurements should I collect?
|
|
5
|
111
|
September 15, 2026
|
|
RTX PRO 4000 Blackwell: FP16/FP8/FP6/FP4 throughput with FP32 accumulation?
|
|
1
|
140
|
September 11, 2026
|
|
How should we tune Jetson AGX Thor for a CPU-bound web server workload (700 concurrent users), and can the GPU contribute anything?
|
|
1
|
92
|
September 11, 2026
|
|
Switchless 2x 4x 6x GB10 Clusters Serving GLM, DeepSeek, and Qwen
|
|
20
|
1767
|
September 7, 2026
|
|
Motif-3 315B core on one DGX Spark: 83.56 GiB, 316.7 pp / 16.5 tg, reproducible build
|
|
1
|
261
|
September 5, 2026
|
|
Qwen3.8-27B on single DGX Spark: 1.9x stock NVFP4 (2.5x FP8) at 16 concurrent, ~50 tok/s on a single request
|
|
7
|
1290
|
September 1, 2026
|
|
Qwen 3.8 27B + DFlash2
|
|
12
|
4459
|
August 27, 2026
|
|
Cold plate thermal benefit of sub-45°C coolant (down to 4°C) — junction temp / throttling / lifespan data?
|
|
0
|
56
|
August 25, 2026
|
|
Qwen3.8-27B-NVFP4 @ TP=4 running at >80 tok/s sustained & peak at 92 tok/s
|
|
8
|
1600
|
August 23, 2026
|
|
Green-context SM provisioning vs the primary context
|
|
2
|
191
|
August 14, 2026
|
|
Jetpack7.2 T5000 性能测试问题
|
|
7
|
292
|
August 4, 2026
|
|
Is there a particular order to how warps are statically assigned to SM sub partitions?
|
|
2
|
310
|
July 28, 2026
|
|
FUSA Coe on AGX Thor - unexpected results
|
|
1
|
107
|
July 27, 2026
|
|
Accelerated vs non-accelerated HSB latency
|
|
3
|
146
|
July 24, 2026
|
|
Nvfortran OpenACC performance regression: WENO5 reconstruction kernel ~1.7x slower since 24.7 (24.5 good, 24.7 first bad)
|
|
2
|
106
|
July 16, 2026
|
|
AVRCK 3.0: Reducing Inference Cost by Routing Causal Gaps Across Local and Cloud Models
|
|
0
|
96
|
July 14, 2026
|
|
Slow wi-fi on orin nano devkit (RTL8822CE 802.11ac)
|
|
11
|
369
|
June 30, 2026
|
|
GB10 really does hit ~1 PFLOP NVFP4 (2:4 sparse) — measured, with an open-source tool to reproduce it
|
|
22
|
2580
|
June 25, 2026
|
|
Flux.2 Klein 9B on DGX Spark: 2.5x Faster Inference and 59% Lower VRAM with Vitoom Nunchaku
|
|
0
|
512
|
June 25, 2026
|
|
Qwen3.5-122B-A10B on single Spark: up to 51 tok/s (v2.1 — patches + quick-start + benchmark)
|
|
434
|
29759
|
June 24, 2026
|
|
DGX Spark Performance Degradation - GPU Power Draw Issue
|
|
69
|
5788
|
June 15, 2026
|
|
Just another ASUS GX10 NCCL all_gather_perf thread... mpirun... please read if you have an ASUS model multinode setup
|
|
4
|
839
|
May 23, 2026
|
|
Lossless 7.67× LoRA / 8.35× Full FT speedup for Qwen3.5 on DGX Spark (GB10, sm_121a)
|
|
3
|
695
|
May 20, 2026
|
|
T5000 - Sample to use full performance
|
|
3
|
295
|
May 11, 2026
|