|
How to report a bug
|
|
2
|
20471
|
May 27, 2024
|
|
Follow-up: lost 289th bit in the reduction mod 2^256-2^32-977 after nj_umul256wide - correct fix costs 1.2 % on sm_89, cheaper way?
|
|
3
|
24
|
September 25, 2026
|
|
Implementation of CUDA extended precision floating point operations
|
|
11
|
95
|
September 24, 2026
|
|
CUDA Tiling API and optimal tile shapes
|
|
2
|
27
|
September 23, 2026
|
|
cudaGraphLaunch crashes or reports a launch failure for multi-million-node graphs
|
|
0
|
26
|
September 23, 2026
|
|
The tractor driver will return...this time with V-PACK
|
|
4
|
91
|
September 22, 2026
|
|
Announcement: opencl-api-cpp - A 'sister' library to cuda-api-wrappers, but for OpenCL
|
|
0
|
50
|
September 21, 2026
|
|
Prefetch.tensormap: elected lane vs full warp using CUTLASS/CuTe on SM90 and SM100
|
|
1
|
43
|
September 21, 2026
|
|
CUDA Graph updates report success but retain stale cluster dimensions on driver 580.159.03
|
|
0
|
36
|
September 18, 2026
|
|
Deep-fold: NF4 GEMV and CUDA graphs on RTX 3080 - feedback on performance and portability
|
|
0
|
34
|
September 18, 2026
|
|
Thin SVD Support for Polar SVD algorithm (cusolverDnXgesvdp)
|
|
4
|
94
|
September 16, 2026
|
|
Documentation or specifications on ULP precision for half (FP16) and nv_bfloat16 (BF16)?
|
|
7
|
155
|
September 15, 2026
|
|
RTX 4060 QLoRA training: 38% higher throughput but 7°C higher temperature what measurements should I collect?
|
|
5
|
85
|
September 15, 2026
|
|
Pennyroyal v2.5: Qwen3.8 on RTX PRO 6000 with SGLang — online FP8, NVMe PLE and Docker
|
|
0
|
153
|
September 15, 2026
|
|
General Question: Should Models Be Structured to Fit Under Frame Budgets when running ML work on display GPU
|
|
2
|
85
|
September 13, 2026
|
|
Optimized Qwen3.8 Flash-Next on 1x RTX PRO 6000: 171 tok/s, 524K, and HiCache/NIXL persistence
|
|
4
|
2211
|
September 12, 2026
|
|
RTX PRO 4000 Blackwell: FP16/FP8/FP6/FP4 throughput with FP32 accumulation?
|
|
1
|
126
|
September 11, 2026
|
|
What is official FP4:FP8 FLOPS ratio on Blackwell?
|
|
8
|
239
|
September 9, 2026
|
|
Terminate_client under MPS: intermittently kills untouched clients, and intermittently never returns
|
|
0
|
38
|
September 9, 2026
|
|
Under MPS, reclaiming a tenant wedged in the cuDNN sm100 SDPA decode kernel silently stalls other tenants — no error, no recovery
|
|
0
|
26
|
September 9, 2026
|
|
A CTA parked on a counted barrier cannot be preempted
|
|
0
|
39
|
September 9, 2026
|
|
Optimizing OpenCV camera stream latency and face detection on RTX hardware for kiosk check-ins
|
|
1
|
89
|
September 8, 2026
|
|
[580.105.08] cuMemSetAccess returns OOM near 512K aggregate VMM mappings across GPUs with free VRAM
|
|
0
|
82
|
September 7, 2026
|
|
How does the operand collector gate FFMA issue on Ampere (sm_86)?
|
|
45
|
509
|
September 5, 2026
|
|
PRPLL NTT now supports both CUDA and OpenCL
|
|
3
|
459
|
September 1, 2026
|
|
B300 SXM6: FabricManager & NVLSM Link Up / Master, but NVLink P2P Traffic Fails under Load
|
|
2
|
152
|
September 1, 2026
|
|
MONOLYTH representation density, 960 GB Grace capacity, and reproduced multi-GB/s execution
|
|
1
|
91
|
August 31, 2026
|
|
[Project Share] Modular Projection Sieve: Θ(√N/log N) memory prime sieving with potential for GPU acceleration
|
|
11
|
147
|
August 31, 2026
|
|
Weekend project: Cut the maximum error in atanhf() in half without negative impact on performance
|
|
12
|
231
|
August 30, 2026
|
|
Experimental reversible-computation architecture with verified exact reversal: is CUDA a meaningful evaluation target?
|
|
18
|
227
|
August 30, 2026
|