|
What is official FP4:FP8 FLOPS ratio on Blackwell?
|
|
8
|
61
|
September 9, 2026
|
|
Cicc segfault on CUDA 12.9 with CCCL
|
|
1
|
22
|
September 9, 2026
|
|
Terminate_client under MPS: intermittently kills untouched clients, and intermittently never returns
|
|
0
|
10
|
September 9, 2026
|
|
Under MPS, reclaiming a tenant wedged in the cuDNN sm100 SDPA decode kernel silently stalls other tenants — no error, no recovery
|
|
0
|
14
|
September 9, 2026
|
|
A CTA parked on a counted barrier cannot be preempted
|
|
0
|
19
|
September 9, 2026
|
|
Thin SVD Support for Polar SVD algorithm (cusolverDnXgesvdp)
|
|
0
|
16
|
September 9, 2026
|
|
Optimized Qwen3.8 Flash-Next on 1x RTX PRO 6000: 171 tok/s, 524K, and HiCache/NIXL persistence
|
|
2
|
1282
|
September 9, 2026
|
|
Documentation or specifications on ULP precision for half (FP16) and nv_bfloat16 (BF16)?
|
|
0
|
24
|
September 8, 2026
|
|
Optimizing OpenCV camera stream latency and face detection on RTX hardware for kiosk check-ins
|
|
1
|
37
|
September 8, 2026
|
|
[580.105.08] cuMemSetAccess returns OOM near 512K aggregate VMM mappings across GPUs with free VRAM
|
|
0
|
47
|
September 7, 2026
|
|
How does the operand collector gate FFMA issue on Ampere (sm_86)?
|
|
45
|
412
|
September 5, 2026
|
|
CUDA Setup and Installation
|
|
1
|
51
|
September 2, 2026
|
|
PRPLL NTT now supports both CUDA and OpenCL
|
|
3
|
354
|
September 1, 2026
|
|
B300 SXM6: FabricManager & NVLSM Link Up / Master, but NVLink P2P Traffic Fails under Load
|
|
2
|
105
|
September 1, 2026
|
|
WSL2 / WDDM: a vectorised (≥64-bit) global load at zero free VRAM takes the Windows host down — ~250-line pure-CUDA reproducer, one flag flips it
|
|
0
|
50
|
August 31, 2026
|
|
MONOLYTH representation density, 960 GB Grace capacity, and reproduced multi-GB/s execution
|
|
1
|
73
|
August 31, 2026
|
|
[Project Share] Modular Projection Sieve: Θ(√N/log N) memory prime sieving with potential for GPU acceleration
|
|
11
|
107
|
August 31, 2026
|
|
Weekend project: Cut the maximum error in atanhf() in half without negative impact on performance
|
|
12
|
186
|
August 30, 2026
|
|
Arch_id.h: error: template constraint not satisfied (CUDA 13.3)
|
|
1
|
50
|
August 30, 2026
|
|
Experimental reversible-computation architecture with verified exact reversal: is CUDA a meaningful evaluation target?
|
|
18
|
190
|
August 30, 2026
|
|
Fedora 43 and NVCC / Cuda13.1 error "exception specification is incompatible" rsqrt / rsqrtf
|
|
10
|
2713
|
August 29, 2026
|
|
NVIDIA got R-CORE. AMD will get . And Intel gets the same R-Core Gen1 — three giants, one core. Break it
|
|
2
|
99
|
August 29, 2026
|
|
Evidence of a four-system reversible CUDA cascade with exact CPU/GPU agreement
|
|
0
|
34
|
August 28, 2026
|
|
NVIDIA L40S 'falls off the bus' (Xid 79) after llama.cpp enters an endless ///// token loop at long context — reproduces on drivers 610.57.04 and 595
|
|
1
|
125
|
August 27, 2026
|
|
Intra-warp Out-of-Order Scheduling Behavior on Hopper & Blackwell
|
|
7
|
151
|
August 27, 2026
|
|
"Limited feature set" for Minor Version Compatibility
|
|
0
|
32
|
August 26, 2026
|
|
How to declare the memory space for pointers stored in an extern __shared__ array?
|
|
2
|
69
|
August 25, 2026
|
|
MXFP6 W6A8 on RTX 5090 / SM120: Qwen3.8-27B quality, throughput, and memory trade-offs
|
|
2
|
192
|
August 24, 2026
|
|
The CUDA 13.4 nvcc compiler still cannot enable the C++23 standard on Windows
|
|
3
|
108
|
August 24, 2026
|
|
Feedback Wanted: Bare-Metal C++ Framework for Sub-Cycle (33ms) Data Center Load Remediation
|
|
0
|
52
|
August 22, 2026
|