|
Multidimensional subscript operator for cuda::std::mdspan in 13.4 doesnt work on purpose?
|
|
2
|
46
|
September 12, 2026
|
|
Documentation or specifications on ULP precision for half (FP16) and nv_bfloat16 (BF16)?
|
|
6
|
78
|
September 11, 2026
|
|
RTX PRO 4000 Blackwell: FP16/FP8/FP6/FP4 throughput with FP32 accumulation?
|
|
1
|
47
|
September 11, 2026
|
|
General Question: Should Models Be Structured to Fit Under Frame Budgets when running ML work on display GPU
|
|
1
|
33
|
September 11, 2026
|
|
Nvcc 13.4 rejects conforming libstdc++ <string> with g++ 16
|
|
2
|
54
|
September 11, 2026
|
|
The tractor driver will return...this time with V-PACK
|
|
3
|
44
|
September 10, 2026
|
|
Compatible NVIDIA GPU drivers on Windows for CUDA Toolkit 13.0+
|
|
7
|
1097
|
September 10, 2026
|
|
CUDA Q error message on windows 10
|
|
3
|
307
|
September 10, 2026
|
|
Optimized Qwen3.8 Flash-Next on 1x RTX PRO 6000: 171 tok/s, 524K, and HiCache/NIXL persistence
|
|
3
|
1442
|
September 10, 2026
|
|
What is official FP4:FP8 FLOPS ratio on Blackwell?
|
|
8
|
104
|
September 9, 2026
|
|
Cicc segfault on CUDA 12.9 with CCCL
|
|
1
|
35
|
September 9, 2026
|
|
Terminate_client under MPS: intermittently kills untouched clients, and intermittently never returns
|
|
0
|
15
|
September 9, 2026
|
|
Under MPS, reclaiming a tenant wedged in the cuDNN sm100 SDPA decode kernel silently stalls other tenants — no error, no recovery
|
|
0
|
16
|
September 9, 2026
|
|
A CTA parked on a counted barrier cannot be preempted
|
|
0
|
22
|
September 9, 2026
|
|
Thin SVD Support for Polar SVD algorithm (cusolverDnXgesvdp)
|
|
0
|
22
|
September 9, 2026
|
|
Optimizing OpenCV camera stream latency and face detection on RTX hardware for kiosk check-ins
|
|
1
|
55
|
September 8, 2026
|
|
[580.105.08] cuMemSetAccess returns OOM near 512K aggregate VMM mappings across GPUs with free VRAM
|
|
0
|
56
|
September 7, 2026
|
|
How does the operand collector gate FFMA issue on Ampere (sm_86)?
|
|
45
|
430
|
September 5, 2026
|
|
CUDA Setup and Installation
|
|
1
|
62
|
September 2, 2026
|
|
PRPLL NTT now supports both CUDA and OpenCL
|
|
3
|
380
|
September 1, 2026
|
|
B300 SXM6: FabricManager & NVLSM Link Up / Master, but NVLink P2P Traffic Fails under Load
|
|
2
|
117
|
September 1, 2026
|
|
WSL2 / WDDM: a vectorised (≥64-bit) global load at zero free VRAM takes the Windows host down — ~250-line pure-CUDA reproducer, one flag flips it
|
|
0
|
60
|
August 31, 2026
|
|
MONOLYTH representation density, 960 GB Grace capacity, and reproduced multi-GB/s execution
|
|
1
|
78
|
August 31, 2026
|
|
[Project Share] Modular Projection Sieve: Θ(√N/log N) memory prime sieving with potential for GPU acceleration
|
|
11
|
117
|
August 31, 2026
|
|
Weekend project: Cut the maximum error in atanhf() in half without negative impact on performance
|
|
12
|
199
|
August 30, 2026
|
|
Arch_id.h: error: template constraint not satisfied (CUDA 13.3)
|
|
1
|
57
|
August 30, 2026
|
|
Experimental reversible-computation architecture with verified exact reversal: is CUDA a meaningful evaluation target?
|
|
18
|
198
|
August 30, 2026
|
|
Fedora 43 and NVCC / Cuda13.1 error "exception specification is incompatible" rsqrt / rsqrtf
|
|
10
|
2738
|
August 29, 2026
|
|
NVIDIA got R-CORE. AMD will get . And Intel gets the same R-Core Gen1 — three giants, one core. Break it
|
|
2
|
111
|
August 29, 2026
|
|
Evidence of a four-system reversible CUDA cascade with exact CPU/GPU agreement
|
|
0
|
40
|
August 28, 2026
|