|
Performance tuning the half-precision exponential function hexp()
|
|
1
|
7
|
September 29, 2026
|
|
How should errors reported through an mbarrier be associated with individual fabric operations?
|
|
1
|
16
|
September 29, 2026
|
|
A newbie with an AI reported BUG: ubuntu2604/x86_64 CUDA repo: Packages index has 5 malformed entries, apt update fails
|
|
3
|
250
|
September 29, 2026
|
|
Follow-up: lost 289th bit in the reduction mod 2^256-2^32-977 after nj_umul256wide - correct fix costs 1.2 % on sm_89, cheaper way?
|
|
2
|
50
|
September 25, 2026
|
|
GPU-PV (dxgkrnl): CUDA workload slows ~1.7× over 3–4 days of host uptime with accumulating dxgvmb_send_sync_msg: wait_for_completion failed in dmesg;
|
|
0
|
28
|
September 25, 2026
|
|
CUDA GPU not getting detected after Python environment update
|
|
0
|
15
|
September 24, 2026
|
|
Implementation of CUDA extended precision floating point operations
|
|
11
|
113
|
September 24, 2026
|
|
Linking issues using CUDA 13.2 on NVIDIA Blackwell GB100
|
|
0
|
20
|
September 24, 2026
|
|
New Version Numbers
|
|
0
|
40
|
September 24, 2026
|
|
WSL2 / WDDM: a vectorised (≥64-bit) global load at zero free VRAM takes the Windows host down — ~250-line pure-CUDA reproducer, one flag flips it
|
|
1
|
100
|
September 23, 2026
|
|
CUDA Tiling API and optimal tile shapes
|
|
2
|
39
|
September 23, 2026
|
|
[bug] cuda 13.4.59 crashed: assertion failed: invalid entity code for constraint configuration of attribute device_builtin
|
|
0
|
19
|
September 23, 2026
|
|
cudaGraphLaunch crashes or reports a launch failure for multi-million-node graphs
|
|
0
|
34
|
September 23, 2026
|
|
The tractor driver will return...this time with V-PACK
|
|
4
|
96
|
September 22, 2026
|
|
Announcement: opencl-api-cpp - A 'sister' library to cuda-api-wrappers, but for OpenCL
|
|
0
|
51
|
September 21, 2026
|
|
Prefetch.tensormap: elected lane vs full warp using CUTLASS/CuTe on SM90 and SM100
|
|
1
|
49
|
September 21, 2026
|
|
CUDA Graph updates report success but retain stale cluster dimensions on driver 580.159.03
|
|
0
|
39
|
September 18, 2026
|
|
Deep-fold: NF4 GEMV and CUDA graphs on RTX 3080 - feedback on performance and portability
|
|
0
|
38
|
September 18, 2026
|
|
cudaMallocManaged is 13× slower than cudaMalloc on Windows — here's what I found
|
|
1
|
58
|
September 18, 2026
|
|
Thin SVD Support for Polar SVD algorithm (cusolverDnXgesvdp)
|
|
4
|
103
|
September 16, 2026
|
|
Ubuntu2604 x86_64_Packages is corrupt
|
|
0
|
74
|
September 16, 2026
|
|
Documentation or specifications on ULP precision for half (FP16) and nv_bfloat16 (BF16)?
|
|
7
|
165
|
September 15, 2026
|
|
RTX 4060 QLoRA training: 38% higher throughput but 7°C higher temperature what measurements should I collect?
|
|
5
|
105
|
September 15, 2026
|
|
Pennyroyal v2.5: Qwen3.8 on RTX PRO 6000 with SGLang — online FP8, NVMe PLE and Docker
|
|
0
|
186
|
September 15, 2026
|
|
Nvcc 13.4 rejects conforming libstdc++ <string> with g++ 16
|
|
4
|
157
|
September 15, 2026
|
|
PTX‑AS chooses to emit STS / STAS instructions concentratedly within the unrolled loop body
|
|
1
|
40
|
September 15, 2026
|
|
CUDA 13.4.1 nvcc 13.4.59 rejects complete Visual Studio Build Tools 2026 with “Host compiler targets unsupported OS”
|
|
0
|
84
|
September 14, 2026
|
|
Compute-admission delay between two periodic processes
|
|
1
|
72
|
September 14, 2026
|
|
General Question: Should Models Be Structured to Fit Under Frame Budgets when running ML work on display GPU
|
|
2
|
92
|
September 13, 2026
|
|
Optimized Qwen3.8 Flash-Next on 1x RTX PRO 6000: 171 tok/s, 524K, and HiCache/NIXL persistence
|
|
4
|
2370
|
September 12, 2026
|