|
How to report a bug
|
|
2
|
20156
|
May 27, 2024
|
|
Green-context SM provisioning vs the primary context
|
|
1
|
53
|
July 23, 2026
|
|
Weekend project: Highly efficient acos() implementations with reduced accuracy
|
|
1
|
117
|
July 22, 2026
|
|
Weekend project: Exploring the feasibility of replacing MUFU.RCP and MUFU. RSQ
|
|
5
|
321
|
July 22, 2026
|
|
Native Incremental and Temporal PageRank for Dynamic Graphs – Design Considerations
|
|
0
|
18
|
July 22, 2026
|
|
Using INT2 for Coarse Prediction Within a Stabilized FP4/FP8 Substrate (Native Hopper Implementation)
|
|
1
|
43
|
July 21, 2026
|
|
Questions for two undocumented VMM API behavior
|
|
2
|
71
|
July 17, 2026
|
|
When Activation Checkpointing Calls a Stateful Quantizer Twice
|
|
0
|
48
|
July 17, 2026
|
|
Does Blackwell support INT4 native?
|
|
13
|
1983
|
July 16, 2026
|
|
How to Calculate the Bank Conflicts
|
|
18
|
280
|
July 14, 2026
|
|
Weekend project: Very accurate double-precision sincos() implementation for a restricted domain
|
|
1
|
146
|
July 13, 2026
|
|
Question about `thread_scope_system` release/acquire store/load when `cudaDevP2PAttrNativeAtomicSupported == 0`
|
|
0
|
39
|
July 13, 2026
|
|
Slow kernel if I don't specify bounds at compile time
|
|
1
|
45
|
July 12, 2026
|
|
Sub/virtual warps or serialize access?
|
|
3
|
47
|
July 12, 2026
|
|
Code samples for GTC S81772: Don’t Leave Tensors on the Table: Programming and Optimizing Tensor Cores
|
|
0
|
41
|
July 11, 2026
|
|
CUDA 13 Fraud AI Benchmark: CPU vs CUDA 12 vs CUDA 13 model-quality comparison
|
|
3
|
131
|
July 10, 2026
|
|
Truncated loop-exit compare in PTX prevents ptxas from unrolling loops
|
|
0
|
49
|
July 7, 2026
|
|
Ldmatrix access pattern to shared memory
|
|
1
|
66
|
July 4, 2026
|
|
Cutlass or Ptx Method To Copy Data From Tensor-Memory Of SM Unit To Registers Efficiently
|
|
12
|
133
|
July 3, 2026
|
|
About NVIDIA ILC / Compute Data Compression
|
|
4
|
113
|
July 2, 2026
|
|
CUDA graph deadlocks in nested loops on Ada, Hopper and Blackwell
|
|
1
|
69
|
June 30, 2026
|
|
About Shared-Memory Bandwidth Usage Of Tensor-Core In Blackwell Architecture
|
|
12
|
218
|
June 29, 2026
|
|
Failure call of cudaIpcOpenEventHandle
|
|
1
|
44
|
June 28, 2026
|
|
Periodic 40–50 ms Kernel Launch Stall When Running Small Kernels on 5+ GPUs Concurrently
|
|
5
|
93
|
June 28, 2026
|
|
Why is cudaMallocAsync named as an asynchronous function?
|
|
4
|
79
|
June 25, 2026
|
|
Sm_90a: only one WGMMA of a 2-mma group overlaps the accumulator XOR-reduction — why?
|
|
5
|
109
|
June 24, 2026
|
|
Guidance on NVIDIA GPU compute support for a custom operating system
|
|
2
|
66
|
June 23, 2026
|
|
Epsilon in CUDA kernel
|
|
2
|
62
|
June 23, 2026
|
|
Fragment layout for mma.sync.aligned.m16n8k64 with FP4 E2M1 and block scaling on SM120
|
|
2
|
227
|
June 23, 2026
|
|
Please help me correct my understanding of H100 Tensor Core and the WGMMA instruction
|
|
1
|
69
|
June 23, 2026
|