|
About the GPU-Accelerated Libraries category
|
|
0
|
5635
|
February 1, 2020
|
|
cuBLASLt returns silently wrong FP16/BF16/FP8 GEMM results on Hopper under CUDA_MPS_ACTIVE_THREAD_PERCENTAGE (algoId 66, cluster kernels)
|
|
2
|
52
|
October 6, 2026
|
|
cuDSS 0.8.0 uniform batch (UBATCH_SIZE > ~160): NaN solutions depending on the content of unrelated device memory
|
|
4
|
76
|
October 6, 2026
|
|
Sharing an Open-Source (CC0) Architecture for Chip-Level MicroChannel Liquid Cooling (MicroChannel-DirectFlow) & Exploring Collaboration
|
|
0
|
21
|
October 6, 2026
|
|
nvCOMP 5.3.0.16: nvcompBatchedZstdDecompressAsync never returns on a corrupted zstd frame
|
|
0
|
16
|
October 1, 2026
|
|
cuSolver - cusolverDnZheevjBatched API informations regarding lda and n
|
|
2
|
72
|
September 30, 2026
|
|
CUDA 13.2 Double-Precision Math Bugs on sm_80: Device Implementation Defects in jn/yn and Constant Folding Errors in llrint/lrint/copysign
|
|
5
|
82
|
September 30, 2026
|
|
Early CUDA port of our SOTA Sobol generator already beats cuRANDDx 2.3× on a Tesla T4
|
|
0
|
23
|
September 28, 2026
|
|
CUDA error while running NAMD simulations
|
|
0
|
18
|
September 24, 2026
|
|
nvCOMP 5.3: ANS FLOAT8_E4M3 single-byte decompression leaves device_statuses unchanged
|
|
0
|
27
|
September 23, 2026
|
|
Discrepancy in CUDA Fast Math Intrinsics (__sincosf, __sinf, __tanf) Between Compile-Time Constant Folding and Dynamic Runtime Execution
|
|
6
|
85
|
September 23, 2026
|
|
__viaddmax_s32 and __viaddmin_s32 return incorrect results on sm_80 when mixing constant INT_MIN with dynamic runtime arguments (CUDA 13.2)
|
|
3
|
53
|
September 23, 2026
|
|
[Bug] Inconsistent rounding results in __double2int_rn (and related intrinsics) between runtime variables and compile-time constants
|
|
2
|
79
|
September 21, 2026
|
|
Wrong FP16 results when a fused convolution, activation and residual add reads a strided concat view
|
|
0
|
33
|
September 18, 2026
|
|
Resolution problem after firmware update (UEFI device firmware)
|
|
1
|
49
|
September 16, 2026
|
|
BAR Allocation Failed & IOMMU Conflicts: Dual GPU (RTX 5060 + 4060) on Ryzen 5800X/B550 - "No Space" Errors
|
|
3
|
571
|
September 13, 2026
|
|
Discrepency in cudss ubatch solve
|
|
2
|
90
|
September 12, 2026
|
|
Surprising behaviour for cudss, uniform batch mode
|
|
1
|
84
|
September 10, 2026
|
|
Householder preprocessing before cusolverDnDsyevj
|
|
0
|
51
|
August 28, 2026
|
|
Output bit pattern mismatch between dynamic argument (arg) and kernel const for unary operator-(__half2) with QNaN payload
|
|
0
|
47
|
August 27, 2026
|
|
CST in ibgda
|
|
1
|
219
|
August 27, 2026
|
|
Call to cublasDgetrsBatched from a Fortran code fails
|
|
2
|
67
|
August 26, 2026
|
|
GPUDirect Storage performance decreases with multiple processes compared to a single process
|
|
0
|
43
|
August 20, 2026
|
|
Discrepancy between Documentation and Actual Hardware/Compiler Behavior for __vset* Intrinsics
|
|
1
|
56
|
August 18, 2026
|
|
cuFile only running compat mode for RTX 6000 pro & Samsung 9100 Pro SSD on ASUS TUF Gaming B850
|
|
7
|
345
|
August 3, 2026
|
|
Need AWS Activate Organization ID for NVIDIA Inception member - iruKa Education (Vietnam)
|
|
0
|
63
|
August 1, 2026
|
|
CUDA-enabled HPC node running OpenFOAM on Ubuntu Noble
|
|
0
|
81
|
July 30, 2026
|
|
Green Context SM Partitioning on RTX PRO 4000 Blackwell — Hard 48 SM Cap and DRAM Bandwidth Contention
|
|
0
|
150
|
July 29, 2026
|
|
Library not found on Ubuntu 24.04 install
|
|
1
|
262
|
July 28, 2026
|
|
Single process with multiple threads pure RDMA data transfer on multiple nodes
|
|
1
|
109
|
July 28, 2026
|