|
cuBLASLt sm_120 (Blackwell): TF32 split-K nvjet kernel raises "Warp Barrier Arrival Mismatch" — intermittent illegal access / GPU hang (RTX 5090)
|
|
2
|
118
|
July 17, 2026
|
|
Nvmath-python 1.0 is Available!
|
|
0
|
25
|
July 16, 2026
|
|
cuBLAS severe underperformance on cublasSgemm for RTX 3060 Laptop GPU
|
|
1
|
70
|
May 14, 2026
|
|
cusolverDnXsyevd status 6 + XID 31 MMU fault at n=50000, FP64 real, CUDA 13.2
|
|
2
|
71
|
April 28, 2026
|
|
Machine readable specifications for compute libraries
|
|
0
|
23
|
April 22, 2026
|
|
cublasDx batched gather gemm
|
|
3
|
63
|
April 20, 2026
|
|
cublasSgemmGroupedBatched requires host-side synchronization after preceding TRSM on A5000 (device-side ordering insufficient)
|
|
1
|
33
|
April 19, 2026
|
|
cuBLAS batched FP32 SGEMM dispatcher picks suboptimal kernel on RTX 5090 (sm_120)
|
|
0
|
67
|
April 10, 2026
|
|
Results divergence between cuBLAS Sgemm and cuSPARSE BLOCKED-ELL SpMM
|
|
1
|
52
|
February 9, 2026
|
|
cuBLASDx large matrix multiplication performance
|
|
3
|
100
|
February 9, 2026
|
|
Is CublasDX compatible with per-block global-pitch or stride values in a batched-gemm kernel?
|
|
2
|
64
|
January 15, 2026
|
|
Support for per-multiplication m, n, k, lda, ldb, and ldc in batched gemm
|
|
0
|
43
|
January 3, 2026
|
|
Will Cublas support arbitrary (row-major) pitched memory for A, B, C matrices in future?
|
|
3
|
76
|
January 2, 2026
|
|
Does cublaslt batch mode for Pointer Arrays apply for scaling factors as well?
|
|
0
|
40
|
January 2, 2026
|
|
Example code of Outer Vector Scaling for FP8 data types
|
|
0
|
57
|
December 1, 2025
|
|
Pointers align requirement for api:cublasGemmBatchedEx
|
|
1
|
67
|
November 26, 2025
|
|
cuSPARSELt: Strict Output Layout Constraints for Optimal Performance in Sparse-Dense GEMM
|
|
2
|
134
|
November 21, 2025
|
|
Static CUDA Build with Opencv
|
|
5
|
168
|
November 6, 2025
|
|
Switch from "sm90_xmma_gemm.._cublas"/ "void cutlass::Kernel<cutlass_80_tensorop_.." kernels with CUDA-12.1 to "nvjet_tst..." kernels with CUDA-12.8
|
|
0
|
263
|
October 26, 2025
|
|
Exception Error cublasSgetrsBatched while cublasSgetrfBatched has no issues (cuda12.8)
|
|
0
|
71
|
September 24, 2025
|
|
Why is cuBLAS cublasDgemm slower than my naive GEMM kernel?
|
|
1
|
134
|
September 15, 2025
|
|
cublasSgemm crash with multi-thread,multi-context on t4,cublas12.4.2
|
|
0
|
56
|
September 2, 2025
|
|
Why am I 2:4 sparse slower than dense in the decode stage of LLaMA2‑7B?
|
|
0
|
115
|
August 1, 2025
|
|
cuSPARSE generic SpSM much slower than legacy csrsm2
|
|
5
|
329
|
June 30, 2025
|
|
Symmetric Matrix Inverse not correct with cusolverDnDsytri
|
|
0
|
97
|
June 30, 2025
|
|
cuDNN vs cuBLAS performance on GEMMs
|
|
0
|
150
|
June 19, 2025
|
|
Calling cublasSnrm2 inside a graph with WHILE conditional node?
|
|
0
|
60
|
June 6, 2025
|
|
Nvlink error : Undefined reference to 'cublasZgemm_v2' in ******.obj'
|
|
19
|
2316
|
May 1, 2025
|
|
How to set a fixed tile size in cublas?
|
|
1
|
156
|
April 26, 2025
|
|
Seg fault on program end when using NVSHMEM and cuBLAS
|
|
2
|
176
|
April 19, 2025
|