|
NVIDIA nvmath-python v1.0 is now GA
|
|
0
|
11
|
July 31, 2026
|
|
INT8 Tensor Core corruption on Turing sm_75 cuBLAS llama.cpp BPE tokenizer
|
|
3
|
71
|
July 31, 2026
|
|
Green Context SM Partitioning on RTX PRO 4000 Blackwell — Hard 48 SM Cap and DRAM Bandwidth Contention
|
|
0
|
23
|
July 29, 2026
|
|
Nvmath-python 1.0 is officially here
|
|
0
|
27
|
July 23, 2026
|
|
cuBLASLt sm_120 (Blackwell): TF32 split-K nvjet kernel raises "Warp Barrier Arrival Mismatch" — intermittent illegal access / GPU hang (RTX 5090)
|
|
2
|
156
|
July 17, 2026
|
|
Installing Pytorch with Python 3.10 for CUDA on NVIDIA TITAN BLACK
|
|
0
|
58
|
July 17, 2026
|
|
GPU Utilization Bottleneck: Single-Process with 8 Streams vs. Multi-Process with 1 Stream each (DeepStream 7.1 / PyServiceMaker)
|
|
13
|
240
|
July 17, 2026
|
|
Nvmath-python 1.0 is Available!
|
|
0
|
39
|
July 16, 2026
|
|
Deterministic Xid 120 GSP kernel panic on RTX PRO 6000 Blackwell under sustained cuBLAS workload (595.84)
|
|
1
|
111
|
July 14, 2026
|
|
Ckg-nvidia-ai — NVIDIA AI stack as a traversable MCP knowledge graph (NIM, NeMo, AgentIQ, Isaac, 20 domains)
|
|
3
|
155
|
July 9, 2026
|
|
vLLM FP8 models unusable on AGX Thor (SM 11.0): kernels compiled for sm100f only — This kernel only supports sm100f. → CUBLAS_STATUS_INTERNAL_ERROR
|
|
2
|
99
|
July 8, 2026
|
|
About Shared-Memory Bandwidth Usage Of Tensor-Core In Blackwell Architecture
|
|
12
|
245
|
June 29, 2026
|
|
Triton Inference Server Support Matrix lists incorrect PyTorch version for release 26.05
|
|
0
|
60
|
June 20, 2026
|
|
AI TOPS of Thor-U measured based on CUTLASS
|
|
0
|
30
|
June 18, 2026
|
|
V100 small-M Q4_K GEMM bottleneck: raw GGUF layout vs prepacked weight cache?
|
|
0
|
74
|
June 10, 2026
|
|
Scenario 1: Custom CUDA Kernel, cuBLAS errors, memory allocation issues, stream context corruption (gate_up_silu 3-in-1 fused Kernel cuBLAS status=13)
|
|
0
|
56
|
June 10, 2026
|
|
Llama.cpp can't work properly with docker. Multi-modal functionality fails with a CUDA internal error
|
|
8
|
479
|
June 9, 2026
|
|
Z-Image Turbo NVFP4
|
|
1
|
765
|
May 31, 2026
|
|
pyTorch Installation on Jetson Orin Nano
|
|
1
|
158
|
May 29, 2026
|
|
Request for sm_121-tuned kernels in cuDNN/cuBLAS — DGX Spark training throughput gap
|
|
4
|
286
|
May 23, 2026
|
|
One way to setup Jetson Orin Nano for OpenCV and Yolo with Cuda May 2026
|
|
1
|
258
|
May 22, 2026
|
|
Pinned memory uploads not being asynchronous on RTX 5060 Ti
|
|
6
|
146
|
May 21, 2026
|
|
cuBLAS severe underperformance on cublasSgemm for RTX 3060 Laptop GPU
|
|
1
|
73
|
May 14, 2026
|
|
Nsight system showing more memory than reality
|
|
7
|
109
|
May 13, 2026
|
|
RTX 4070 (AD104) GSP firmware crash (Xid 120 @ pc:0x1a92c96) under sustained CUDA workload — Windows BSOD + Linux GPU reset
|
|
0
|
145
|
May 11, 2026
|
|
Issues generating 64T64R testMAC vectors via cuMAC (thread-block limit & 32-bit integer overflow)
|
|
1
|
91
|
May 11, 2026
|
|
RTX Pro 6000 Backwell Card Crash
|
|
5
|
693
|
May 8, 2026
|
|
cuFFT (libcufft) crashes on H100 in Confidential Computing (CC) mode
|
|
1
|
69
|
April 29, 2026
|
|
cusolverDnXsyevd status 6 + XID 31 MMU fault at n=50000, FP64 real, CUDA 13.2
|
|
2
|
73
|
April 28, 2026
|
|
Parallel cuBLAS distributions - which one is the canonical one?
|
|
0
|
47
|
April 26, 2026
|