Is anyone else getting drastic performance difference (greater than 100GFLOPS) in CUBLAS 3.0 DGEMM between matrix dimensions that are multiple of 48 and matrix dimensions that are non-multiple of 48?
BlahCuda
1
This is true irregardless of your GPU. CUBLAS’s implementation of DGEMM’s has a “sweet spot” at this size.
Related topics
| Topic | Replies | Views | Activity | |
|---|---|---|---|---|
| What's the best matrix size for cublasSgemm performance ? | 1 | 1746 | February 16, 2017 | |
| Low CuBLAS performance | 3 | 597 | January 15, 2019 | |
| CUBLAS Configuration The use of CUBLAS for small matrix | 3 | 3842 | April 4, 2007 | |
| poor cgemm performance with cuda 3.0 | 12 | 5626 | June 2, 2010 | |
| cublas problem with very big matrixes and cublasDgemm slow | 2 | 1128 | February 23, 2017 | |
| Performance query Odd results profiling GPU speed of matrix multiplication using cublas | 1 | 1535 | February 12, 2010 | |
| cublasSgemm gives incorrect result with big matrix | 0 | 429 | June 26, 2020 | |
| cuBLAS sgemm is slow | 4 | 2636 | June 26, 2017 | |
| CGEMM problems | 14 | 6972 | February 2, 2011 | |
| CUBLAS grids and threads division | 7 | 3995 | June 18, 2018 |