Hi CUDA Team and Community,
We identified five correctness failures when testing CUDA double-precision math intrinsics on sm_80 (NVIDIA A30 / A100) using CUDA 13.2 (nvcc -O3).
These failures stem from two distinct root causes:
-
Core Device Math Algorithm Defects (
jn,yn): The underlying device math implementation contains logic bugs for boundary inputs, causing both dynamic runtime execution (arg) and compile-time evaluation (const) to return incorrect values. -
Compiler Constant-Folding / Optimization Discrepancies (
llrint,lrint,copysign): Host-side constant folding during compilation yields results that diverge from correct dynamic runtime outputs.
Environment
-
CUDA Toolkit: CUDA 13.2 (
nvcc -O3) -
Target Architecture:
sm_80(NVIDIA A30 / A100)
Detailed Analysis of Observed Failures
Category 1: Device Implementation Defects (jn, yn)
Both dynamic runtime execution (arg) and compile-time evaluation (const) fail for these functions because the bug is located directly inside the device math algorithm (e.g., misidentifying near-zero or signed zero as pole/domain branch errors).
-
yn(2, -0.0)$\rightarrow$ Returns+infinstead of-inf-
CUDA Specification Contract Violation: The CUDA Math API documentation explicitly states:
yn(n, ±0) returns -inf. -
Standard Compliance: C Standard (Annex F.10.8.5) and
glibcboth specify-inf(0xfff0000000000000). -
Actual Output: Dynamic and static calls return
+inf(0x7ff0000000000000), directly violating CUDA’s written API contract.
-
-
jn(2, +0.0 / -0.0 / 2^-1074)$\rightarrow$ Returns-qNaNinstead of+0.0-
Mathematical Definition & Standards: For order $n=2$, $J_2(0) = 0$, $J_2(-0) = J_2(+0) = +0$, and $J_2(2^{-1074}) \approx x^2/8$, which underflows to
+0. C Standard (Annex F.10.8.4) andglibcspecify+0for $n>0$. -
CUDA Specification: CUDA documentation explicitly lists all NaN-returning conditions ($x=\text{NaN}$, $n<0$, or $x=+\infty$). Inputs $x \in \{+0, -0, 2^{-1074}\}$ for $n=2$ do not match any specified NaN condition.
-
Actual Output: Dynamic and static calls return
-qNaN, indicating that zero/underflow inputs incorrectly trigger a domain/pole branch check error inside the device implementation.
-
Category 2: Compiler Constant-Folding Discrepancies (llrint, lrint, copysign)
-
llrint/lrint(x = 0x1.0000000000001p-1):-
Input value
0x3fe0000000000001ULLis slightly larger than0.5. -
Dynamic execution (
arg) correctly rounds half-up to1. -
Compile-time constant folding (
const) incorrectly truncates/folds to0.
-
-
copysign(x = qNaNorsNaN,y = -1.0):-
Dynamic execution with runtime arguments correctly sets the sign bit to negative (
0xfff8.../0xfff0...). -
Mixed execution (
xdynamic,y = -1.0literal) loses the negative sign bit and returns a positive NaN (0x7ff8.../0x7ff0...).
-
Reproducer & Details
- Godbolt Link: Compiler Explorer
Questions for NVIDIA
-
Bug Confirmation: Are these five cases confirmed as bugs in CUDA 13.2 (either inside the device math algorithm or host-side constant folding)?
-
Normative Reference Standard: In cases where host constant folding diverges from device execution, should dynamic device execution always be treated as the normative runtime baseline?