NVIDIA Report.docx (13.5 KB)
Hello NVIDIA Team,
We have completed a full health audit of a dual NVIDIA DGX Spark cluster and would like engineering clarification regarding a ConnectX power-related message.
Environment:
-
Dual NVIDIA DGX Spark systems
-
Dual-rail RoCE configured and operational
-
NCCL operational
-
torchrun operational
-
Distributed workloads operational
Validated Results:
-
Health Audit completed successfully
-
Phase 1 through Phase 10: PASS
-
Post-reboot validation: PASS
-
Recovery validation: PASS
-
Thermal stability: PASS
-
Power stability: PASS
-
No XID errors observed
-
No NVRM failures observed
Performance Baseline:
-
Sustained dual-rail NCCL throughput: 23.2053 GB/s
-
No measurable regression after reboot
Observed Message:
mlx5_core:
Detected insufficient power on the PCIe slot (27W)
This message is observed on both DGX Spark systems.
Observed Impact:
-
No demonstrated impact on NCCL
-
No demonstrated impact on RoCE
-
No demonstrated impact on torchrun
-
No demonstrated impact on distributed workloads
-
Systems remain fully operational
Questions:
-
Is this expected behavior on DGX Spark systems?
-
Is this a known firmware or platform reporting artifact?
-
Is the ConnectX adapter operating in any reduced-power or degraded mode?
-
Is any functionality disabled because of this condition?
-
Is any firmware, BIOS, BMC, ConnectX, or platform update recommended?
-
Can NVIDIA provide the official interpretation of this message for DGX Spark customers?
We are intentionally preserving the currently validated operational baseline and would like guidance before making any configuration or firmware changes.
Thank you.
PS, For more information and assistance, please see the attached file