ConnectX-7 Won't Come Online

Connecting two DGX Sparks via the ConnectX-7 using a NADDOD approved Nvidia QSPF 112 cable, was shipped back to verify it was good Naddod tested said it was good but tested another one and shipped that to me. They tested on their own DGX and and no issues. These are brand new devices from different lot and different suppliers. Interesting enough I couldn’t even see the ports until I disabled the hotplug fix located in dgx-spark-mlnx-hotplug

I’ve tried swapping the ports I’ve verified the cable is seating properly on both sides.

Open to any suggestions

Connect exactly the same ports on both sparks: Spark 1 port 1 to Spark 2 port 1 or Spark 1 port 2 to Spark 2 port 2

You can use our community tool to configure the cluster and network for you:

The Spark has CX7 power savings that disables the CX7 ports when no cable is connected. After connecting the cable, you should see the ports. dgx-spark-mlnx-hotplug enables this feature.
Please follow the playbook to configure the network connection: Connect Two Sparks | DGX Spark

Nope the kernel is locking out the CX7 with mailbox errors and never bring them online

Fresh builds of the latest and latest firmware 3 different cables from FS.com Naddod and Microcenter all tested good by them just never get the CX7’s to come online same error messages on all 3 regardless of cables used the system never even pops a message about the cables being inserted because it can never securely communicate with the CX7 interfaces

[ 0.105843] platform NVDA8800:00: failed to claim resource 0: [mem 0x05170000-0x051cffff]

[ 0.105848] acpi NVDA8800:00: platform device creation failed: -16

[ 0.105919] platform NVDA8900:00: failed to claim resource 0: [mem 0xc8000000-0xd7ffffff]

[ 0.105922] acpi NVDA8900:00: platform device creation failed: -16

[ 1.120555] pci 000f:01:00.0: DOE: [2c8] failed to reset mailbox with abort command : -5

[ 1.120563] pci 000f:01:00.0: DOE: [2c8] failed to create mailbox: -5

[ 3982.500213] nvidia-modeset: WARNING: GPU:0: HDMI FRL link training failed.

[ 0.105843] platform NVDA8800:00: failed to claim resource 0: [mem 0x05170000-0x051cffff]

[ 0.105848] acpi NVDA8800:00: platform device creation failed: -16

[ 0.105919] platform NVDA8900:00: failed to claim resource 0: [mem 0xc8000000-0xd7ffffff]

[ 0.105922] acpi NVDA8900:00: platform device creation failed: -16

[ 1.120555] pci 000f:01:00.0: DOE: [2c8] failed to reset mailbox with abort command : -5

[ 1.120563] pci 000f:01:00.0: DOE: [2c8] failed to create mailbox: -5

[ 3982.500213] nvidia-modeset: WARNING: GPU:0: HDMI FRL link training failed.

I understand, in that case, please run nvidia-bug-report and send me the resulting bundle so I can have the engineering team analyze it.