There’s a low-probability PCIe error upon power-up. The connection is recognized for PCIe 3.0 and x4, but after training, an AER (Automatic Error Response) occurs, and the connection drops.
We want to rule out ORIN module power supply and PCIe CLK issues.
We’d like to ask if the PCIe UPHY0 and UPHY2 in the ORIN NANO 8G chip are powered by the same power supply?
Or are they powered by the same power supply? Our SSD uses PCIe2 without problems, but the FPGA uses PCIe0 with a low-probability, intermittent power-up error. We want to confirm if these two PCIe IPs share the same power supply within the SoC.
We’ve ruled out power fluctuations and drops causing intermittent PCIe issues. If they share the same power supply, we can assume the problem isn’t caused by the ORIN PCIe power supply.
Regarding CLK, during RC (Research and Development) work, are PCIe0 and PCIe2 from the same module source? Or are they separate channels, or is the path selected by a selector?
Or are there multiple PCIe CLKs internally distributed based on the input CLK source within the SoC chip? Similarly, rule out PCIe CLK errors;
Is it a custom carrier board or it is NVIDIA DevKit ?
Can you provide details of PCIe usage like which control and how many lanes, what are side band signal ?
Our custom-designed carrier board connects to the FPGA via PCIe 3.0 (PCIe 0 x4). Ctrl #4 x4 mode is PCIe 3.0 (PCIe 0); PCIe0_RST* is enabled; PCIe0_WAKE* is disabled; PCIe0_CLKREQ* is set to a fixed value. ASPM is off by default.
I have ruled out most possible causes and focused my investigation on the power supply and clock frequency (CLK) of the ORIN NANO 8G. I have the following questions:
-
Are the PCIe 0-3 (UPHY0, UPHY2) of the ORIN NANO SoC powered by the same power domain? Or can different power supplies be configured to power different UPHY0 and UPHY2?
-
Do the SoC’s PCIe clock frequencies come from the same power supply? How are they designed? And how are they allocated to different PCIes? Could you briefly explain? Providing this information will help me troubleshoot this low-probability PCIe boot error, as we’ve never encountered this situation before on an M.2 SSD (PCIe2 controller #7).
If the PCIe power domains are the same, I can rule out the possibility that all PCIes are affected by the power supply voltage drop, not just one of them. The same logic applies to clock frequency analysis.
- Are the PCIe 0-3 (UPHY0, UPHY2) of the ORIN NANO SoC powered by the same power domain? Or can different power supplies be configured to power different UPHY0 and UPHY2?
Yes, UPHY0 and UPHY2 are powered by the same power domain.
The FPGA is a RP (Root Port) or EP (Endpoint) ? Can you confirm when the FPGA is powered on and how long it takes for the FPGA bitstream / firmware to be fully loaded to respond if the FPGA is an Endpoint. Similarly for FPGA is RP, please make sure FPGA starts link training after Jetson Orin Nano is fully powered, initialized, and stable.
- Do the SoC’s PCIe clock frequencies come from the same power supply? How are they designed? And how are they allocated to different PCIes? Could you briefly explain? Providing this information will help me troubleshoot this low-probability PCIe boot error, as we’ve never encountered this situation before on an M.2 SSD (PCIe2 controller #7).
Please provide details of your system with clock mode.
Thank for the attached waveform. It was very helpful in confirming the power-on priority between the FPGA and the Jetson module.
To investigate this further, we need some additional confirmation:
-
Is the FPGA located on the same PCB as the Jetson module, or is it connected as a separate module via a connector?
-
How was the integrity of the FPGA verified?
-
It appears that the FPGA is currently configured to operate at PCIe Gen3 speed. Could you please test it at a slower speed, such as PCIe Gen1 or Gen2?
- Is the FPGA located on the same PCB as the Jetson module, or is it connected as a separate module via a connector?
The FPGA runs on a custom-designed carrier board, to which the ORIN expansion card is connected via a DDR4 SODIMM 260P connector.
2.How was the integrity of the FPGA verified?
Checking the FPGA’s DONE signal confirms that the FPGA has booted. The problem arises after PCIe 3.0 x4 is recognized and training is complete; a large number of AER errors begin to appear after communication. This situation is extremely unlikely to occur after a single power-on.
3.It appears that the FPGA is currently configured to operate at PCIe Gen3 speed. Could you please test it at a slower speed, such as PCIe Gen1 or Gen2?
I haven’t yet tested changing the code to PCIe Gen2 or Gen1; I will test that later.
I want to know if the ORIN NANO module’s PCIe lanes internally output a PCIe CLK. This would allow me to compare it to the other two PCIe lanes under PCIe 3.0. If they come from the same source, and the other PCIe lanes are normal, the likelihood of the ORIN NANO module having a PCIe CLK anomaly is extremely low; a CLK from the same source is unlikely to cause an anomalous output on a single lane. We have also investigated other FPGA-related hardware, including insertion loss, FPGA power supply, etc. Since the internal circuitry of the ORIN NANO is not visible, we need assistance with PCIe clock design.
- Are the PCIe 0-3 (UPHY0, UPHY2) of the ORIN NANO SoC powered by the same power domain? Or can different power supplies be configured to power different UPHY0 and UPHY2?
Yes, UPHY0 and UPHY2 are powered by the same power domain.
To avoid misunderstanding:
I’d like to understand further: Are UPHY0 and UPHY2 powered by the same power chip? Are they connected to the same power network?
For example, both are 3.3V, one 3.3V-A and the other 3.3V-B, which are in the same power domain, but can be powered and controlled separately through two different power networks and power chips.
UPHY0 and UPHY2 are powered by a common voltage regulator.
Please test with PCIe Gen1 and Gen2 to rule out a signal integrity issue and confirm whether the same issue occurs.
- Is the FPGA located on the same PCB as the Jetson module, or is it connected as a separate module via a connector?
The FPGA runs on a custom-designed carrier board, to which the ORIN expansion card is connected via a DDR4 SODIMM 260P connector.
2.How was the integrity of the FPGA verified?
Checking the FPGA’s DONE signal confirms that the FPGA has booted. The problem arises after PCIe 3.0 x4 is recognized and training is complete; a large number of AER errors begin to appear after communication. This situation is extremely unlikely to occur after a single power-on.
3.It appears that the FPGA is currently configured to operate at PCIe Gen3 speed. Could you please test it at a slower speed, such as PCIe Gen1 or Gen2?
I haven’t yet tested changing the code to PCIe Gen2 or Gen1; I will test that later.
I want to know if the ORIN NANO module’s PCIe lanes internally output a PCIe CLK. This would allow me to compare it to the other two PCIe lanes under PCIe 3.0. If they come from the same source, and the other PCIe lanes are normal, the likelihood of the ORIN NANO module having a PCIe CLK anomaly is extremely low; a CLK from the same source is unlikely to cause an anomalous output on a single lane. We have also investigated other FPGA-related hardware, including insertion loss, FPGA power supply, etc. Since the internal circuitry of the ORIN NANO is not visible, we need assistance with PCIe clock design.
To quickly answer your question, PCIe C4 and PCIe C7 use different PLLs.
SSC is enabled by default on the Jetson PCIe reference clock, so please confirm whether the FPGA endpoint supports SSC.
Also, please confirm whether the 100 MHz reference clock is supplied by the Jetson Orin Nano module in Common Clock mode.