NVSwitch attestation from inside a TDX guest: nvattest 1.2.0 fails with NSCQ_RC_WARNING: RDT init failure (Code 1)

Setup

  • HGX node with 8x H100 80GB HBM3, driver 595.71.05, tenant VM is an Intel TDX guest (Ubuntu 24.04 image, kernel 6.8.0-110-generic, /dev/tdx_guest present).
  • GPUs run in multi-GPU Protected PCIe mode (nvidia-smi conf-compute reads CC State OFF, Multi-GPU Mode Protected PCIe, GPUs Ready State Ready).
  • GPU attestation works from inside the guest with nv-attestation-sdk in REMOTE mode with get_evidence(options={“ppcie_mode”: False}): NRAS returns “Attestation Successful” for all 8 GPUs (token 17522 bytes). That part is fine.

What fails

nvattest attest --device nvswitch from inside the same guest. libnvidia-nscq is absent from the image; after installing libnvidia-nscq 615.71.09 the failure moves one layer down:

[switch/nscq_client.cpp:195] Failed to create NSCQ session: NSCQ_RC_WARNING: RDT init failure (Code 1)
[switch/evidence.cpp:178] Failed to initialize NSCQ
Error 600: NSCQ Initialization Failed

The four switches are visible from the guest: [10de:22a3] on the bus, /dev/nvidia-nvswitch0 to /dev/nvidia-nvswitch3 present. The Python SDK, when asked for switch evidence on 2026-09-10, came back with errorCode 4005 INVALID_EVIDENCE.

Questions

  1. In Protected PCIe mode, is NVSwitch attestation expected to work from inside the tenant guest at all, or is it host side only (Fabric Manager / NVOS)? The infrastructure operator tells me nvattest attest --device nvswitch works on their 8x H200 from the host.
  2. If it is supported from the guest, which libnvidia-nscq version is expected with driver 595.71.05, and does the guest need anything beyond the /dev/nvidia-nvswitch* nodes (Fabric Manager access, a specific kernel module option)?
  3. Side issue: in the same guest, nvattest attest --device gpu fails intermittently on the local OCSP check (“Fatal libcurl error code: SSL connect error (35)” against ocsp.ndis.nvidia.com, then “Failed to generate certificate chain claims for RIM with id NV_GPU_DRIVER_GH100_595.71.05”), while the SDK in REMOTE mode always succeeds because NRAS runs the revocation checks on its side. Is there a supported way to make nvattest rely on NRAS for revocation instead of local OCSP?

Full log of the run, including the exact commands, is published here:
https://voltagegpu.com/blog/two-proofs/evidence/nvswitch-2026-09-16/README.txt

For context, we publish tenant side attestation bundles (TDX quote plus NRAS tokens bound to the same challenge) for the single GPU SKUs, produced with an open source verifier, GitHub - Jabsama/voltage-verify: Bind an Intel TDX quote and an NVIDIA GPU attestation to the workload you meant to run, then verify the bundle without trusting the cloud provider. · GitHub . Until this NSCQ question is answered we state publicly that the NVLink fabric on 8 GPU nodes is not attested, and I would rather write the correct thing than guess.

Julien Aubry, VoltageGPU