Ready2Run Parabricks DeepVariant on HealthOmics: silent 0-byte VCF from TensorRT engine version mismatch

Reporting a failure I observed in the Ready2Run “Parabricks Germline DeepVariant”
workflow on AWS HealthOmics: the run reports success but produces a 0-byte VCF.

Summary

Running the Ready2Run “Parabricks Germline DeepVariant” workflow on AWS HealthOmics,
the run reports SUCCESS but the output VCF is 0 bytes. The fq2bam (alignment) stage
completes fine and produces a valid BAM/BAI; only variant calling fails.

Root cause

In the task log, the deepvariant stage fails at the call_variants step with a TensorRT
engine deserialization error:

NvInfer ERROR: 1: Serialization (assertion safeVersionRead == safeSerializationVersion
failed. Version tag does not match. Current Version: 0, Serialized Engine Version: 96)

DeepVariant in Parabricks runs CNN inference through a TensorRT engine file. The engine
baked into this Ready2Run image appears to have been serialized with a TensorRT version
that doesn’t match the runtime in the image, so it fails to deserialize. The wrapper still
exits 0, so HealthOmics marks the run “successful” and writes a header-only, 0-byte VCF.

This matches the general TensorRT “safeVersionRead == safeSerializationVersion” symptom
reported across other NVIDIA projects — an engine/runtime version mismatch.

Why it’s particularly bad

It’s a silent failure: a “successful” run with an empty output. Anyone not checking
output file sizes could ship nothing without noticing.

Workaround that worked

Instead of the prebuilt Ready2Run image, I built a clean Parabricks image and ran
DeepVariant as a HealthOmics private workflow against the already-produced BAM:

  1. docker pull nvcr.io/nvidia/clara/clara-parabricks:4.5.1-1 from NGC
  2. push it to a private ECR repo; grant omics.amazonaws.com pull on the repo
  3. a small WDL task that runs pbrun deepvariant on the existing BAM, with an explicit
    non-empty-output guard (test -s VCF || exit 1) so a silent 0-byte VCF fails the run

With the clean image (Parabricks 4.5.1-1), call_variants proceeds normally — ProgressMeter
advances across all chromosomes through chrX/chrY, the run logs “Deepvariant is finished”,
and we get a valid (non-empty) VCF. Total DeepVariant time was ~22 minutes on a 4×A10G
instance.

Questions for NVIDIA/AWS

  • Can the TensorRT engine in the Ready2Run DeepVariant image be rebuilt to match the
    image’s runtime?
  • Can the wrapper be fixed to return a non-zero exit code on engine-load failure, so the
    run fails loudly instead of emitting a 0-byte VCF?

Environment: observed on a Ready2Run Parabricks DeepVariant (50X) run in us-east-1, GRCh38,
in June 2026.