Reporting a failure I observed in the Ready2Run “Parabricks Germline DeepVariant”
workflow on AWS HealthOmics: the run reports success but produces a 0-byte VCF.
Summary
Running the Ready2Run “Parabricks Germline DeepVariant” workflow on AWS HealthOmics,
the run reports SUCCESS but the output VCF is 0 bytes. The fq2bam (alignment) stage
completes fine and produces a valid BAM/BAI; only variant calling fails.
Root cause
In the task log, the deepvariant stage fails at the call_variants step with a TensorRT
engine deserialization error:
NvInfer ERROR: 1: Serialization (assertion safeVersionRead == safeSerializationVersion
failed. Version tag does not match. Current Version: 0, Serialized Engine Version: 96)
DeepVariant in Parabricks runs CNN inference through a TensorRT engine file. The engine
baked into this Ready2Run image appears to have been serialized with a TensorRT version
that doesn’t match the runtime in the image, so it fails to deserialize. The wrapper still
exits 0, so HealthOmics marks the run “successful” and writes a header-only, 0-byte VCF.
This matches the general TensorRT “safeVersionRead == safeSerializationVersion” symptom
reported across other NVIDIA projects — an engine/runtime version mismatch.
Why it’s particularly bad
It’s a silent failure: a “successful” run with an empty output. Anyone not checking
output file sizes could ship nothing without noticing.
Workaround that worked
Instead of the prebuilt Ready2Run image, I built a clean Parabricks image and ran
DeepVariant as a HealthOmics private workflow against the already-produced BAM:
- docker pull nvcr.io/nvidia/clara/clara-parabricks:4.5.1-1 from NGC
- push it to a private ECR repo; grant omics.amazonaws.com pull on the repo
- a small WDL task that runs
pbrun deepvarianton the existing BAM, with an explicit
non-empty-output guard (test -s VCF || exit 1) so a silent 0-byte VCF fails the run
With the clean image (Parabricks 4.5.1-1), call_variants proceeds normally — ProgressMeter
advances across all chromosomes through chrX/chrY, the run logs “Deepvariant is finished”,
and we get a valid (non-empty) VCF. Total DeepVariant time was ~22 minutes on a 4×A10G
instance.
Questions for NVIDIA/AWS
- Can the TensorRT engine in the Ready2Run DeepVariant image be rebuilt to match the
image’s runtime? - Can the wrapper be fixed to return a non-zero exit code on engine-load failure, so the
run fails loudly instead of emitting a 0-byte VCF?
Environment: observed on a Ready2Run Parabricks DeepVariant (50X) run in us-east-1, GRCh38,
in June 2026.