Deepvariant fails with cram file, OK with bam

Dear all

I would be grateful for any help with this problem that’s taken many hours

My call to pbrun deepvariant on a cram file fails with this error

PB Info 2025-Dec-02 18:33:43] Starting DeepVariant[PB Info 2025-Dec-02 18:33:43] Running with 2 GPU devices, each with 4 group instances and 6 workers[PB Error 2025-Dec-02 18:33:52][src/PBCramFile.cpp:416] cram_decode_slice_pb failed with a status of -1, expected sts == 0, exiting.[PB Info 2025-Dec-02 18:33:52] ProgressMeter -	Current-Locus	Elapsed-Minutes[PB Debug 2025-Dec-02 18:33:52][./inc2/common.h:112] NvInfer: CUDA initialization failure with error: 35terminate called after throwing an instance of ‘std::length_error’what():  basic_string::_M_create

I now have this test file. I start with a sam file and convert to bam. pbrun-dv runs fine. I convert to cram passing the reference sequence and now deepvariant fails. I can convert the cram file back to bam and it all works OK. As a sanity check, if I deliberately make an error in the reference name, then pbrun fails with a very different error.

This is my command

/usr/local/parabricks/pbrun deepvariant     --in-bam see.cram    --ref chm13v2.0.fa     --gvcf     --out-variants AGS0202.g.vcf.gz  --keep-tmp --verbose  --logfile LOGFILE

With other data sets I’ve used cram files so I know it’s something about this file. The only obvious thing I could see in this cram file was that it had UR tags that previous crams didn’t. I removed them and that didn’t help.

FWIW, my header looks like this

```

@HD	VN:1.6	SO:coordinate
@SQ	SN:chr1	LN:248387328	M5:e469247288ceb332aee524caec92bb22
@SQ	SN:chr2	LN:242696752	M5:e5ff97665b6025191d4d7dfa2cec4cf0
@SQ	SN:chr3	LN:201105948	M5:0012f8a55c07f0b61e199736105e7257
@SQ	SN:chr4	LN:193574945	M5:b7e17a4f043aab5bc45cdfd5cccf9deb
@SQ	SN:chr5	LN:182045439	M5:e069fd69dce480b348767fa9827671b4
@SQ	SN:chr6	LN:172126628	M5:ae714983adbd673c7c84b0b1daa91f7e
@SQ	SN:chr7	LN:160567428	M5:3ff4435e65cf8cdd69a3150d09dd649f
@SQ	SN:chr8	LN:146259331	M5:29aeec58cb54d9132e2283839f9fcc9c
@SQ	SN:chr9	LN:150617247	M5:419dad1fb58ae389ad8828536b5d78cd
@SQ	SN:chr10	LN:134758134	M5:fbf3087282695f275d460157fc7864a5
@SQ	SN:chr11	LN:135127769	M5:7f144a0b4800c341a4b4f1518838cc59
@SQ	SN:chr12	LN:133324548	M5:5086cf368e83b9cefd7bb19ef0f40d3c
@SQ	SN:chr13	LN:113566686	M5:6f9949b26d5ece38c6cdefd48a516907
@SQ	SN:chr14	LN:101161492	M5:7ad36d43a1ab059e190ab918f27cb5bf
@SQ	SN:chr15	LN:99753195	M5:914e0733351aaa9a4354ca8efb527401
@SQ	SN:chr16	LN:96330374	M5:fcc04e14d89b12c757d56c177261e248
@SQ	SN:chr17	LN:84276897	M5:8627043bd24255e9f356e0c99ece251f
@SQ	SN:chr18	LN:80542538	M5:b6a85aabc31b2d57f707aa679a8612c5
@SQ	SN:chr19	LN:61707364	M5:c2c74cb0832b34b9e8eebf4b744a88c1
@SQ	SN:chr20	LN:66210255	M5:f1a31f30d3da50893470ff47bf0b5a38
@SQ	SN:chr21	LN:45090682	M5:760794bbbe965d4d1e3884bf09d1f306
@SQ	SN:chr22	LN:51324926	M5:0068189b6b7c95e860cce0815662293f
@SQ	SN:chrX	LN:154259566	M5:1f2616c4cfa30104517e44dd3c426453
@SQ	SN:chrY	LN:62460029	M5:775564f397e3690ae4e4f0372537ff49
@SQ	SN:chrM	LN:16569	M5:ec493a132ac4823aa696e37109f64972
@PG	ID:bwa-mem2	PN:bwa-mem2	VN:2.2.1	CL:bwa-mem2 mem -t 18 t2t/chm13v2.0.fa AGS0202_1.fq.gz AGS0202_2.fq.gz
@PG	ID:samtools	PN:samtools	PP:bwa-mem2	VN:1.22	CL:samtools view -@4 -b -o AGS0202.bam
@PG	ID:samtools.1	PN:samtools	PP:samtools	VN:1.22	CL:samtools fixmate -@6 -m AGS0202.bam -
@PG	ID:samtools.2	PN:samtools	PP:samtools.1	VN:1.22	CL:samtools sort -@8 -m 12G
@PG	ID:samtools.3	PN:samtools	PP:samtools.2	VN:1.22	CL:samtools markdup -@10 - -
@PG	ID:samtools.4	PN:samtools	PP:samtools.3	VN:1.22	CL:samtools view -o see.bam see.sam
@PG	ID:samtools.5	PN:samtools	PP:samtools.4	VN:1.22	CL:samtools view -h -T t2t/chm13v2.0.fa -o see.cram see.bam
@PG	ID:samtools.6	PN:samtools	PP:samtools.5	VN:1.22	CL:samtools view -H see.cram
@PG	ID:samtools.7	PN:samtools	PP:samtools.6	VN:1.22	CL:samtools reheader - see.cram
@PG	ID:samtools.8	PN:samtools	PP:samtools.7	VN:1.22	CL:samtools view -H output.cram
@RG	ID:LP2100147-DNA_E02_L1	LB:RGLB	PU:1	SM:AGS0202


Hi @shaze , thank you for your question. Could you please share the cram and reference files so that we can try to reproduce the issue? Please also share details about the system/GPUs you are using if possible.

Thanks Abhair, much appreciated.

  1. Rocky 9 server with 256GB of RAM, two L4 GPUs. version Running clara-parabricks_4.6.0-1 through apptainer
  2. Further testing shows this pattern: If I run my fq through fq2bam to produce cram and then run deepvariant, there’s no problem. If I produce BAM, the BAM file works fine. But if I then use samtools (v22) to convert to cram I get a significantly smaller CRAM that fails on deepvariant
  3. I am using T2T but I have tested with GRCh38 and I get the same behaviour.
  4. I have created a sample data set for you here http://www.bioinf..wits.ac.za/~scott/pb/pb.tar.gz (1.2GB, mainly because of reference genome). The sample is a public data set (not the one I originally used).
  5. I have included the fq.gz used. E_fq2bam_t.cram is the cram produced directly via fq2bam. E_sam_t.cram was produced first by producing a BAM file with fq2bam and then converting to cram with this.
  6. In my real data, I am using standard bwa-mem2 on CPUs to produce CRAM and then using samtools to convert to CRAM so it seems to be something about the way samtools produces cram. (As an aside, I’d love to use fq2bam, but the memory requirements appear to be well over 300GB of RAM on the CPU side — no problem with RAM on the device, and in fact with two L4s I don’t have to use --low-memory. bwa-mem2 only requires about 90GB of RAM. )
  7. I can use the standard deepvariant release on the crams I produce.

Please let me know if you need more information.

Hi @shaze , thanks for uploading the files.

The error thrown by deepvariant is because of the version of the CRAM file that is supplied as an input. samtools (v22) produces CRAM files consistent with version 3.1 by default. However, parabricks deepvariant currently only supports CRAM files up to version 3.0.

To resolve the error, please specify the version of the CRAM file needed when using samtools to create the file. (Steps are described here: CRAM Benchmarks)

Also thank you for your comment about fq2bam memory usage. We will take this into account for future development. Please feel free to respond with further questions or close this ticket if your problem is resolved.

Many thanks for the update. I very much appreciate the prompt response.

The memory change would be useful. Our GPU machines are multi-purpose and most applications fit well into 256GB of RAM. We don’t have resources to upgrade all to 512GB of RAM, but if we don’t then we won’t have capacity for larger projects – hence moving to CPUs.

As an aside, I’d love to use fq2bam, but the memory requirements appear to be well over 300GB of RAM on the CPU side — no problem with RAM on the device, and in fact with two L4s I don’t have to use --low-memory

@shaze , when you saw such high memory use was this when fq2bam was engaging recovery mode? I ask because we fixed a host memory issue with the CPU fallback in recovery mode in nvcr.io/nvidia/clara/clara-parabricks:4.6.0-2.

Additionally, we have some parameters which can be used to tune memory use (both host and device): fq2bam (BWA-MEM + GATK) - NVIDIA Docs .

Thank you for the update. Yes – I think so – there are lots of such messages.

I have tried the new version and it does make a difference. I can now run fq2bam using about 240GB with two GPUs (if I try more, then I run out of RAM).

I will test out the different optimisation parameters that you pointed out. Perhaps I am doing something wrong but (a) when I use Nvidia-smi -l most of time time the GPUs appear to be idle. Also I see no difference in performance whether I use two L4 GPUs (with -low-memory) or two L40S GPUs (not using that), which is not what I expected based on the benchmarks tests that you’ve done.

Hi @shaze , it is possible that the performance is lower than expected because fq2bam constantly needs to enter recovery mode. Do you have any public versions of the data? I could look and see if there are alternate parameters which would work better for this data and allow it to be processed on GPU for every batch.