I would be grateful for any help with this problem that’s taken many hours
My call to pbrun deepvariant on a cram file fails with this error
PB Info 2025-Dec-02 18:33:43] Starting DeepVariant[PB Info 2025-Dec-02 18:33:43] Running with 2 GPU devices, each with 4 group instances and 6 workers[PB Error 2025-Dec-02 18:33:52][src/PBCramFile.cpp:416] cram_decode_slice_pb failed with a status of -1, expected sts == 0, exiting.[PB Info 2025-Dec-02 18:33:52] ProgressMeter - Current-Locus Elapsed-Minutes[PB Debug 2025-Dec-02 18:33:52][./inc2/common.h:112] NvInfer: CUDA initialization failure with error: 35terminate called after throwing an instance of ‘std::length_error’what(): basic_string::_M_create
I now have this test file. I start with a sam file and convert to bam. pbrun-dv runs fine. I convert to cram passing the reference sequence and now deepvariant fails. I can convert the cram file back to bam and it all works OK. As a sanity check, if I deliberately make an error in the reference name, then pbrun fails with a very different error.
With other data sets I’ve used cram files so I know it’s something about this file. The only obvious thing I could see in this cram file was that it had UR tags that previous crams didn’t. I removed them and that didn’t help.
Hi @shaze , thank you for your question. Could you please share the cram and reference files so that we can try to reproduce the issue? Please also share details about the system/GPUs you are using if possible.
Rocky 9 server with 256GB of RAM, two L4 GPUs. version Running clara-parabricks_4.6.0-1 through apptainer
Further testing shows this pattern: If I run my fq through fq2bam to produce cram and then run deepvariant, there’s no problem. If I produce BAM, the BAM file works fine. But if I then use samtools (v22) to convert to cram I get a significantly smaller CRAM that fails on deepvariant
I am using T2T but I have tested with GRCh38 and I get the same behaviour.
I have created a sample data set for you here http://www.bioinf..wits.ac.za/~scott/pb/pb.tar.gz (1.2GB, mainly because of reference genome). The sample is a public data set (not the one I originally used).
I have included the fq.gz used. E_fq2bam_t.cram is the cram produced directly via fq2bam. E_sam_t.cram was produced first by producing a BAM file with fq2bam and then converting to cram with this.
In my real data, I am using standard bwa-mem2 on CPUs to produce CRAM and then using samtools to convert to CRAM so it seems to be something about the way samtools produces cram. (As an aside, I’d love to use fq2bam, but the memory requirements appear to be well over 300GB of RAM on the CPU side — no problem with RAM on the device, and in fact with two L4s I don’t have to use --low-memory. bwa-mem2 only requires about 90GB of RAM. )
I can use the standard deepvariant release on the crams I produce.
The error thrown by deepvariant is because of the version of the CRAM file that is supplied as an input. samtools (v22) produces CRAM files consistent with version 3.1 by default. However, parabricks deepvariant currently only supports CRAM files up to version 3.0.
To resolve the error, please specify the version of the CRAM file needed when using samtools to create the file. (Steps are described here: CRAM Benchmarks)
Also thank you for your comment about fq2bam memory usage. We will take this into account for future development. Please feel free to respond with further questions or close this ticket if your problem is resolved.
Many thanks for the update. I very much appreciate the prompt response.
The memory change would be useful. Our GPU machines are multi-purpose and most applications fit well into 256GB of RAM. We don’t have resources to upgrade all to 512GB of RAM, but if we don’t then we won’t have capacity for larger projects – hence moving to CPUs.
As an aside, I’d love to use fq2bam, but the memory requirements appear to be well over 300GB of RAM on the CPU side — no problem with RAM on the device, and in fact with two L4s I don’t have to use --low-memory
@shaze , when you saw such high memory use was this when fq2bam was engaging recovery mode? I ask because we fixed a host memory issue with the CPU fallback in recovery mode in nvcr.io/nvidia/clara/clara-parabricks:4.6.0-2.
Thank you for the update. Yes – I think so – there are lots of such messages.
I have tried the new version and it does make a difference. I can now run fq2bam using about 240GB with two GPUs (if I try more, then I run out of RAM).
I will test out the different optimisation parameters that you pointed out. Perhaps I am doing something wrong but (a) when I use Nvidia-smi -l most of time time the GPUs appear to be idle. Also I see no difference in performance whether I use two L4 GPUs (with -low-memory) or two L40S GPUs (not using that), which is not what I expected based on the benchmarks tests that you’ve done.
Hi @shaze , it is possible that the performance is lower than expected because fq2bam constantly needs to enter recovery mode. Do you have any public versions of the data? I could look and see if there are alternate parameters which would work better for this data and allow it to be processed on GPU for every batch.