Computes per-base read coverage by folding a per-contig difference array inside
the multithreaded C++ reader and run-length encoding the result, without ever
materialising the alignments as an R GAlignments object. This is the fast path
for coverage / bigWig generation on large BAMs: the standard route
(GenomicAlignments::coverage(readGAlignments(file))) builds a full
GAlignments in R first, which is the dominant cost on 100M+-read files.
Usage
bam_coverage(
file,
param = NULL,
threads = 1L,
BPPARAM = BiocParallel::bpparam(),
auto_threads = FALSE
)Arguments
- file
A BAM file path, Rsamtools::BamFile, or a vector /
BamFileListof them.- param
Optional Rsamtools::ScanBamParam (or a compatible list). Only its
flagandmapqFilterare honoured, applied per read in C++. With the defaultNULLthe filter matchesreadGAlignments()(drop only unmapped reads). Note thatwhich(region) restriction is not applied.- threads
Number of OpenMP threads for within-file decoding.
- BPPARAM
A BiocParallel::BiocParallelParam for across-file parallelism when
filehas more than one element.- auto_threads
Logical; if
TRUE, balance the thread / worker split.
Value
An IRanges::RleList (SimpleRleList) of integer coverage, one element
per BAM-header sequence in header order, each of length equal to the header
sequence length (contigs with no reads are all-zero runs). For multiple input
files, a named list of such RleLists.
Details
The result is byte-identical (identical()) to
GenomicAlignments::coverage(GenomicAlignments::readGAlignments(file)): CIGAR
operations M, =, X and D contribute to coverage, N introduces a gap,
and I/S/H/P are ignored; only unmapped reads (SAM flag 0x4) are
dropped by default – secondary, supplementary, duplicate and QC-fail reads are
counted, exactly as readGAlignments() does.
