Ferlab-Ste-Justine/quality-control-pipeline is a Nextflow (DSL2) bioinformatics pipeline for comprehensive quality control of genomic sequencing data. It accepts FASTQ reads, BAM/CRAM alignments, and VCF variant files — running a full suite of QC tools across each data type — and aggregates all results into a single MultiQC HTML report. It can also ingest pre-computed DRAGEN per-sample metrics in place of recomputing them. It is designed to handle multi-sample cohorts with mixed sequencing strategies (WGS, WES, targeted panels) and supports pedigree-based sample identity checks.
- BAM/CRAM merging — merge multi-run alignment files per sample (Samtools)
- FASTQ QC — raw read quality metrics (FastQC) and sample identity checking (ngsCheckMate)
- Alignment QC — alignment statistics (Samtools stats), WGS metrics (Picard CollectWgsMetrics), sequencing depth (Mosdepth), and DNA contamination estimation (VerifyBamID2)
- Per-region coverage — per-gene coverage summaries for up to two custom BED region sets
- Sample identity & relatedness — genetic relatedness checking, per-family or cohort-wide (Somalier)
- VCF QC — variant counts by type and zygosity, Ts/Tv ratio (BCFtools)
- Report aggregation — all QC results consolidated into a single interactive report (MultiQC)
Alignment and variant metrics can alternatively be sourced from pre-computed DRAGEN metrics CSVs via --dragen_metrics_dir.
Note
If you are new to Nextflow and nf-core, please refer to this page on how to set up Nextflow. Make sure to test your setup with -profile test before running the workflow on actual data.
Prepare a samplesheet CSV describing your samples (see usage docs for full column reference):
participant,sample,familyId,fileType,file1,file2
P001,S001,FAM1,FASTQ,/data/S001_R1.fastq.gz,/data/S001_R2.fastq.gz
P002,S002,FAM1,CRAM,/data/S002.cram,/data/S002.crai
P003,S003,FAM2,VCF,/data/S003.vcf.gz,/data/S003.vcf.gz.tbiThen run the pipeline:
nextflow run Ferlab-Ste-Justine/quality-control-pipeline \
-profile docker \
--input samplesheet.csv \
--outdir ./results \
--fasta /path/to/reference.fa \
--somalier_sites /path/to/sites.vcf.gzFor all available parameters, see docs/usage.md. For a description of the output files, see docs/output.md.
Warning
Provide pipeline parameters via the CLI or a Nextflow -params-file. Custom config files specified with -c must only be used for tuning process resource specifications or infrastructural tweaks, not for pipeline parameters.
Ferlab-Ste-Justine/quality-control-pipeline was originally written by Georgette Femerling, Lysiane Bouchard.
If you would like to contribute to this pipeline, please see the contributing guidelines.
An extensive list of references for the tools used by the pipeline can be found in the CITATIONS.md file.
This pipeline uses code and infrastructure developed and maintained by the nf-core community, reused here under the MIT license.
The nf-core framework for community-curated bioinformatics pipelines.
Philip Ewels, Alexander Peltzer, Sven Fillinger, Harshil Patel, Johannes Alneberg, Andreas Wilm, Maxime Ulysse Garcia, Paolo Di Tommaso & Sven Nahnsen.
Nat Biotechnol. 2020 Feb 13. doi: 10.1038/s41587-020-0439-x.