Skip to content

About

No description, website, or topics provided.

Resources

Contributing

Stars

0 stars

Watchers

4 watching

Forks

Latest commit

 

History

126 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Ferlab-Ste-Justine/quality-control-pipeline

GitHub Actions CI Status GitHub Actions Linting Status nf-test

Nextflow run with conda run with docker run with singularity

Introduction

Ferlab-Ste-Justine/quality-control-pipeline is a Nextflow (DSL2) bioinformatics pipeline for comprehensive quality control of genomic sequencing data. It accepts FASTQ reads, BAM/CRAM alignments, and VCF variant files — running a full suite of QC tools across each data type — and aggregates all results into a single MultiQC HTML report. It can also ingest pre-computed DRAGEN per-sample metrics in place of recomputing them. It is designed to handle multi-sample cohorts with mixed sequencing strategies (WGS, WES, targeted panels) and supports pedigree-based sample identity checks.

Pipeline steps

  1. BAM/CRAM merging — merge multi-run alignment files per sample (Samtools)
  2. FASTQ QC — raw read quality metrics (FastQC) and sample identity checking (ngsCheckMate)
  3. Alignment QC — alignment statistics (Samtools stats), WGS metrics (Picard CollectWgsMetrics), sequencing depth (Mosdepth), and DNA contamination estimation (VerifyBamID2)
  4. Per-region coverage — per-gene coverage summaries for up to two custom BED region sets
  5. Sample identity & relatedness — genetic relatedness checking, per-family or cohort-wide (Somalier)
  6. VCF QC — variant counts by type and zygosity, Ts/Tv ratio (BCFtools)
  7. Report aggregation — all QC results consolidated into a single interactive report (MultiQC)

Alignment and variant metrics can alternatively be sourced from pre-computed DRAGEN metrics CSVs via --dragen_metrics_dir.

Usage

Note

If you are new to Nextflow and nf-core, please refer to this page on how to set up Nextflow. Make sure to test your setup with -profile test before running the workflow on actual data.

Prepare a samplesheet CSV describing your samples (see usage docs for full column reference):

participant,sample,familyId,fileType,file1,file2
P001,S001,FAM1,FASTQ,/data/S001_R1.fastq.gz,/data/S001_R2.fastq.gz
P002,S002,FAM1,CRAM,/data/S002.cram,/data/S002.crai
P003,S003,FAM2,VCF,/data/S003.vcf.gz,/data/S003.vcf.gz.tbi

Then run the pipeline:

nextflow run Ferlab-Ste-Justine/quality-control-pipeline \
   -profile docker \
   --input samplesheet.csv \
   --outdir ./results \
   --fasta /path/to/reference.fa \
   --somalier_sites /path/to/sites.vcf.gz

For all available parameters, see docs/usage.md. For a description of the output files, see docs/output.md.

Warning

Provide pipeline parameters via the CLI or a Nextflow -params-file. Custom config files specified with -c must only be used for tuning process resource specifications or infrastructural tweaks, not for pipeline parameters.

Credits

Ferlab-Ste-Justine/quality-control-pipeline was originally written by Georgette Femerling, Lysiane Bouchard.

Contributions and Support

If you would like to contribute to this pipeline, please see the contributing guidelines.

Citations

An extensive list of references for the tools used by the pipeline can be found in the CITATIONS.md file.

This pipeline uses code and infrastructure developed and maintained by the nf-core community, reused here under the MIT license.

The nf-core framework for community-curated bioinformatics pipelines.

Philip Ewels, Alexander Peltzer, Sven Fillinger, Harshil Patel, Johannes Alneberg, Andreas Wilm, Maxime Ulysse Garcia, Paolo Di Tommaso & Sven Nahnsen.

Nat Biotechnol. 2020 Feb 13. doi: 10.1038/s41587-020-0439-x.

About

No description, website, or topics provided.

Resources

Contributing

Stars

0 stars

Watchers

4 watching

Forks

Releases

Packages

Used by

Contributors

Languages