A Step-by-Step Guide to Designing a Germline Variant Calling Pipeline from Raw Sequencing Data
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Germline variant calling pipelines require a modular approach, progressing from raw FASTQ quality assessment (e.g., FastQC, MultiQC) to read alignment (e.g., BWA-MEM), post-alignment preprocessing (e.g., GATK MarkDuplicates, BaseRecalibrator), variant calling (e.g., GATK HaplotypeCaller), and finally variant filtering.
- Accurate alignment to a specific reference genome build (e.g., GRCh37, GRCh38) is critical, as misalignment in repetitive or homopolymer regions is a common failure pattern that impacts downstream variant accuracy.
- Post-alignment preprocessing, including marking duplicate reads and base quality score recalibration using known variant sites, is essential to correct systematic errors and improve the reliability of variant calls.
- GATK HaplotypeCaller in GVCF mode followed by joint genotyping is recommended for multi-sample cohorts to enhance genotype accuracy, particularly at low-coverage sites, and to facilitate cohort-wide variant discovery.
- Variant filtering, whether through hard filters or machine learning recalibration, involves a critical sensitivity versus precision tradeoff, with annotations like Fisher strand bias and mapping quality rank sum being key discriminators of true variants.
- Reproducibility is paramount, necessitating meticulous documentation of tool versions, parameters, reference genomes, and input file checksums, often facilitated by workflow managers and containerization technologies.
Germline variant calling is the process of identifying inherited sequence differences between a sample and a reference genome, starting from raw sequencing reads. This article provides a modular blueprint for building a germline variant calling pipeline from raw FASTQ files to a filtered VCF file, with concrete decision points at each stage. The intended reader is a biology student, researcher, laboratory professional, or life-science practitioner who has sequencing data and needs a defensible, reproducible analysis framework. The pipeline described here uses BWA-MEM for alignment and GATK HaplotypeCaller for variant discovery, with preprocessing steps that follow community best practices. Each section covers the purpose of the step, the specific tool choices, the parameters that matter, the records you should keep, and the common failure patterns you will encounter.
Scope and Reader Context
This article addresses the specific problem of building a germline variant calling pipeline from raw sequencing data. You have FASTQ files from a whole genome sequencing (WGS) or whole exome sequencing (WES) experiment, and you need to produce a variant call format (VCF) file that contains high-confidence single nucleotide variants (SNVs) and small insertions or deletions (indels). The pipeline described here is for germline variants, meaning variants that are present in the germline genome and inherited, not somatic variants that arise in tumor tissue. The distinction matters because somatic variant calling requires matched normal samples and different filtering logic, as discussed in the limitations section.
The workflow has five major stages: raw data quality assessment, read alignment to a reference genome, post-alignment preprocessing, variant calling, and variant filtering. Each stage has multiple tool options and parameter choices. The key insight from the literature is that no single pipeline is optimal across the entire genome, and pipeline design requires tradeoffs based on your application [<a href="#ref-1">1</a>]. This means you must make deliberate choices at each stage and document those choices for reproducibility.
At a Glance: Pipeline Stage Decision Table
| Pipeline Stage | Primary Tool Choice | Key Decision Point | Output File | Common Failure Pattern |
|---|---|---|---|---|
| Raw data QC | FastQC, MultiQC | Adapter contamination and base quality thresholds | HTML report, trimmed FASTQ | Trimming too aggressively removes biological signal |
| Read alignment | BWA-MEM | Reference genome version and index | SAM/BAM file | Misalignment in repetitive or homopolymer regions |
| Post-alignment preprocessing | GATK MarkDuplicates, BaseRecalibrator | Duplicate removal and base quality score recalibration | Recalibrated BAM file | Overly strict duplicate removal reduces coverage |
| Variant calling | GATK HaplotypeCaller | GVCF mode versus single-sample calling | GVCF or VCF file | Low-quality variants in difficult-to-map regions |
| Variant filtering | GATK VariantRecalibrator or hard filters | Sensitivity versus precision tradeoff | Filtered VCF file | Over-filtering removes true variants in low-complexity regions |
Raw Data Inputs and Quality Assessment
FASTQ File Structure and What It Tells You
The pipeline begins with FASTQ files, which contain the raw sequencing reads and their associated quality scores. Each read has four lines: a sequence identifier, the nucleotide sequence, a plus sign, and the quality scores encoded as ASCII characters. The quality scores represent the probability that a base call is incorrect, and they are the primary input for downstream quality filtering decisions.
Before you run any alignment, you need to assess the quality of your raw data. The NCBI provides access to sequence read archives and quality assessment resources that can help you understand the expected format and quality metrics for your sequencing platform [<a href="#ref-2">2</a>]. The EMBL-EBI training materials cover the fundamentals of working with sequencing data and quality assessment, which is useful if you are new to this type of analysis [<a href="#ref-3">3</a>].
Quality Metrics You Must Record
For each FASTQ file, record the following metrics before proceeding:
- Total number of reads and total bases
- Per-base quality scores across read length
- GC content distribution
- Adapter contamination levels
- Duplication rates
- Overrepresented sequences
These metrics tell you whether your sequencing run was successful and whether any library preparation issues occurred. A high duplication rate suggests over-amplification during library preparation. Adapter contamination indicates that the insert size was too short for the read length. Both issues affect downstream variant calling accuracy.
Trimming Decisions and Their Consequences
If adapter contamination is present, you need to trim the adapters before alignment. The decision to trim is not automatic. Trimming removes bases from the ends of reads, and if you trim too aggressively, you remove biological sequence that could contain variants. The standard approach is to trim adapters and low-quality bases only at the read ends, leaving the internal sequence intact.
The Galaxy Training Network provides accessible tutorials on quality control and trimming workflows that demonstrate the decision points in this stage [<a href="#ref-4">4</a>]. These tutorials show how to interpret quality reports and make trimming decisions based on the data instead of applying a universal threshold.
Records for the Quality Assessment Stage
Create a quality assessment log that records:
- The version of the quality assessment tool used
- The date of the analysis
- The input file names and their checksums
- The quality metrics listed above
- The trimming parameters used, if trimming was performed
- The number of reads removed by trimming
This log becomes part of your pipeline documentation and is essential for reproducing the analysis later.
Reference Genome Selection and Indexing
Choosing the Correct Reference Genome
The reference genome is the backbone of your variant calling pipeline. Every read will be aligned to this reference, and every variant will be reported relative to its coordinates. The choice of reference genome version affects the coordinates in your VCF file, the annotation databases you can use, and the accuracy of alignment in difficult regions.
For human data, the most commonly used references are GRCh37 and GRCh38. These are different builds of the human genome, and they have different coordinates for many variants. If you are comparing your results to public databases, you must use the same reference build that those databases use. The NCBI provides access to reference genome assemblies and their documentation [<a href="#ref-2">2</a>].
Reference Indexing Requirements
Before alignment, the reference genome must be indexed for the aligner you plan to use. BWA-MEM requires a BWA index, which consists of several files that allow the aligner to search the reference quickly. The GATK tools require a dictionary file and an index file for the reference FASTA. These indexing steps are prerequisites for alignment and variant calling, and they must be performed once for each reference genome version.
The Impact of Reference Choice on Variant Calling Accuracy
The choice of reference genome affects variant calling accuracy in specific genomic contexts. Research using interpretable machine learning to predict germline variant calling errors has shown that hard-to-map regions and homopolymer regions contribute disproportionately to errors [<a href="#ref-1">1</a>]. The same research demonstrated that graph-based references can improve recall in hard-to-map regions compared to linear references [<a href="#ref-1">1</a>]. This means that if your study focuses on regions with high sequence complexity or segmental duplications, you should consider whether a linear reference is sufficient or whether a graph-based approach would improve your results.
Records for the Reference Genome Stage
Record the following in your pipeline documentation:
- The exact reference genome version and build
- The source of the reference FASTA file
- The checksum of the reference file
- The versions of all index files created
- The date the reference was downloaded and indexed
These records ensure that anyone reviewing your analysis can identify the exact reference used and reproduce the alignment.
Read Alignment with BWA-MEM
How BWA-MEM Works and Why It Is the Default Choice
BWA-MEM is a widely used aligner for germline variant calling pipelines. It aligns reads to the reference genome by finding the best matching position for each read, allowing for mismatches and small indels. The algorithm is designed to handle the error profiles of Illumina sequencing data, which is the most common platform for germline variant studies.
The choice of aligner is one of the most consequential decisions in the pipeline. Different aligners have different strengths and weaknesses in different genomic contexts. The literature on pipeline optimization emphasizes that choosing appropriate tools and parameters is the key problem in achieving optimal precision and recall [<a href="#ref-5">5</a>]. BWA-MEM is a reasonable default because it is well documented, widely tested, and produces alignments that work well with the GATK variant calling tools.
Alignment Parameters That Matter
The default parameters for BWA-MEM are appropriate for most standard germline variant calling applications. The parameters that you might need to adjust include:
- The minimum seed length, which affects sensitivity for shorter reads
- The band width, which affects the ability to detect longer indels
- The number of threads, which affects runtime but not results
For standard Illumina reads of 100 to 150 base pairs, the default parameters are appropriate. If you are working with longer reads or a different sequencing platform, you should consult the aligner documentation and test different parameter settings on a subset of your data.
The SAM and BAM File Formats
The output of BWA-MEM is a SAM file, which is a text format that contains the alignment information for each read. The SAM file is typically converted to a BAM file, which is the binary compressed version of the same information. The BAM file is the standard input for downstream preprocessing and variant calling tools.
The SAM format includes a header section with information about the reference genome and the alignment parameters, and an alignment section with one line per read. Each alignment line includes the read name, the reference position, the mapping quality, and the CIGAR string, which describes how the read aligns to the reference.
Common Alignment Failure Patterns
The most common alignment failures occur in repetitive regions, homopolymer stretches, and segmental duplications. In these regions, reads may map to multiple locations with equal or nearly equal scores, and the aligner must choose one location. The mapping quality score reflects this uncertainty, and reads with low mapping quality are often filtered out in downstream steps.
Research on variant calling errors has shown that difficult-to-map regions and homopolymer regions are major contributors to errors in germline variant calling [<a href="#ref-1">1</a>]. This means that even with a well-performing aligner, you will have systematic errors in these regions. You should expect lower sensitivity in these regions and plan your downstream filtering accordingly.
Records for the Alignment Stage
For the alignment stage, record:
- The version of BWA-MEM used
- The reference genome version and index files used
- The alignment parameters, including any deviations from defaults
- The number of reads aligned and the alignment rate
- The number of reads with mapping quality below your threshold
- The insert size distribution for paired-end reads
The alignment rate is a key quality metric. For good quality WGS data, you expect an alignment rate above 95 percent. A lower alignment rate suggests contamination, adapter problems, or a reference mismatch.
Post-Alignment Preprocessing
Marking and Removing Duplicate Reads
Duplicate reads are reads that originate from the same DNA fragment during library preparation or sequencing. They are identified by their identical outer coordinates and orientation. Duplicate reads do not represent independent observations of the genome, and they can bias variant allele frequency estimates if not removed.
The standard approach is to mark duplicates instead of remove them. Marking duplicates flags the reads in the BAM file so that downstream tools can ignore them, but the reads remain in the file. This allows you to assess the duplication rate and to re-analyze the data with different duplicate handling if needed.
The duplication rate is an important quality metric. High duplication rates reduce the effective coverage of your sequencing experiment. For WGS data, duplication rates above 20 percent suggest a library preparation problem. For WES data, higher duplication rates are expected because of the targeted amplification steps.
Base Quality Score Recalibration
Base quality score recalibration is a process that adjusts the quality scores assigned by the sequencing instrument based on empirical evidence from the data. The recalibration model accounts for covariates such as the position in the read, the nucleotide context, and the sequencing machine cycle. The goal is to produce quality scores that more accurately reflect the true error rate.
The recalibration process requires a set of known variant sites to use as training data. These known sites are typically obtained from public databases such as dbSNP or the 1000 Genomes Project. The NCBI provides access to these databases and their documentation [<a href="#ref-2">2</a>].
The decision to perform base quality score recalibration depends on your variant caller. GATK HaplotypeCaller is designed to work with recalibrated quality scores, and the GATK best practices recommend recalibration. If you are using a different variant caller, you should check whether recalibration is expected or recommended.
The Impact of Preprocessing on Variant Calling Accuracy
Preprocessing decisions have a direct impact on variant calling accuracy. Duplicate reads that are not marked will inflate the read depth at variant sites, which can lead to false positive calls. Base quality scores that are not recalibrated can lead to systematic errors in variant quality scores, which affects the filtering step.
The literature on clinically responsive genomic analysis pipelines shows that pipeline improvements, including changes to variant calling and quality control processes, can increase diagnostic rates and reduce costs [<a href="#ref-6">6</a>]. This finding underscores the importance of careful preprocessing in clinical applications, where the cost of a missed variant is high.
Records for the Preprocessing Stage
Record the following for the preprocessing stage:
- The version of the duplicate marking tool used
- The duplication rate for each sample
- The version of the base quality score recalibration tool used
- The known variant sites file used for recalibration
- The number of bases recalibrated and the covariates used
These records allow you to assess whether the preprocessing steps were performed correctly and to compare results across samples or experiments.
Variant Calling with GATK HaplotypeCaller
How HaplotypeCaller Works
GATK HaplotypeCaller is a variant caller that uses a local assembly approach. For each region of the genome, it assembles the reads into haplotypes, then compares the haplotypes to the reference to identify variants. This approach is more accurate than simple pileup-based calling because it can handle complex variants and local reassembly.
HaplotypeCaller can be run in two modes: single-sample calling and cohort calling with GVCF files. In single-sample mode, the tool processes one sample at a time and produces a VCF file for that sample. In cohort mode, the tool produces a GVCF file for each sample, and the GVCF files are combined in a joint genotyping step.
GVCF Mode and Joint Genotyping
The GVCF mode is the recommended approach for projects with multiple samples. In GVCF mode, HaplotypeCaller produces a GVCF file that contains the genotype likelihoods for every position in the genome, including positions that are not variant. The GVCF files from all samples are then combined using GenotypeGVCFs, which performs joint genotyping across the cohort.
Joint genotyping has several advantages over single-sample calling. It improves the accuracy of genotype calls at low coverage sites, it allows for the detection of variants that are present in only one sample, and it produces a single VCF file for the entire cohort. The tradeoff is that joint genotyping requires more computational resources and storage.
Variant Calling Parameters and Their Effects
The key parameters for HaplotypeCaller include:
- The minimum base quality score for a base to be considered
- The minimum mapping quality score for a read to be considered
- The ploidy of the sample, which is 2 for diploid germline samples
- The confidence threshold for emitting variant calls
The default parameters are appropriate for most germline variant calling applications. The ploidy parameter is critical: for diploid germline samples, ploidy should be set to 2. If you set the wrong ploidy, the genotype calls will be incorrect.
The Output VCF File Structure
The VCF file produced by HaplotypeCaller contains a header section and a data section. The header describes the file format, the reference genome, the sample names, and the INFO and FORMAT fields. The data section has one line per variant, with the chromosome, position, reference allele, alternate allele, quality score, and filter status.
The quality score in the VCF file is a Phred-scaled probability that the variant is not present in the sample. Higher quality scores indicate higher confidence. The filter status indicates whether the variant passed the filtering criteria or failed one or more filters.
Common Variant Calling Failure Patterns
The most common variant calling failures are false positives in difficult-to-map regions and false negatives in low-coverage regions. False positives occur when reads from paralogous regions are misaligned to the wrong location, creating the appearance of a variant. False negatives occur when the coverage is too low to support a confident genotype call.
Research on variant calling errors has shown that no single pipeline is optimal across the entire genome, and that different genomic contexts have different error profiles [<a href="#ref-1">1</a>]. This means that you should expect some errors in every variant call set, and you should design your filtering strategy to balance sensitivity and precision based on your application.
Records for the Variant Calling Stage
Record the following for the variant calling stage:
- The version of HaplotypeCaller used
- The mode used (single-sample or GVCF)
- The ploidy setting
- The confidence thresholds used
- The number of variants called before filtering
- The transition to transversion ratio for SNVs
The transition to transversion ratio is a useful quality metric. For human WGS data, the expected ratio is approximately 2.0 to 2.1. A significantly lower ratio suggests an excess of false positive calls.
Variant Filtering and Quality Control
The Purpose of Variant Filtering
Variant filtering is the process of separating high-confidence variants from low-confidence variants. The goal is to remove false positives while retaining true variants. The filtering strategy you choose depends on your application. A clinical diagnostic pipeline requires high precision to avoid reporting false variants, while a discovery pipeline may tolerate lower precision to achieve higher sensitivity.
The literature on pipeline optimization emphasizes that the key problem is choosing appropriate tools and selecting the best parameters for optimal precision and recall [<a href="#ref-5">5</a>]. This applies directly to the filtering stage, where the choice of filter thresholds determines the balance between sensitivity and precision.
Hard Filters versus Machine Learning Recalibration
There are two main approaches to variant filtering: hard filters and machine learning recalibration. Hard filters apply fixed thresholds to variant annotations such as quality score, read depth, and mapping quality. Machine learning recalibration uses a training set of known variants to learn the relationship between variant annotations and variant accuracy.
The GATK best practices recommend machine learning recalibration for datasets with a sufficient number of variants, typically more than 30,000 SNVs. For smaller datasets, hard filters are more appropriate. The choice between the two approaches depends on your data size and your tolerance for false positives and false negatives.
Filtering Annotations That Matter
The most informative annotations for filtering germline variants include:
- Quality by depth, which is the quality score divided by the read depth
- Fisher strand bias, which measures the strand bias of the alternate allele
- Mapping quality rank sum, which compares the mapping quality of reads supporting the reference and alternate alleles
- Read position rank sum, which compares the position of the alternate allele within the reads
These annotations capture different types of errors. Strand bias indicates that the alternate allele is supported primarily by reads from one strand, which is a common artifact. Low mapping quality rank sum indicates that the alternate allele is supported by reads with lower mapping quality, which suggests misalignment.
The Sensitivity and Precision Tradeoff
The filtering stage is where you make the most consequential tradeoff decisions. Increasing the stringency of filters reduces false positives but also reduces true positives. Decreasing stringency has the opposite effect. The optimal balance depends on your application.
Research on variant calling errors has shown that different pipelines have different error profiles in different genomic contexts [<a href="#ref-1">1</a>]. This means that the optimal filtering strategy for one genomic region may not be optimal for another. If your study focuses on specific genomic regions, you should evaluate the filtering strategy in those regions specifically.
Records for the Filtering Stage
Record the following for the filtering stage:
- The filtering approach used (hard filters or machine learning recalibration)
- The specific filter thresholds applied
- The number of variants that passed and failed each filter
- The transition to transversion ratio after filtering
- The number of variants in dbSNP and the number of novel variants
The number of variants in dbSNP is a useful quality metric. For human WGS data, you expect the majority of high-confidence variants to be present in dbSNP. A low overlap with dbSNP suggests either a novel population or an excess of false positives.
Reproducibility and Pipeline Documentation
The Importance of Reproducible Workflows
Reproducibility is the ability to recreate the same results from the same input data using the same pipeline. In bioinformatics, reproducibility requires more than saving the commands you ran. It requires capturing the exact versions of all tools, the parameters used, the reference genome version, and the input file versions.
The nf-core documentation provides standards for community pipelines that emphasize reproducibility, including version control, containerization, and automated testing [<a href="#ref-7">7</a>]. These standards are useful even if you are not using nf-core pipelines, because they define what reproducibility means in practice.
Containerization and Version Control
Containerization is the practice of packaging all the tools and their dependencies into a single image that can be run on any system. Containers ensure that the tool versions and their dependencies are identical across different computing environments. This is essential for reproducibility because tool versions can change behavior in subtle ways.
Version control is the practice of tracking changes to your pipeline code and configuration files. The Carpentries lessons provide foundational training in version control with Git, which is the standard tool for this purpose [<a href="#ref-8">8</a>]. Version control allows you to track when and why you changed your pipeline, and it allows others to review your pipeline history.
The Role of Workflow Managers
Workflow managers are tools that automate the execution of multi-step pipelines. They handle the dependencies between steps, the parallelization of jobs, and the tracking of intermediate files. Popular workflow managers include Nextflow, Snakemake, and CWL.
The nf-core documentation describes how community pipelines are built using Nextflow and how they are configured for different computing environments [<a href="#ref-7">7</a>]. Workflow managers are not strictly required for germline variant calling, but they become essential as the number of samples or the complexity of the pipeline increases.
Records for Reproducibility
For full reproducibility, you need to record:
- The exact versions of all tools used in the pipeline
- The parameters used for each tool
- The reference genome version and source
- The input file names and checksums
- The computing environment, including the operating system and hardware
- The workflow manager version and configuration, if used
These records should be stored with the pipeline code and the output files. They allow anyone to reproduce your analysis and to assess whether the results are trustworthy.
Common Failure Patterns and Troubleshooting
Low Alignment Rates
A low alignment rate, typically below 90 percent for WGS data, indicates a problem with the input data or the reference genome. Possible causes include adapter contamination, sample contamination, a mismatch between the sequencing platform and the reference, or a problem with the reference genome files.
The first troubleshooting step is to examine the quality assessment reports for the raw data. If adapter contamination is present, trim the adapters and re-align. If the quality scores are low, consider whether the sequencing run was successful. If the reference genome is the wrong version or build, download the correct reference and re-index.
High Duplication Rates
A high duplication rate, typically above 20 percent for WGS data, indicates over-amplification during library preparation. This reduces the effective coverage of the experiment and can lead to false positive variant calls.
The duplication rate is determined during the duplicate marking step. If the duplication rate is high, you cannot fix the problem by re-analysis. You need to prepare a new library with less amplification. For the current data, you can proceed with the analysis but you should note the reduced effective coverage in your records.
Excessively High or Low Variant Counts
The number of variants called depends on the sequencing platform, the coverage, and the variant caller. For human WGS data, you expect approximately 3 to 4 million SNVs and 500,000 indels per sample. A significantly higher count suggests false positives, and a significantly lower count suggests false negatives.
The transition to transversion ratio is a useful diagnostic. A ratio below 2.0 suggests an excess of false positive calls. A ratio above 2.2 is unusual and may indicate a problem with the reference or the variant caller.
Poor Concordance with Known Variants
If you have a sample with known variants, such as a reference sample or a sample that has been genotyped on an array, you can compare your variant calls to the known variants. Poor concordance indicates a problem with the pipeline.
The concordance rate is the percentage of known variants that are detected by your pipeline. A concordance rate below 95 percent suggests a problem with sensitivity. The discordant variants should be examined to determine whether they are false negatives in your pipeline or errors in the known variant set.
Limitations and Interpretation Boundaries
Germline versus Somatic Variant Calling
The pipeline described in this article is for germline variant calling. Somatic variant calling, which is used in cancer research, requires a different approach. Somatic variants are present only in the tumor tissue, and they must be distinguished from germline variants and sequencing artifacts.
The literature on tumor-only variant calling shows that distinguishing somatic mutations from germline variants is a challenging problem that requires specialized approaches [<a href="#ref-9">9</a>]. If you are working with tumor samples, you should use a somatic variant calling pipeline instead of the germline pipeline described here.
The Impact of Sequencing Platform and Coverage
The pipeline described here is optimized for Illumina sequencing data. Other sequencing platforms, such as Oxford Nanopore or Pacific Biosciences, have different error profiles and require different alignment and variant calling strategies. The literature on variant calling errors has shown that different sequencing platforms have different error profiles in different genomic contexts [<a href="#ref-1">1</a>].
Coverage is another critical factor. The pipeline described here assumes sufficient coverage for confident genotype calls, typically 30x or higher for WGS. At lower coverage, the sensitivity for detecting variants decreases, and the genotype accuracy decreases. If your data has low coverage, you should adjust your expectations and your filtering strategy accordingly.
Difficult Genomic Regions
The pipeline described here has known limitations in difficult genomic regions, including homopolymer stretches, segmental duplications, and regions with high sequence complexity. Research has shown that these regions contribute disproportionately to variant calling errors [<a href="#ref-1">1</a>].
If your study focuses on these regions, you should consider additional approaches, such as graph-based references or targeted sequencing with different technologies. You should also be cautious about interpreting variants in these regions, as they may be false positives.
The Need for Validation
Variant calling pipelines produce predictions, not confirmed variants. The variants identified by your pipeline should be validated using an independent method, such as Sanger sequencing or an orthogonal sequencing platform, before they are used for clinical decisions or high-stakes research conclusions.
The literature on clinically responsive genomic analysis pipelines emphasizes the importance of diagnostic sensitivity and the cost-effectiveness of pipeline improvements [<a href="#ref-6">6</a>]. In clinical applications, the cost of a false negative is high, and validation is essential.
Professional Escalation Criteria
When to Seek Expert Help
You should consider seeking expert help in the following situations:
- You are working with a non-standard sequencing platform or library preparation method
- Your data has unusual quality metrics that you cannot explain
- You are setting up a pipeline for clinical diagnostic use
- You are working with difficult genomic regions and need specialized approaches
- Your variant calling results are inconsistent with known biology or previous experiments
The EMBL-EBI training resources provide learning pathways for bioinformatics that can help you build the skills needed to troubleshoot your pipeline [<a href="#ref-3">3</a>]. The Galaxy Training Network provides accessible tutorials for specific analysis steps [<a href="#ref-4">4</a>]. These resources are useful for building your skills, but they do not replace expert consultation for complex problems.
The Role of Community Resources
Community resources can help you troubleshoot your pipeline. The Bioconductor project provides packages for genomic analysis and reproducible research, with documentation and support forums [<a href="#ref-10">10</a>]. The nf-core community provides pipelines and documentation that follow community standards [<a href="#ref-7">7</a>].
When you encounter a problem, search the documentation and support forums for the specific tools you are using. Many common problems have been encountered and solved by other users. If you cannot find a solution, post a question with a minimal reproducible example that includes your input data, your commands, and your error messages.
Documentation for Escalation
When you escalate a problem, provide the following documentation:
- The exact versions of all tools used
- The input data files and their checksums
- The commands you ran and the parameters you used
- The error messages or unexpected results
- The quality metrics for each pipeline stage
This documentation allows the expert to reproduce your problem and to identify the cause. Without this documentation, the expert cannot help you effectively.
Frequently Asked Questions
What is the difference between germline and somatic variant calling?
Germline variant calling identifies variants that are inherited and present in all cells of the body. Somatic variant calling identifies variants that arise during a person's lifetime and are present only in specific tissues, such as tumors. The two approaches require different experimental designs and different bioinformatics pipelines. Germline variant calling typically uses a single sample, while somatic variant calling requires a matched normal sample to distinguish somatic mutations from germline variants [<a href="#ref-9">9</a>].
Why is BWA-MEM the recommended aligner for germline variant calling?
BWA-MEM is widely used because it produces accurate alignments for Illumina sequencing data and it is well integrated with the GATK variant calling tools. The choice of aligner is one of the most consequential decisions in the pipeline, and different aligners have different strengths and weaknesses in different genomic contexts [<a href="#ref-5">5</a>]. BWA-MEM is a reasonable default for standard germline variant calling applications.
What is the purpose of base quality score recalibration?
Base quality score recalibration adjusts the quality scores assigned by the sequencing instrument to more accurately reflect the true error rate. The recalibration model accounts for covariates such as read position, nucleotide context, and sequencing machine cycle. The recalibrated quality scores are used by the variant caller to compute variant quality scores, and inaccurate quality scores lead to inaccurate variant quality scores.
How many variants should I expect from a human whole genome sequencing sample?
For human WGS data, you expect approximately 3 to 4 million single nucleotide variants and 500,000 small insertions or deletions per sample. The exact number depends on the population ancestry of the sample and the sequencing platform. The transition to transversion ratio is a useful quality metric, with an expected value of approximately 2.0 to 2.1 for human data.
What is the difference between hard filters and machine learning recalibration for variant filtering?
Hard filters apply fixed thresholds to variant annotations such as quality score, read depth, and mapping quality. Machine learning recalibration uses a training set of known variants to learn the relationship between variant annotations and variant accuracy. Machine learning recalibration is recommended for datasets with a sufficient number of variants, typically more than 30,000 SNVs. For smaller datasets, hard filters are more appropriate.
Why is reproducibility important in variant calling pipelines?
Reproducibility allows you and others to recreate the same results from the same input data using the same pipeline. This is essential for scientific validity and for clinical applications where results may be used for patient care decisions. Reproducibility requires capturing the exact versions of all tools, the parameters used, the reference genome version, and the input file versions [<a href="#ref-7">7</a>].
What should I do if my variant calling results have a low transition to transversion ratio?
A transition to transversion ratio below 2.0 for human WGS data suggests an excess of false positive calls. You should examine the variants that are likely false positives, such as those with low quality scores, low read depth, or strand bias. You should also check the quality metrics from earlier pipeline stages to identify potential sources of error.
When should I use a workflow manager for my variant calling pipeline?
A workflow manager is useful when you have multiple samples, multiple pipeline steps, or a need to track intermediate files and parameters. Workflow managers automate the execution of multi-step pipelines and handle dependencies between steps. The nf-core documentation provides standards for community pipelines that emphasize reproducibility and automated testing [<a href="#ref-7">7</a>]. For a single sample analysis, a workflow manager may be unnecessary, but it becomes essential as the scale of the analysis increases.
Related Bioinformatics Guides
- Genomic Data Processing: From Raw Sequencing to Analysis-Ready Files
- Single-Cell Sequencing Analysis Pipeline: From Raw Data to Biological Insights
- RNA Sequencing Data Analysis: From Raw Reads to Differential Expression
- Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations
- Genomic Data Analysis Tools: A Comparative Guide for Researchers
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
[1] [StratoMod: predicting sequencing and variant calling errors with interpretable machine learning.](https://pubmed.ncbi.nlm.nih.gov/39397114). Communications biology, 2024. [2] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [3] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [4] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [5] [ToTem: a tool for variant calling pipeline optimization.](https://pubmed.ncbi.nlm.nih.gov/29940847). BMC bioinformatics, 2018. [6] [Clinically Responsive Genomic Analysis Pipelines: Elements to Improve Detection Rate and Efficiency.](https://pubmed.ncbi.nlm.nih.gov/33962052). The Journal of molecular diagnostics : JMD, 2021. [7] [nf-core Documentation](https://nf-co.re/docs). nf-core. [8] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [9] [Fast, accurate, and racially unbiased pan-cancer tumor-only variant calling with tabular machine learning.](https://pubmed.ncbi.nlm.nih.gov/36611079). NPJ precision oncology, 2023. [10] [Bioconductor](https://bioconductor.org/). Bioconductor Project.This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.