Quality Control and Preprocessing for Single-Cell Multi-Omics: Best Practices for ATAC + RNA and Methylation + RNA

By Dr. Zubair Khalid, DVM, MS, PhD ·

Quality Control and Preprocessing for Single-Cell Multi-Omics: Best Practices for ATAC + RNA and Methylation + RNA

Key Takeaways

  • Joint quality control for single-cell multi-omics (ATAC+RNA, Methylation+RNA) necessitates integrated metrics, as cells passing individual modality thresholds may fail others; filtering decisions must account for joint information rather than independent cutoffs.
  • Critical ATAC+RNA quality metrics include fraction of reads in peaks and transcription start site (TSS) enrichment score, alongside RNA metrics like UMI counts and mitochondrial percentage, to identify nuclei with degraded chromatin or poor transcript capture.
  • For Methylation+RNA, bisulfite conversion efficiency is paramount, as incomplete conversion leads to false methylation calls, distorting downstream analyses; this is assessed alongside CpG coverage and RNA metrics.
  • A tiered cell classification system (Tier 1: high confidence, Tier 2: borderline/requires review, Tier 3: excluded) based on joint metric profiles and a weighted joint quality score provides a more nuanced approach than binary pass/fail filtering.
  • Reproducibility mandates meticulous documentation of all filtering decisions, including metric distributions, applied thresholds, rationale, and the number of cells removed at each step, often supported by computational notebooks or scripts.
  • Common failure patterns include applying RNA-only thresholds, ignoring modality-specific batch effects, over-filtering, incomplete documentation, and neglecting the biological relationship between modalities (e.g., gene expression vs. promoter accessibility).

Single-cell multi-omics experiments that jointly profile chromatin accessibility and gene expression (ATAC + RNA) or DNA methylation and gene expression (Methylation + RNA) require quality control procedures that differ substantially from those used for single-modality datasets. The central problem is that a cell passing quality thresholds for one modality may fail for another, and the filtering decision must account for joint information instead of independent modality-specific cutoffs. This article provides concrete quality control workflows, modality-specific metrics, filtering strategies, and documentation practices for researchers processing joint single-cell datasets.

Scope and Reader Context

This guidance applies to researchers who have already generated or acquired single-cell multi-omics data and need to perform quality control before downstream analysis such as clustering, trajectory inference, or integration with other datasets. The focus is on two common joint profiling combinations: ATAC + RNA (often produced by commercial Multiome platforms or combined snRNA-seq and snATAC-seq experiments) and Methylation + RNA (produced by protocols such as single-cell nucleosome methylation and transcription sequencing, scNMT). The quality control principles described here also apply to related multi-modal designs that combine gene expression with protein abundance or immune receptor repertoires, though the specific metrics differ.

The intended reader is a graduate student, postdoctoral researcher, laboratory scientist, or bioinformatics practitioner who understands basic single-cell RNA sequencing analysis but has not yet developed a systematic approach to joint modality quality control. The practical outcome is a defensible filtering strategy that produces a high-confidence cell population for downstream analysis, with clear records of decisions made and metrics tracked.

Why Joint Quality Control Differs from Single-Modality Analysis

Single-cell RNA sequencing quality control typically evaluates three core metrics: total unique molecular identifiers (UMIs) per cell, number of genes detected per cell, and percentage of reads mapping to mitochondrial genes. Cells with very low counts are presumed to be empty droplets or damaged cells, while cells with very high counts may be doublets. This framework has been extensively documented in bioinformatics workflows and training materials from the Galaxy Training Network and Bioconductor project.

Joint profiling introduces additional complexity because each modality has its own failure modes. A nucleus may have excellent RNA content but degraded chromatin accessibility, or vice versa. The BD Rhapsody system documentation describes how multimodal capture requires quality control at multiple stages, including imaging-based assessment of cell viability and multiplet rates before sequencing. The imaging step provides an independent check that sequencing-based metrics cannot fully replace.

For ATAC + RNA data, the key additional metrics include the fraction of reads in peaks, transcription start site (TSS) enrichment score, and the ratio of reads mapping to promoter regions versus other genomic features. For Methylation + RNA data, bisulfite conversion efficiency becomes a critical metric because incomplete conversion produces false methylation calls that can distort downstream analysis.

The practical consequence is that filtering on RNA metrics alone will retain cells with poor chromatin data, and filtering on chromatin metrics alone will retain cells with poor transcriptome data. Joint filtering requires either intersection approaches, where a cell must pass both modality thresholds, or union approaches with explicit documentation of why a cell with one failing modality was retained.

Core Principles for Multi-Omics Quality Control

Principle 1: Define the Biological Question Before Setting Thresholds

Quality control thresholds are not universal constants. The appropriate minimum number of genes per cell depends on whether the study aims to identify rare cell populations, characterize developmental trajectories, or compare cell states across conditions. A study focused on abundant, transcriptionally distinct cell types can use more aggressive filtering than a study attempting to capture rare progenitor populations.

The same logic applies to chromatin accessibility metrics. A study of enhancer regulation may require higher TSS enrichment because promoter-proximal signal is the focus, while a study of transposable element expression may tolerate lower TSS scores because the biological signal of interest is elsewhere in the genome.

Principle 2: Track Metrics Before and After Filtering

Quality control is not a single step but a continuous process. The distribution of each metric should be examined before filtering, after initial threshold application, and after any additional filtering based on joint criteria. This documentation allows researchers to assess whether filtering introduced bias and provides a record for reproducibility.

The scQCEA framework demonstrates this approach by generating interactive quality control reports that compare metrics across samples and visualize quality scores. The framework also includes expression-based quality control that distinguishes true biological variation from background noise, which is particularly relevant for multi-omics data where technical noise can differ between modalities.

Principle 3: Use Multiple Metrics Instead of a Single Threshold

No single quality metric captures all failure modes. A cell with high UMI counts but low TSS enrichment may represent a nucleus with intact RNA but degraded chromatin. A cell with high TSS enrichment but low gene counts may represent a nucleus with intact chromatin but failed reverse transcription. Joint evaluation of multiple metrics provides more information than any individual threshold.

Principle 4: Document All Filtering Decisions

Reproducibility requires that every filtering decision be recorded, including the metric distributions that motivated the choice, the exact thresholds applied, and the number of cells removed at each step. This documentation should be included in the methods section of any manuscript and should be available as a computational notebook or script. The nf-core documentation describes community standards for reproducible bioinformatics pipelines, including configuration and usage practices that support documentation and reproducibility.

At a Glance: Modality-Specific Quality Metrics and Filtering Strategies

Modality PairPrimary Quality MetricsCommon Filtering ApproachKey Failure Mode Addressed
ATAC + RNARNA: UMI counts, gene counts, mitochondrial percentage. ATAC: TSS enrichment, fraction reads in peaks, nucleosome banding patternIntersection filtering with modality-specific thresholds, followed by joint outlier detectionRetaining cells with intact RNA but degraded chromatin, or vice versa
Methylation + RNARNA: UMI counts, gene counts, mitochondrial percentage. Methylation: bisulfite conversion rate, CpG coverage, methylation call rateSequential filtering with methylation-specific thresholds applied after RNA filteringFalse methylation calls from incomplete bisulfite conversion
RNA + Protein (CITE-seq or similar)RNA: UMI counts, gene counts. Protein: antibody capture counts, number of protein features detectedThreshold on RNA metrics first, then evaluate protein detection ratesCells with failed antibody staining that appear normal by RNA metrics

Modality-Specific Quality Metrics for ATAC + RNA

RNA Modality Metrics

The RNA component of a joint ATAC + RNA experiment should be evaluated using the same core metrics applied to standard single-cell RNA sequencing data. These include total UMI counts per cell, number of genes detected, and the fraction of reads mapping to mitochondrial genes. The guidelines for single-cell sequencing analysis in Alzheimer's disease research describe quality control and normalization as the first major analytical direction, with specific attention to the challenges of human postmortem tissue where RNA quality varies substantially between samples.

For single-nucleus RNA sequencing, which is the typical RNA component in joint ATAC + RNA experiments, mitochondrial read fraction is often less informative than in whole-cell preparations because mitochondria are largely excluded from nuclei. However, the metric should still be tracked because high mitochondrial reads in a nuclear preparation can indicate cytoplasmic contamination or cell lysis during isolation. The plant single-cell and spatial transcriptomics review notes that single-nucleus RNA sequencing improves representation of recalcitrant lineages and reduces stress signatures while remaining compatible with multi-omics profiling, but also requires attention to organellar and intronic metrics.

ATAC Modality Metrics

The ATAC component requires metrics that assess chromatin accessibility quality. The transcription start site enrichment score measures the accumulation of reads at promoter regions relative to flanking regions. A high TSS enrichment indicates that the transposition reaction preferentially targeted open chromatin at promoters, which is the expected pattern for successful ATAC-seq.

The fraction of reads in peaks measures the proportion of sequencing reads that fall within called peaks. This metric reflects the signal-to-noise ratio of the experiment. Low values indicate either excessive background signal or a transposition reaction that did not specifically target open chromatin.

The nucleosome banding pattern, assessed by fragment size distribution, provides information about the periodicity of nucleosome positioning. A successful ATAC-seq experiment shows a strong mononucleosomal peak at approximately 150 to 200 base pairs, with additional peaks at multiples of this size. The mitochondrial single-cell ATAC-seq protocol describes quality control metrics for chromatin accessibility profiling, including considerations for extending the approach to other mammalian tissues.

Joint Metrics for ATAC + RNA

The relationship between RNA expression and chromatin accessibility at specific loci provides an additional quality check. For genes with strong expression, the corresponding promoter should generally show accessible chromatin. Systematic mismatches between expression and accessibility can indicate technical problems in one modality.

The lineage tracing study using combined snRNA and ATAC sequencing in kidney injury models demonstrates how joint profiling can identify transcription factor activity that is not apparent from either modality alone. The study generated a final dataset of 83,315 high-quality nuclei after quality control and doublet removal, illustrating that joint filtering can retain sufficient cell numbers for complex downstream analyses.

Modality-Specific Quality Metrics for Methylation + RNA

Methylation Modality Metrics

Bisulfite conversion efficiency is the most critical quality metric for methylation profiling. Incomplete conversion results in false methylation calls at unmethylated cytosines, which can substantially distort downstream analysis. Conversion efficiency is typically assessed using spike-in controls or by measuring methylation at known unmethylated regions.

CpG coverage per cell reflects the depth of methylation information available. Low coverage limits the ability to assess methylation at specific loci and increases the noise in global methylation estimates. The scNMT protocol update for multi-omics profiling of mouse brain and pancreatic organoids describes the experimental workflow for joint methylation and expression profiling, with quality control considerations embedded in the protocol.

Joint Metrics for Methylation + RNA

The relationship between promoter methylation and gene expression provides a joint quality metric. In general, high promoter methylation is associated with reduced gene expression, though this relationship is context-dependent and not universal. Systematic violations of this expected pattern can indicate technical problems in either modality.

The osteoarthritis genomics study used single-cell multi-omics data to identify signal enrichment in embryonic skeletal development pathways, integrating transcriptome, proteome, and epigenome profiles. This study demonstrates how joint profiling can reveal regulatory relationships that are not apparent from expression data alone, but the quality of these inferences depends on the reliability of both modalities.

Practical Workflow for Joint Quality Control

Step 1: Generate Per-Cell Metric Matrices

The first step is to compute quality metrics for each cell in each modality. For the RNA modality, this includes UMI counts, gene counts, and mitochondrial read fraction. For the ATAC modality, this includes TSS enrichment, fraction of reads in peaks, and fragment size distribution. For the methylation modality, this includes bisulfite conversion rate, CpG coverage, and methylation call rate.

These metrics should be computed using the standard tools for each modality. The Bioconductor project provides packages for single-cell analysis that include quality control functions, and the Galaxy Training Network offers accessible tutorials for implementing these workflows.

Step 2: Examine Metric Distributions

Before applying any thresholds, examine the distribution of each metric across all cells. Histograms and violin plots are useful for identifying outliers and understanding the overall quality of the experiment. Cells with extreme values for any metric should be flagged for closer examination.

The scQCEA framework provides interactive visualization of quality metrics across samples, allowing researchers to compare quality distributions between batches or conditions. This comparison is particularly important for multi-omics data where batch effects can differ between modalities.

Step 3: Apply Modality-Specific Thresholds

Apply initial thresholds for each modality independently. The specific thresholds will depend on the data quality and the biological question, but common approaches include:

For RNA: Remove cells with fewer than a minimum number of UMIs or genes, and remove cells with very high mitochondrial read fractions.

For ATAC: Remove cells with low TSS enrichment scores and low fractions of reads in peaks.

For Methylation: Remove cells with low bisulfite conversion rates and low CpG coverage.

The key decision is whether to use intersection filtering, where a cell must pass all modality thresholds, or union filtering, where a cell passes if it meets criteria for at least one modality. Intersection filtering is more conservative and produces a higher-confidence cell population, but it may remove cells that are biologically informative in one modality. Union filtering retains more cells but requires careful documentation of why cells with one failing modality were retained.

Step 4: Apply Joint Outlier Detection

After initial thresholding, apply joint outlier detection to identify cells that pass individual thresholds but show unusual combinations of metrics across modalities. For example, a cell with very high RNA quality but very low ATAC quality may represent a technical artifact instead of a genuine biological state.

The cross-platform transcriptomic data integration study in osteoarthritis synovium demonstrates the importance of quality control before statistical integration. The study retained 139 of 153 subjects after quality control, showing that even well-designed studies require careful filtering before downstream analysis.

Step 5: Assess Doublet Rates

Doublets, where two cells are captured in the same droplet or well, are a major source of noise in single-cell experiments. The BD Rhapsody system includes imaging-based quality control for viability and multiplet detection before sequencing, providing an independent assessment that complements computational doublet detection.

For joint ATAC + RNA data, doublet detection can leverage information from both modalities. A doublet may show unusually high RNA counts combined with high ATAC signal, or it may show a mixture of cell type markers from different lineages. The microfluidic biochip review describes how single-cell isolation technologies incorporate quality control features to ensure reliable cell isolation for downstream multi-omics analysis.

Step 6: Document and Report

Record all filtering decisions, including the metric distributions examined, the thresholds applied, and the number of cells removed at each step. This documentation should be included in the methods section of any manuscript and should be available as a computational notebook or script.

The nf-core documentation describes community standards for reproducible bioinformatics pipelines, including configuration and usage practices that support documentation and reproducibility. Adopting these standards for quality control workflows ensures that filtering decisions are transparent and reproducible.

Options and Tradeoffs in Filtering Strategies

Intersection Filtering

Intersection filtering requires that a cell pass quality thresholds for all modalities. This approach produces the highest-confidence cell population and is appropriate when the downstream analysis requires reliable data from all modalities. The tradeoff is that cells with poor quality in one modality are removed even if they contain valuable information in another modality.

Union Filtering

Union filtering retains cells that pass quality thresholds for at least one modality. This approach maximizes cell retention and is appropriate when the downstream analysis can use information from either modality independently. The tradeoff is that cells with poor quality in one modality may introduce noise into joint analyses.

Sequential Filtering

Sequential filtering applies thresholds in a specific order, typically starting with the modality that has the most reliable quality metrics. For example, RNA quality thresholds might be applied first, followed by ATAC or methylation thresholds on the remaining cells. This approach is practical when one modality is more established or when the primary biological question depends more heavily on one modality.

Weighted Approaches

Weighted approaches combine quality metrics from multiple modalities into a single score, with weights reflecting the relative importance of each modality for the biological question. This approach is more flexible than threshold-based filtering but requires careful calibration and documentation.

The single-cell multi-omics analysis for drug discovery describes how joint profiling can inform drug target identification, but the quality of these inferences depends on the filtering strategy. The choice of filtering approach should be guided by the specific biological question and the tolerance for false positives versus false negatives.

Records and Measurements for Quality Control

Required Records

Maintain the following records for each experiment:

  1. Sample metadata, including tissue source, preparation date, and any experimental conditions
  2. Sequencing statistics, including total reads, read length, and sequencing platform
  3. Per-cell quality metrics for each modality, computed before and after filtering
  4. Filtering thresholds applied and the rationale for each threshold
  5. Number of cells removed at each filtering step
  6. Final cell counts and the distribution of quality metrics in the retained population

Recommended Measurements

In addition to the required records, the following measurements provide useful context for quality control decisions:

  1. Comparison of quality metrics across batches or conditions to identify systematic differences
  2. Assessment of marker gene expression or accessibility to confirm that expected cell types are present
  3. Evaluation of the relationship between modalities, such as the correlation between expression and accessibility at specific loci
  4. Sensitivity analysis to assess how downstream results change with different filtering thresholds

The guidelines for single-cell sequencing analysis in Alzheimer's disease research emphasize the importance of quality control and normalization as the foundation for all downstream analyses. The study implemented recommended workflows for each major analytical direction and applied them to a large single-nucleus RNA sequencing dataset, providing a model for systematic quality control.

Common Failure Patterns in Multi-Omics Quality Control

Failure Pattern 1: Applying RNA-Only Thresholds to Joint Data

Researchers familiar with single-cell RNA sequencing analysis often apply the same thresholds to the RNA component of a joint dataset without considering the ATAC or methylation components. This approach can retain cells with poor chromatin or methylation data, introducing noise into joint analyses.

Failure Pattern 2: Ignoring Modality-Specific Batch Effects

Batch effects can differ between modalities. A batch effect that affects RNA quality may not affect ATAC quality, and vice versa. Ignoring modality-specific batch effects can lead to incorrect conclusions about biological differences between conditions.

Failure Pattern 3: Over-Filtering Based on Joint Metrics

Joint filtering can remove large numbers of cells if thresholds are too aggressive. The kidney injury lineage tracing study retained 83,315 high-quality nuclei after quality control and doublet removal, but this required careful optimization of filtering thresholds. Over-filtering can remove rare cell populations and reduce statistical power for downstream analyses.

Failure Pattern 4: Incomplete Documentation of Filtering Decisions

Failure to document filtering decisions makes it impossible to reproduce the analysis or to assess whether filtering introduced bias. This is particularly problematic for multi-omics data where the filtering strategy can substantially affect downstream results.

Failure Pattern 5: Ignoring the Relationship Between Modalities

The relationship between modalities provides important quality information. For example, cells with high expression of a gene but no chromatin accessibility at the promoter may indicate a technical problem. Ignoring these relationships can allow technical artifacts to pass quality control.

Failure Pattern 6: Inadequate Handling of Ambient RNA Contamination

Ambient RNA from lysed cells can contaminate droplets or wells, producing spurious expression signals that pass standard quality filters. The plant single-cell and spatial transcriptomics review provides practical guidance for mitigating ambient RNA and interpreting organellar and intronic metrics, which is directly applicable to joint multi-omics datasets where contamination can affect both modalities differently.

Limitations of Current Quality Control Approaches

Threshold Selection Remains Subjective

Despite the availability of computational tools for quality control, the selection of specific thresholds remains a subjective decision that depends on the biological question and the quality of the data. Different researchers analyzing the same dataset may make different filtering decisions, leading to different downstream results.

Doublet Detection Is Imperfect

Computational doublet detection methods have limited accuracy, particularly for cells from similar lineages or for doublets involving cells with similar transcriptional profiles. The imaging-based quality control available on some platforms provides additional information but is not available for all experimental designs.

Quality Metrics Do Not Capture All Failure Modes

The standard quality metrics for each modality capture common failure modes but do not capture all possible technical problems. For example, batch effects within a sample, index hopping during sequencing, or contamination from other species may not be detected by standard quality metrics.

Joint Metrics Are Still Evolving

The field of single-cell multi-omics is relatively young, and the optimal approaches for joint quality control are still being developed. The scNMT protocol update and the mitochondrial single-cell ATAC-seq protocol represent recent efforts to establish standards for specific joint profiling approaches, but consensus guidelines for joint quality control across all platforms are not yet available.

Platform-Specific Considerations

Different commercial platforms have different quality characteristics. The BD Rhapsody system uses microwell-based cartridges with imaging-based quality control, while droplet-based platforms rely more heavily on computational metrics. Researchers should understand the specific quality control features of their platform and adjust their workflows accordingly.

Safety and Regulatory Context

Data Management and Privacy

Single-cell multi-omics data from human samples may contain sensitive information that requires careful data management. The NCBI Data Resources provide official descriptions of database systems and search resources that support secure data deposition and access. Researchers should follow institutional and national guidelines for data sharing and privacy protection.

Reproducibility Standards

Funding agencies and journals increasingly require reproducible analysis workflows. The nf-core documentation describes community standards for pipeline development and configuration that support reproducibility. The Carpentries lessons provide foundational training in computing and data skills that support reproducible research practices.

Ethical Considerations

Single-cell multi-omics data can reveal information about cell types, developmental states, and disease mechanisms that may have ethical implications. Researchers should consider the potential uses and misuses of their data and follow institutional review board requirements for human subjects research.

Professional Escalation Criteria

When to Seek Additional Expertise

Consult with a bioinformatics specialist or the platform manufacturer when:

  1. Quality metrics show unexpected patterns that cannot be explained by standard failure modes
  2. The relationship between modalities is systematically disrupted across many cells
  3. Batch effects differ substantially between modalities
  4. The filtering strategy removes an unexpectedly large fraction of cells
  5. Downstream analyses produce results that are inconsistent with known biology

When to Consider Repeating the Experiment

Consider repeating the experiment when:

  1. The fraction of cells passing quality control is very low, indicating a fundamental problem with the experimental protocol
  2. Quality metrics are consistently poor across all cells, suggesting a problem with reagents or sequencing
  3. The relationship between modalities is disrupted in a way that cannot be corrected by filtering
  4. The biological question requires high-quality data from both modalities, and one modality is consistently failing

A Decision Framework for Joint Modality Filtering with Escalation Triggers

The practical challenge in multi-omics quality control is not computing metrics but deciding what to do when metrics conflict within the same cell. A nucleus with excellent RNA complexity but marginal TSS enrichment presents a decision that threshold-based workflows do not resolve. This section provides a structured decision framework that assigns cells to explicit quality tiers, defines escalation triggers for ambiguous cases, and links each decision to a documentation record that can be audited during manuscript review or data deposition.

Tiered Cell Classification Instead of Binary Pass or Fail

Binary filtering, where a cell either passes or fails, forces every cell into one of two categories. This approach hides the intermediate states that are common in joint profiling experiments. A more useful framework assigns each cell to one of three quality tiers based on its joint metric profile.

Tier 1 cells pass all modality-specific thresholds and show consistent quality across modalities. These cells enter downstream analysis without further review. Tier 2 cells pass thresholds for one modality but show borderline values for another, or they pass all thresholds but show an unusual cross-modality relationship. These cells require individual review before inclusion. Tier 3 cells fail thresholds for one or more modalities and are excluded unless a documented biological justification exists for retention.

The tiered approach addresses a core limitation of intersection and union filtering. Intersection filtering removes Tier 2 cells that may contain valuable biological information. Union filtering retains Tier 3 cells that introduce technical noise. Tiered classification preserves the analytical decision for the researcher instead of hiding it in a binary threshold.

Assignment Criteria for Each Tier

Tier 1 assignment requires that a cell meet all of the following conditions. For the RNA modality, the cell must exceed the minimum UMI count, minimum gene count, and fall below the maximum mitochondrial read fraction established for the experiment. For the ATAC modality, the cell must exceed the minimum TSS enrichment score and minimum fraction of reads in peaks. For the methylation modality, the cell must exceed the minimum bisulfite conversion rate and minimum CpG coverage. Additionally, the cell must not be flagged as a doublet by computational detection methods.

Tier 2 assignment applies when a cell meets thresholds for at least one modality but shows one of the following patterns. The cell falls within a defined buffer zone below the threshold for one modality, typically within 10 to 20 percent of the cutoff value. The cell passes all thresholds but shows a cross-modality inconsistency, such as high expression of a gene with no corresponding chromatin accessibility at its promoter. The cell passes all thresholds but has a metric value that falls more than three median absolute deviations from the population median for any single metric.

Tier 3 assignment applies when a cell fails thresholds for more than one modality, fails a single modality by a wide margin, or is flagged as a doublet with supporting evidence from both modalities. The mitochondrial single-cell ATAC-seq protocol describes how joint profiling can identify cells with discordant mitochondrial genotype and chromatin accessibility, which would fall into Tier 3 when the discordance reflects technical artifact instead of biological variation.

The Joint Quality Score as a Decision Aid

A joint quality score provides a single numerical value that summarizes quality across modalities and supports tier assignment. The score is computed by normalizing each modality-specific metric to a common scale, then combining the normalized values with weights that reflect the biological importance of each modality for the specific study.

For ATAC + RNA data, a practical scoring approach normalizes RNA gene counts and ATAC TSS enrichment to z-scores across all cells, then computes a weighted sum. The weights should be set based on the biological question. A study focused on gene regulatory networks might weight ATAC metrics at 0.6 and RNA metrics at 0.4, while a study focused on cell type identification might use equal weights. The guidelines for single-cell sequencing analysis in Alzheimer's disease research describe quality control and normalization as the foundation for 14 major analytical directions, and the weighting decision should be documented as part of this foundation.

For Methylation + RNA data, the joint score should incorporate bisulfite conversion rate as a multiplicative factor instead of an additive term. A cell with 95 percent conversion efficiency has a substantially different error profile than a cell with 99 percent conversion, and this difference compounds across the thousands of CpG sites measured per cell. The scNMT protocol update for multi-omics profiling of mouse brain and pancreatic organoids describes the experimental workflow where conversion efficiency is a primary quality determinant.

The joint quality score does not replace modality-specific thresholds. It provides a supplementary view that helps identify cells where the combination of metrics is unusual even when individual metrics pass. A cell with a high joint score but a single failing metric should be reviewed manually instead of automatically included or excluded.

Escalation Triggers for Manual Review

Manual review is required when a cell falls into Tier 2 or when the joint quality score conflicts with tier assignment based on individual thresholds. The following triggers should prompt individual cell examination.

Trigger 1 applies when a cell passes all thresholds but has a joint quality score in the bottom 5 percent of the population. This pattern suggests that the cell has acceptable individual metrics but poor overall quality, which may indicate a technical artifact that individual thresholds do not capture.

Trigger 2 applies when a cell shows a strong cross-modality inconsistency. For ATAC + RNA data, this means high expression of a gene with no chromatin accessibility at the promoter, or accessible chromatin at a promoter with no corresponding expression. For Methylation + RNA data, this means low promoter methylation with low gene expression, or high promoter methylation with high gene expression. The osteoarthritis genomics study used single-cell multi-omics data to identify signal enrichment in developmental pathways, and the authors would have encountered such inconsistencies during quality control of their 489,975 case dataset.

Trigger 3 applies when a cell is flagged as a doublet by one computational method but not by another. The BD Rhapsody system documentation describes imaging-based multiplet detection that provides an independent assessment before sequencing. When imaging and computational methods disagree, the cell should be reviewed manually with attention to whether the marker gene profile suggests a mixture of cell types.

Trigger 4 applies when a cell passes all thresholds but comes from a batch or condition where the overall quality metrics are substantially lower than other batches. The cross-platform transcriptomic data integration study in osteoarthritis synovium retained 139 of 153 subjects after quality control, demonstrating that sample-level quality variation requires attention even when individual cells pass thresholds.

Documentation Records for Each Decision

Each tier assignment and escalation decision requires a documentation record that can be audited. The record should include the cell identifier, the values for all modality-specific metrics, the tier assignment, the trigger that prompted review if applicable, the decision made, and the rationale for that decision.

A practical documentation format is a table with one row per cell and columns for each metric, the tier assignment, and a decision code. The decision code should be one of a fixed set: include, exclude, include with caveat, or exclude with note. The caveat and note fields should contain free text explaining the rationale.

The scQCEA framework provides an example of systematic quality documentation. The framework generates interactive reports of quality metrics for multi-omics data and includes automated cell type annotation for expression-based quality control. Adopting a similar reporting structure ensures that quality decisions are transparent and reproducible.

The nf-core documentation describes community standards for reproducible bioinformatics pipelines, including configuration and usage practices. The documentation records for tier assignments should be integrated into the pipeline configuration so that they are generated automatically and stored with the analysis outputs.

Common Failure Patterns in Tier Assignment

The tiered framework fails in predictable ways when implemented without attention to context. The most common failure is setting buffer zones too wide, which moves a large fraction of cells into Tier 2 and creates an impractical manual review burden. Buffer zones should be set based on the observed metric distributions, not on arbitrary percentages. If more than 20 percent of cells fall into Tier 2, the buffer zones are too wide or the primary thresholds are poorly calibrated.

A second failure pattern is applying the same tier assignment criteria across batches without checking whether metric distributions differ between batches. The plant single-cell and spatial transcriptomics review notes that single-nucleus RNA sequencing reduces stress signatures compared to protoplast-based methods, but batch effects can still differ between modalities. Tier assignment criteria should be evaluated separately for each batch, and the evaluation should be documented.

A third failure pattern is treating the joint quality score as a substitute for modality-specific thresholds. The score is a decision aid, not a replacement for examining individual metrics. A cell with a high joint score can still have a failing metric that indicates a specific technical problem.

A fourth failure pattern is inconsistent application of escalation triggers. If one researcher reviews Tier 2 cells and another does not, the resulting datasets are not comparable. Escalation triggers should be applied uniformly, and the application should be documented in the analysis script.

Practical Implementation Steps

Implement the tiered framework in the following sequence. First, compute all modality-specific metrics for every cell using the standard tools for each modality. The Bioconductor project provides packages for single-cell analysis that include quality control functions, and the Galaxy Training Network offers accessible tutorials for implementing these workflows.

Second, examine the distribution of each metric across all cells and across batches. Record the median, interquartile range, and outlier boundaries for each metric. This examination informs the primary thresholds and the buffer zone widths.

Third, assign primary thresholds based on the metric distributions and the biological question. Document the rationale for each threshold choice. The guidelines for single-cell sequencing analysis in Alzheimer's disease research describe quality control and normalization as the first major analytical direction, and the threshold rationale should be recorded with the same rigor as the analytical steps.

Fourth, assign each cell to a tier based on the primary thresholds and buffer zones. Generate the tier assignment table with all metric values and decision codes.

Fifth, apply escalation triggers to identify Tier 2 cells requiring manual review and Tier 3 cells requiring documented justification for retention. Perform the manual review and record the decisions.

Sixth, assess the impact of the tiered filtering on downstream analysis. Compare clustering results, cell type proportions, and marker gene detection with and without Tier 2 cells. This sensitivity analysis provides evidence that the filtering decisions do not drive the biological conclusions.

Seventh, document the complete workflow in a computational notebook or script that can be rerun on the raw data. The Carpentries lessons provide foundational training in computing and data skills that support reproducible research practices, and the documentation should follow these standards.

Records and Measurements for Tier Assignment

Maintain the following records for each experiment. The per-cell metric matrix with all modality-specific values. The tier assignment table with decision codes and rationale. The escalation trigger log showing which cells triggered manual review and the outcome of each review. The sensitivity analysis results comparing downstream outcomes with and without Tier 2 cells. The batch-level metric summaries showing how quality distributions vary across batches.

The NCBI Data Resources provide official descriptions of database systems that support secure data deposition. When depositing multi-omics data, include the tier assignment tables and quality control documentation as supplementary files so that reviewers and readers can assess the filtering decisions.

Professional Escalation Criteria

Escalate to a bioinformatics specialist or platform manufacturer when the tiered framework produces unexpected results. Specific escalation triggers include the following situations. More than 40 percent of cells fall into Tier 2, indicating that the primary thresholds are poorly calibrated or the experiment has systemic quality issues. The joint quality score shows no correlation with any individual modality metric, suggesting that the scoring approach is not capturing the relevant quality information. Tier assignment differs dramatically between batches despite similar experimental conditions, indicating a batch-specific technical problem. Manual review of Tier 2 cells reveals patterns that cannot be explained by known technical artifacts.

Consider repeating the experiment when the fraction of cells in Tier 1 is below 30 percent, when the relationship between modalities is systematically disrupted across most cells, or when the escalation review identifies a protocol problem that cannot be corrected by filtering. The lineage tracing study using combined snRNA and ATAC sequencing in kidney injury models retained 83,315 high-quality nuclei after quality control and doublet removal, demonstrating that optimized protocols can produce large Tier 1 populations. If your experiment produces substantially lower retention, the protocol instead of the filtering strategy may require attention.

Frequently Asked Questions

What is the difference between quality control for single-cell RNA sequencing and quality control for single-cell multi-omics?

Single-cell RNA sequencing quality control evaluates RNA-specific metrics such as UMI counts, gene counts, and mitochondrial read fraction. Single-cell multi-omics quality control must additionally evaluate modality-specific metrics such as TSS enrichment and fraction of reads in peaks for ATAC data, or bisulfite conversion rate and CpG coverage for methylation data. The filtering decision must consider joint information across modalities instead of applying independent thresholds.

How do I choose between intersection and union filtering for joint ATAC + RNA data?

Intersection filtering requires that a cell pass quality thresholds for both modalities, producing a higher-confidence cell population but potentially removing cells with valuable information in one modality. Union filtering retains cells that pass thresholds for at least one modality, maximizing cell retention but potentially introducing noise from cells with poor quality in one modality. The choice depends on whether the downstream analysis requires reliable data from both modalities.

What is TSS enrichment and why is it important for ATAC + RNA quality control?

TSS enrichment measures the accumulation of sequencing reads at transcription start sites relative to flanking regions. A high TSS enrichment score indicates that the transposition reaction successfully targeted open chromatin at promoters, which is the expected pattern for high-quality ATAC-seq data. Low TSS enrichment can indicate degraded chromatin, failed transposition, or excessive background signal.

How do I assess bisulfite conversion efficiency for Methylation + RNA data?

Bisulfite conversion efficiency is typically assessed using spike-in controls with known methylation status or by measuring methylation at known unmethylated regions. Incomplete conversion produces false methylation calls that can distort downstream analysis. The scNMT protocol update describes the experimental workflow for joint methylation and expression profiling, including quality control considerations.

What should I do if the RNA and ATAC modalities show inconsistent quality for the same cells?

Inconsistent quality between modalities can indicate technical problems such as degraded chromatin in cells with intact RNA, or failed reverse transcription in cells with intact chromatin. Examine the relationship between modality-specific metrics to identify patterns, and consider whether the inconsistency reflects a technical artifact or a genuine biological state. If the inconsistency is widespread, consider repeating the experiment.

How many cells should I expect to retain after joint quality control?

The number of cells retained depends on the quality of the experimental preparation, the stringency of the filtering thresholds, and the biological complexity of the sample. The kidney injury lineage tracing study retained 83,315 high-quality nuclei after quality control and doublet removal, but this was a large experiment with optimized protocols. There is no universal expected retention rate, and the appropriate rate depends on the specific experiment.

Can I use the same quality control thresholds for different batches or conditions?

Quality control thresholds should be evaluated separately for each batch or condition because technical quality can vary between samples. Applying the same thresholds across batches can introduce bias if one batch has systematically lower quality. Compare metric distributions across batches before applying thresholds, and consider whether batch-specific thresholds are appropriate.

What documentation should I maintain for quality control decisions?

Maintain records of sample metadata, sequencing statistics, per-cell quality metrics before and after filtering, the specific thresholds applied, the rationale for each threshold, the number of cells removed at each step, and the final cell counts. This documentation should be included in the methods section of any manuscript and should be available as a computational notebook or script.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.