A Practical Guide to 10x Multiome ATAC + Gene Expression: From Library Preparation to Integrated Analysis
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- The 10x Multiome ATAC + Gene Expression platform enables simultaneous profiling of chromatin accessibility and transcriptome from the same single nucleus, allowing direct linkage of regulatory elements to gene expression. Key quality metrics at each stage include nuclei yield and integrity, fragment distribution, library complexity, fraction of reads in peaks, and fraction of valid barcodes.
- Successful nuclei isolation is paramount, with fresh tissue generally yielding higher quality nuclei than frozen tissue, though optimized protocols can mitigate freeze-thaw damage; critical quality assessment involves nuclei concentration, integrity, and minimal debris/clumping.
- Library preparation requires careful optimization of Tn5 transposition conditions and PCR cycle number to avoid over- or under-tagmentation and ensure sufficient amplification without introducing bias, with final library assessment focusing on concentration and fragment size distribution.
- Computational processing with Cell Ranger ARC involves alignment, cell calling, and quality control metrics such as number of cells recovered, median genes per cell, and fraction of reads in peaks, followed by rigorous quality filtering for both RNA (e.g., minimum genes, mitochondrial read percentage) and ATAC (e.g., unique fragments, fraction of reads in peaks) components.
- Data integration, often via Seurat's weighted-nearest neighbor (WNN) analysis, combines RNA and ATAC modalities by learning cell-specific modality weights, with evaluation focusing on cluster separation and modality agreement, while downstream analyses include differential gene expression/accessibility, transcription factor activity estimation, and peak-to-gene linking.
- Troubleshooting common failure patterns like low cell recovery, high background, poor integration, data sparsity, or batch effects requires systematic evaluation of experimental and computational steps, with validation of key findings using orthogonal methods such as ChIP-seq or immunofluorescence being essential for robust biological interpretation.
The 10x Multiome ATAC + Gene Expression platform generates paired chromatin accessibility and transcriptome measurements from the same single nucleus, enabling researchers to link regulatory elements to gene expression within individual cells. This workflow covers experimental design, nuclei isolation, library preparation, sequencing considerations, and computational integration using tools such as Seurat's weighted-nearest neighbor analysis. The intended reader is a biology student, researcher, laboratory professional, or life-science practitioner who needs concrete decisions at each step, from nuclei isolation through integrated data interpretation.
At a Glance
The table below summarizes the major workflow stages, key decisions, and quality considerations for a 10x Multiome experiment.
| Workflow Stage | Primary Decision | Key Quality Metric | Common Failure Mode |
|---|---|---|---|
| Nuclei isolation | Fresh versus frozen tissue, mechanical versus enzymatic dissociation | Nuclei yield and integrity, minimal debris and clumping | Excessive debris or clumped nuclei reducing capture efficiency |
| Transposition | Tn5 tagmentation conditions and incubation time | Fragment distribution and accessible chromatin signal | Over-tagmentation producing short fragments with low signal |
| Library preparation | PCR cycle number and indexing strategy | Final library concentration and fragment size distribution | Insufficient or excessive amplification causing low complexity |
| Sequencing | Read depth and sequencing configuration per sample | Fraction of reads in peaks, median genes per nucleus | Low sequencing depth yielding sparse ATAC and RNA data |
| Computational processing | Cell Ranger ARC parameters and reference genome | Number of cells recovered, fraction of valid barcodes | Incorrect cell calling producing empty or doublet droplets |
| Quality filtering | Thresholds for RNA and ATAC metrics | Number of nuclei passing filters, doublet rate | Overly stringent or lenient filtering removing real cells or retaining debris |
| Integration | WNN analysis parameters and dimensionality | Cluster separation and modality agreement | Dominance of one modality in clustering decisions |
| Interpretation | Motif analysis and regulatory network inference | Transcription factor activity scores and peak-to-gene links | Overinterpretation of sparse data without validation |
Platform Capabilities and Biological Applications
The 10x Multiome platform simultaneously profiles gene expression and chromatin accessibility from the same nucleus by combining single-nucleus RNA sequencing with single-nucleus ATAC sequencing. This paired measurement allows researchers to ask questions that neither modality alone can answer, such as which transcription factors regulate cell-type-specific gene expression programs and how chromatin state changes relate to transcriptional output during development or disease.
The platform has been applied across diverse biological systems. A study of the developing human cerebral cortex generated a single-cell atlas of gene expression and chromatin accessibility both independently and jointly, revealing waves of gene regulation by key transcription factors across a nearly continuous differentiation trajectory and distinguishing the expression programs of glial lineages [<a href="#ref-1">1</a>]. In the context of diabetic cardiomyopathy, researchers processed single-cell RNA and ATAC data from 10x Multiome libraries using Cell Ranger ARC v2.0.1, then applied Seurat and Signac filtration, estimated transcription factor activity with chromVAR, and calculated cis-coaccessibility networks using Cicero [<a href="#ref-2">2</a>]. A bovine placenta study used the 10X Genomics multiome platform to determine chromatin accessibility landscapes across diverse cell populations, identifying candidate gene regulatory networks involved in trophoblast differentiation [<a href="#ref-3">3</a>].
These examples illustrate the range of experimental systems amenable to Multiome analysis, from human tissues to agricultural species. The key advantage is the ability to link open chromatin regions to the expression of nearby genes within the same cell, providing direct evidence for regulatory relationships that would require computational inference when using separate single-cell RNA and ATAC datasets.
Experimental Design Considerations
Sample Selection and Biological Replicates
The choice of sample type and number of biological replicates determines the statistical power of downstream analyses. For Multiome experiments, nuclei are the input material, so fresh or frozen tissue must be processed to yield intact nuclei suitable for transposition and reverse transcription. A protocol for isolating nuclei from murine cardiac tissue describes mechanical homogenization, sequential filtration, sucrose cushion purification, and fluorescence-activated nuclei sorting to enable multiomic analysis across various cardiac cell types [<a href="#ref-4">4</a>]. This protocol highlights the importance of removing debris and obtaining a clean nuclear suspension before proceeding to library preparation.
For studies comparing conditions such as disease versus control or treated versus untreated, biological replicates are essential to distinguish genuine biological variation from technical noise. The number of replicates depends on the expected effect size and the heterogeneity of the tissue. A study of gastric epithelial homeostasis collected ten unique gastric samples from wildtype mice and analyzed 31,598 cells after quality control [<a href="#ref-5">5</a>]. A study of peripheral blood mononuclear cells from two pig breeds generated single-cell transcriptomes and chromatin maps to compare immune cell heterogeneity across breeds [<a href="#ref-6">6</a>]. These examples demonstrate that replicate numbers vary by study design, but the principle remains that biological replication supports robust conclusions.
Fresh Versus Frozen Tissue
Fresh tissue generally yields higher quality nuclei with less debris and better preservation of accessible chromatin. However, frozen tissue is often more practical for clinical samples or field collections. The cardiac nuclei isolation protocol specifically describes steps for fresh-frozen murine cardiac ventricular tissue, indicating that frozen tissue can produce acceptable results when processed with appropriate methods [<a href="#ref-4">4</a>]. The key is to minimize freeze-thaw cycles and to optimize the isolation protocol for the specific tissue type.
Cell Number and Expected Recovery
The number of nuclei loaded onto the 10x Multiome platform determines the number of cells recovered after sequencing. Loading too few nuclei reduces data yield, while loading too many increases the risk of doublets, where two nuclei are captured in the same droplet and appear as a single cell with mixed profiles. A prostate cancer study applied the 10x Multiome platform to multiple cell lines and obtained 65,501 high quality single cells across eight cell lines [<a href="#ref-7">7</a>]. This number reflects the combined output from multiple samples, and individual samples typically yield between 5,000 and 15,000 cells depending on loading density and tissue type.
Targeted Enrichment Strategies
Standard Multiome profiling captures genome-wide chromatin accessibility and transcriptome information, but specific genomic regions of interest may require deeper coverage than the default approach provides. The prostate cancer study performed targeted sequencing to enrich sequencing data at prostate cancer risk loci involving 2,730 candidate germline variants and 273 associated genes [<a href="#ref-7">7</a>]. This targeted approach did not increase the number of captured cells but improved eQTL gene expression abundance by about 20% and chromatin accessibility abundance by about 5% [<a href="#ref-7">7</a>]. When the biological question centers on specific risk loci or candidate regulatory regions, a targeted enrichment strategy may be appropriate. When the goal is unbiased discovery of regulatory elements, standard genome-wide profiling remains the appropriate choice.
Nuclei Isolation and Quality Assessment
Isolation Methods
Nuclei isolation methods vary by tissue type and downstream requirements. Mechanical homogenization is common for solid tissues, while enzymatic dissociation may be needed for tissues with dense extracellular matrices. The cardiac protocol describes mechanical homogenization followed by sequential filtration and sucrose cushion purification [<a href="#ref-4">4</a>]. Fluorescence-activated nuclei sorting can further purify nuclei from debris and remove clumps, but it adds time and requires access to a sorter.
For cultured cells, a simpler lysis protocol is often sufficient. The prostate cancer study used multiple cell lines including RWPE1, RWPE2, PrEC, BPH1, DU145, PC3, 22Rv1, and LNCaP, which would require a gentler lysis than solid tissue [<a href="#ref-7">7</a>]. The choice of lysis buffer and detergent concentration affects nuclear integrity and the accessibility of chromatin to Tn5 transposase.
Quality Metrics for Nuclei
Before proceeding to library preparation, assess the nuclear suspension for yield, purity, and integrity. Key metrics include:
- Nuclei concentration, measured by hemocytometer or automated cell counter
- Percentage of intact nuclei, assessed by staining with a nuclear dye such as DAPI or propidium iodide
- Presence of debris or clumps, evaluated by microscopy
- RNA integrity, if assessing transcriptome quality separately
Debris and clumps reduce capture efficiency and increase the fraction of reads lost to ambient contamination. A clean nuclear suspension with minimal debris is essential for high-quality Multiome data.
Tissue-Specific Optimization
Different tissues present distinct challenges for nuclei isolation. The cardiac protocol was developed specifically for fresh-frozen murine cardiac ventricular tissue and includes steps for mechanical homogenization, sequential filtration, sucrose cushion purification, and fluorescence-activated nuclei sorting [<a href="#ref-4">4</a>]. The bovine placenta study used the 10X Genomics multiome platform to determine chromatin accessibility landscapes across diverse cell populations in developing and mature placenta, identifying distinct trophoblast, mesenchyme, endothelial, immune, and epithelial cell populations [<a href="#ref-3">3</a>]. Each tissue type may require adjustments to homogenization intensity, filtration pore sizes, and centrifugation conditions to achieve a clean nuclear preparation.
Transposition and Library Preparation
Tn5 Transposition
The Multiome platform uses Tn5 transposase to tag accessible chromatin regions. The transposition reaction occurs in the nucleus before droplet encapsulation, and the conditions must be optimized for each tissue type. Over-tagmentation produces short fragments with reduced signal, while under-tagmentation yields large fragments that may not be efficiently sequenced.
The prostate cancer study used Tn5 transposase-tagged nuclei from multiple cell lines, and the researchers performed targeted sequencing to enrich sequencing data at prostate cancer risk loci [<a href="#ref-7">7</a>]. This example illustrates that transposition conditions can be adapted to enrich for regions of interest.
Reverse Transcription and cDNA Amplification
After transposition, nuclei are encapsulated in droplets where reverse transcription converts mRNA to cDNA and the transposed DNA fragments are amplified. The PCR cycle number must be optimized to produce sufficient material without over-amplification, which can introduce bias and reduce library complexity.
The final library contains both the ATAC fragments and the cDNA from gene expression, which are separated during sequencing by their distinct read structures. The library preparation protocol includes a cleanup step to remove primers and unwanted fragments before sequencing.
Library Quality Assessment
Before sequencing, assess the final library for concentration and fragment size distribution. The library should show a nucleosomal pattern in the ATAC portion, with a prominent mononucleosomal peak and smaller peaks corresponding to di- and tri-nucleosomes. The cDNA portion should show a broad size distribution corresponding to the transcript lengths.
A Bioanalyzer or TapeStation trace provides the fragment size distribution, and quantitative PCR or fluorometric quantification provides the concentration. Libraries with low concentration or abnormal fragment distributions should be re-amplified or re-prepared before sequencing.
Sequencing Considerations
Read Depth and Configuration
The recommended sequencing depth for Multiome libraries depends on the number of nuclei and the biological questions. Gene expression libraries typically require 20,000 to 50,000 read pairs per nucleus, while ATAC libraries require 25,000 to 50,000 read pairs per nucleus. The prostate cancer study used targeted sequencing to enrich data at risk loci, which improved abundance without increasing the number of captured cells [<a href="#ref-7">7</a>]. This approach may be useful when specific genomic regions are of primary interest.
The sequencing configuration follows the 10x Genomics recommended settings, with paired-end reads covering the transposed fragments and the cDNA. The read lengths are determined by the library structure and the need to identify the cell barcode and unique molecular identifier.
Sequencing Quality Control
After sequencing, assess the quality of the raw data before proceeding to alignment. Key metrics include:
- Percentage of bases with quality score above Q30
- Percentage of reads mapping to the reference genome
- Fraction of reads in peaks for the ATAC portion
- Median genes per nucleus for the RNA portion
Low mapping rates may indicate contamination or poor library quality, while low fractions of reads in peaks suggest excessive background or poor transposition efficiency.
Computational Processing with Cell Ranger ARC
Alignment and Cell Calling
Cell Ranger ARC is the standard pipeline for processing 10x Multiome data. It aligns the ATAC and RNA reads to the reference genome, assigns reads to cell barcodes, and calls cells based on the combined signal from both modalities. The diabetic cardiomyopathy study processed single-cell RNA and ATAC data using Cell Ranger ARC v2.0.1 [<a href="#ref-2">2</a>], and the gastric homeostasis study used Seurat and Signac with weighted-nearest neighbors analysis for joint RNA and ATAC clustering [<a href="#ref-5">5</a>].
Cell calling is a critical step that determines which barcodes are considered cells and which are considered background. The default parameters in Cell Ranger ARC work well for most samples, but manual adjustment may be needed for samples with unusual background levels or cell types with low RNA content.
Output Files and Quality Metrics
Cell Ranger ARC produces several output files, including filtered feature-barcode matrices for RNA and ATAC, peak calls, and summary metrics. Key quality metrics to review include:
- Number of cells recovered
- Median genes per cell
- Median unique fragments per cell
- Fraction of reads in peaks
- Fraction of reads mapped to the genome
- Percentage of reads that are mitochondrial
These metrics provide a first indication of data quality and guide downstream filtering decisions.
Quality Filtering and Data Preprocessing
RNA Quality Filters
The RNA component of Multiome data requires filtering to remove low-quality cells and ambient contamination. Common filters include:
- Minimum number of genes detected per cell, typically 500 to 1,000
- Maximum number of genes detected, to remove potential doublets
- Percentage of mitochondrial reads, with high percentages indicating stressed or dying cells
The gastric homeostasis study analyzed 31,598 cells after quality control with Seurat and Signac [<a href="#ref-5">5</a>], indicating that a substantial fraction of captured cells may be removed during filtering. The thresholds should be adjusted based on the tissue type and the distribution of metrics in the data.
ATAC Quality Filters
The ATAC component requires separate quality filters, including:
- Minimum number of unique fragments per cell
- Fraction of fragments in peaks
- Nucleosome signal, with low signal indicating good accessibility
- Transcription start site enrichment, with high enrichment indicating good signal
The diabetic cardiomyopathy study applied Seurat and Signac filtration to the gene expression and ATAC data [<a href="#ref-2">2</a>], demonstrating that both modalities require independent quality assessment before integration.
Doublet Detection and Removal
Doublets are a major source of noise in single-cell data, and Multiome data are particularly susceptible because two nuclei in the same droplet produce both RNA and ATAC signals from different cells. Doublet detection methods use the combined signal to identify cells with unusually high complexity or mixed profiles. Removing doublets before downstream analysis improves cluster resolution and reduces false findings.
Data Integration with Seurat and Signac
Weighted-Nearest Neighbor Analysis
Seurat's weighted-nearest neighbor analysis integrates the RNA and ATAC modalities by computing a weighted combination of the two data types for each cell. The weights are learned from the data and reflect the informativeness of each modality for each cell. The gastric homeostasis study used weighted-nearest neighbors analysis for joint RNA and ATAC clustering [<a href="#ref-5">5</a>], and the approach is standard for Multiome data analysis.
The WNN analysis produces a joint embedding that can be used for dimensionality reduction, clustering, and visualization. The relative contribution of each modality can be examined to understand which cell types are better defined by RNA or ATAC signal.
Alternative Integration Methods
Several alternative methods exist for integrating single-cell chromatin accessibility and gene expression data. The sciCAN method uses a cycle-consistent adversarial network to integrate scATAC-seq and scRNA-seq data in an unsupervised manner, and it was benchmarked against five existing methods across five datasets [<a href="#ref-8">8</a>]. The single-cell Multi-View Profiler is a deep generative model designed for data that simultaneously measure gene expression and chromatin accessibility, including Multiome from 10x Genomics [<a href="#ref-9">9</a>]. These methods may be useful when the standard Seurat workflow does not adequately integrate the modalities or when the data have specific characteristics such as high sparsity.
Evaluating Integration Quality
After integration, assess whether the joint representation preserves biological relationships and whether the modalities agree on cell identity. The sciCAN study confirmed that the integrated representation preserves biological relationships within the hematopoietic hierarchy when applied to 10x Multiome data [<a href="#ref-8">8</a>]. Visual inspection of the joint embedding, examination of modality weights, and comparison of cluster markers across modalities provide evidence for integration quality.
Downstream Analyses
Differential Gene Expression and Accessibility
Differential gene expression analysis identifies genes whose expression differs between cell types or conditions. The diabetic cardiomyopathy study performed differential gene expression and identified accessible chromatin regions, then estimated transcription factor activity with chromVAR and calculated cis-coaccessibility networks using Cicero [<a href="#ref-2">2</a>]. These analyses revealed altered cell proportions in the disease group, including decreased endothelial cells and macrophages and increased fibroblasts and myocardial cells [<a href="#ref-2">2</a>].
For Multiome data, differential accessibility analysis identifies peaks that differ between conditions. The bovine placenta study identified ATAC-seq peaks that defined open chromatin regions, facilitating the identification of transcription factor binding sites and candidate gene regulatory networks involved with trophoblast differentiation [<a href="#ref-3">3</a>].
Transcription Factor Activity and Regulatory Networks
Transcription factor activity can be estimated from chromatin accessibility data using motif analysis. The gastric homeostasis study used SCENIC+, MultiVelo, EpiCHAOS, and cell plasticity scores to uncover gene regulatory networks, cell state dynamics, and lineage trajectories [<a href="#ref-5">5</a>]. This analysis revealed previously uncharacterized regulatory networks comprising novel transcription factor combinations that define cell identities, including Ppara, Pparg, Arid5b, and Sox5 as candidate regulators of parietal, foveolar, chief, and neck cells [<a href="#ref-5">5</a>].
The memory T cell study used single-nuclei multiome sequencing to characterize dynamic activation responses and reconstruct gene regulatory networks, identifying memory-associated transcription factors including MAF, PRDM1, RUNX2, SMAD3, and KLF6 [<a href="#ref-10">10</a>]. These factors were predicted to orchestrate rapid recall, and KLF6 binding to its predicted target genes was confirmed by chromatin immunoprecipitation sequencing [<a href="#ref-10">10</a>].
Peak-to-Gene Links and Cis-Regulatory Elements
A key advantage of Multiome data is the ability to link peaks to genes within the same cell. Co-accessibility analysis identifies peaks that are accessible in the same cells as their putative target genes, providing evidence for cis-regulatory relationships. The aging brain study correlated open chromatin regions with transcription factor expression to identify age-associated regulatory networks and used co-accessibility to identify linked peaks and genes, revealing a catalog of putative cis-regulatory elements by cell type [<a href="#ref-11">11</a>].
The prostate cancer study used multiomic profiling to associate RNA expression alterations with chromatin accessibility of germline variants at single-cell levels, and cross-validation analysis showed high overlaps between the multiome associations and bulk eQTL findings from the GTEx prostate cohort [<a href="#ref-7">7</a>]. This approach demonstrates how Multiome data can link genetic variants to their regulatory effects.
Cell-Cell Communication and Pathway Analysis
The diabetic cardiomyopathy study performed cell-cell communication analysis and gene-motif correlation to reveal intricate molecular changes during disease progression [<a href="#ref-2">2</a>]. These analyses extend the utility of Multiome data beyond regulatory relationships to include intercellular signaling networks. Cell-cell communication analysis uses ligand-receptor pairs expressed in different cell types to infer signaling interactions, while gene-motif correlation links transcription factor motif accessibility to target gene expression.
Common Failure Patterns and Troubleshooting
Low Cell Recovery
Low cell recovery can result from poor nuclei isolation, low loading density, or inefficient capture. If the number of cells recovered is substantially lower than expected, check the nuclei concentration and viability before loading, and verify that the transposition reaction was efficient. The cardiac protocol emphasizes the importance of a clean nuclear suspension, and fluorescence-activated nuclei sorting can improve purity [<a href="#ref-4">4</a>].
High Background or Ambient Contamination
Ambient RNA and ATAC fragments from lysed cells contribute to background signal and can obscure true biological differences. High background is often caused by excessive debris in the nuclear suspension or by cells lysing during the experiment. Reducing the time between nuclei isolation and encapsulation, and using more stringent washing steps, can reduce ambient contamination.
Poor Integration Between Modalities
When the RNA and ATAC modalities do not agree on cell identity, the integration may be dominated by one modality or may fail to capture the true biological structure. The sciCAN method was developed to address integration challenges and demonstrated consistent performance across datasets [<a href="#ref-8">8</a>]. The scMVP method provides separate imputations for differential analysis and cis-regulatory element identification, which can help mitigate data sparsity issues [<a href="#ref-9">9</a>].
Sparse Data and Missing Values
Single-cell data are inherently sparse, with many genes and peaks not detected in every cell. The prostate cancer study addressed this by performing targeted sequencing to enrich data at risk loci, which improved gene expression abundance by about 20% and chromatin accessibility abundance by about 5% [<a href="#ref-7">7</a>]. Computational imputation methods can also help, but they should be used with caution because they can introduce false signals.
Batch Effects Across Samples
When processing multiple samples or batches, technical variation can confound biological differences. The aging brain study generated single-nucleus multiome profiles from 357 human brain samples and classified cells into seven major cell types using canonical marker genes [<a href="#ref-11">11</a>]. Such large-scale studies require careful batch correction to ensure that clustering reflects biological variation instead of technical artifacts.
Records and Documentation
Laboratory Records
Maintain detailed records of the experimental protocol, including:
- Tissue source, collection date, and storage conditions
- Nuclei isolation method and buffer composition
- Transposition conditions, including enzyme lot and incubation time
- PCR cycle number and amplification conditions
- Library quantification results and fragment size distributions
- Sequencing run parameters and quality metrics
These records support troubleshooting and reproducibility. The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility [<a href="#ref-12">12</a>], and the nf-core documentation describes community pipeline standards for reproducible workflow configuration [<a href="#ref-13">13</a>].
Computational Records
Document the computational environment, including software versions, parameters, and reference genome versions. The Bioconductor project provides official package, workflow, installation, and reproducible genomic-analysis documentation [<a href="#ref-14">14</a>], and the EMBL-EBI Training program offers bioinformatics learning pathways and practical analysis education [<a href="#ref-15">15</a>]. Version control using Git, as taught by The Carpentries lessons [<a href="#ref-16">16</a>], supports reproducible analysis by tracking changes to scripts and parameters.
Data Management and Repositories
The NCBI provides official descriptions of databases, search systems, sequence resources, and analysis services [<a href="#ref-17">17</a>]. Depositing raw and processed Multiome data in appropriate repositories supports transparency and enables reanalysis by other researchers. Data management plans should address file formats, metadata standards, and access restrictions for sensitive samples.
Limitations and Interpretation Caveats
Technical Limitations
Multiome data have several technical limitations that affect interpretation. The data are sparse, with many genes and peaks not detected in every cell. The prostate cancer study noted data sparsity as a common challenge and used targeted sequencing to address it [<a href="#ref-7">7</a>]. The scMVP method was designed to handle data sparsity with imputation [<a href="#ref-9">9</a>], but imputed values should be interpreted as estimates instead of direct measurements.
The capture efficiency for RNA and ATAC differs, and some cell types may be better represented in one modality than the other. The WNN analysis in Seurat accounts for this by weighting each modality based on its informativeness for each cell [<a href="#ref-5">5</a>], but the weights should be examined to understand which cell types are driven by which modality.
Biological Interpretation
The identification of transcription factor motifs in accessible chromatin does not prove that those factors are active in the cells. The memory T cell study confirmed KLF6 binding by chromatin immunoprecipitation sequencing [<a href="#ref-10">10</a>], but such validation is not always feasible. Similarly, peak-to-gene links identified by co-accessibility are correlative and require experimental validation.
The gastric homeostasis study noted that stem cells can be defined not by a separate epigenetic state but by epigenetic plasticity, with open chromatin states of all differentiated cell types in the absence of transcriptional reprogramming [<a href="#ref-5">5</a>]. This finding illustrates that the relationship between chromatin accessibility and gene expression is complex and that accessibility does not always predict expression.
Model Assumptions and Computational Limitations
Integration methods make assumptions about the relationship between modalities. The sciCAN method uses a cycle-consistent adversarial network to integrate data in an unsupervised manner [<a href="#ref-8">8</a>], while scMVP generates common latent representations for dimensionality reduction, cell clustering, and developmental trajectory inference [<a href="#ref-9">9</a>]. These models may not capture all biological complexity, and results should be interpreted within the framework of the chosen method.
Professional Escalation Criteria
When to Seek Expert Assistance
Several situations warrant consultation with a bioinformatics specialist or the platform vendor:
- Consistently low cell recovery or poor data quality across multiple samples
- Unexpected patterns in quality metrics that cannot be explained by the experimental design
- Integration failures where the RNA and ATAC modalities cannot be reconciled
- Need for advanced analyses such as trajectory inference, RNA velocity, or regulatory network reconstruction
The NCBI provides official descriptions of databases, search systems, sequence resources, and analysis services [<a href="#ref-17">17</a>], and the EMBL-EBI Training program offers practical analysis education [<a href="#ref-15">15</a>]. These resources can help researchers identify appropriate tools and methods.
Validation of Findings
Before publishing or acting on Multiome findings, validate key results with orthogonal methods. The memory T cell study used chromatin immunoprecipitation sequencing to confirm transcription factor binding [<a href="#ref-10">10</a>], and the diabetic cardiomyopathy study used immunofluorescent staining on paraffin-embedded tissues to verify findings [<a href="#ref-2">2</a>]. The EMT study used CRISPR/Cas9-mediated loss-of-function studies coupled with in vitro and in vivo functional assays to identify transcription factors regulating specific EMT states [<a href="#ref-18">18</a>]. These validation approaches provide confidence in the computational findings.
A Practical Decision Framework for Multiome Data Analysis
When to Use Which Integration Method
The choice of integration method for paired ATAC and gene expression data depends on the biological question, data characteristics, and available computational resources. Seurat's weighted-nearest neighbor analysis is the standard approach for 10x Multiome data and works well for most datasets where the goal is joint clustering and cell-type identification [<a href="#ref-5">5</a>]. The WNN approach learns modality weights for each cell, allowing the analysis to favor RNA signal in cell types where gene expression is informative and ATAC signal in cell types where chromatin accessibility provides better discrimination.
Alternative methods become appropriate under specific conditions. The sciCAN method uses a cycle-consistent adversarial network to integrate single-cell chromatin accessibility and gene expression data in an unsupervised manner, and it demonstrated consistent performance across datasets with better balance of mutual transferring between modalities than five existing methods [<a href="#ref-8">8</a>]. This method may be preferable when the two modalities show substantial disagreement or when the dataset contains cells that are poorly represented in one modality. The single-cell Multi-View Profiler generates common latent representations for dimensionality reduction, cell clustering, and developmental trajectory inference while providing separate imputations for differential analysis and cis-regulatory element identification [<a href="#ref-9">9</a>]. This method is particularly useful when data sparsity limits the reliability of direct integration.
A practical decision framework for method selection considers three factors. First, assess the degree of modality agreement by comparing cluster assignments from independent RNA and ATAC analyses before integration. If the adjusted Rand index between modality-specific clusterings exceeds 0.7, WNN analysis will likely perform well. If the index falls below 0.5, consider sciCAN or scMVP. Second, evaluate data sparsity by examining the fraction of cells with detectable expression for housekeeping genes and the fraction of peaks with reads in a substantial proportion of cells. Highly sparse data may benefit from scMVP imputation. Third, consider the downstream analysis goals. If the primary objective is trajectory inference or regulatory network reconstruction, methods that preserve continuous structure such as scMVP may be preferable to methods optimized for discrete clustering.
A Structured Troubleshooting Protocol for Integration Failures
Integration failures manifest as poor cluster separation, modality dominance, or biologically implausible groupings. A structured troubleshooting protocol addresses these failures systematically instead of through trial and error.
Begin by examining the modality weights from the WNN analysis. If one modality receives near-zero weights across most cells, the other modality is dominating the integration. This pattern often indicates that the underrepresented modality has poor signal quality. Check the quality metrics for that modality independently. For ATAC data, verify the fraction of reads in peaks and transcription start site enrichment. For RNA data, verify the median genes per cell and the fraction of mitochondrial reads. If one modality fails quality thresholds, reprocess the data with adjusted filtering parameters before attempting integration again.
If both modalities pass quality checks but integration still fails, examine the feature selection step. The integration relies on a shared set of features that capture biological variation in both modalities. If the selected features are dominated by genes or peaks with high technical noise, the integration will reflect technical artifacts instead of biological structure. Re-run feature selection with stricter criteria, such as requiring features to be variable in both modalities or using a larger set of highly variable features.
When integration failures persist, test whether the problem is specific to certain cell types. Subset the data by preliminary cluster assignments and examine modality agreement within each cluster. Cell types with low RNA content, such as quiescent cells or certain immune populations, may be better defined by ATAC signal, while cell types with similar chromatin landscapes may require RNA signal for discrimination. The WNN approach accounts for this by learning cell-specific modality weights [<a href="#ref-5">5</a>], but the weights may not fully compensate when one modality is uninformative for a large fraction of cells.
A Record System for Multiome Experiments
A structured record system supports troubleshooting, reproducibility, and comparison across experiments. The system should capture experimental parameters, quality metrics, and analysis decisions in a format that allows rapid review when problems arise.
For each experiment, record the tissue source, collection date, storage conditions, and any deviations from the standard protocol. Document the nuclei isolation method, including buffer composition, homogenization settings, filtration pore sizes, and centrifugation conditions. The cardiac nuclei isolation protocol describes mechanical homogenization, sequential filtration, sucrose cushion purification, and fluorescence-activated nuclei sorting as key steps [<a href="#ref-4">4</a>], and each of these steps should be documented with specific parameters.
Record transposition conditions, including the Tn5 enzyme lot, incubation time and temperature, and the number of nuclei used. Document PCR cycle numbers for both the cDNA amplification and the final library amplification. Record library quantification results, including concentration and fragment size distribution from the Bioanalyzer or TapeStation trace.
For sequencing, record the platform, flow cell type, read configuration, and the number of read pairs per nucleus. Document the Cell Ranger ARC version and parameters, including the reference genome version and any custom settings for cell calling. Record the quality metrics from the Cell Ranger ARC summary, including the number of cells recovered, median genes per cell, median unique fragments per cell, fraction of reads in peaks, and fraction of reads mapped to the genome.
For downstream analysis, record the software versions for Seurat, Signac, and any additional packages. Document the filtering thresholds applied to both RNA and ATAC data, the number of cells removed at each filtering step, and the rationale for threshold choices. Record the integration method used, the parameters for that method, and the quality metrics used to evaluate integration success.
The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility [<a href="#ref-12">12</a>], and the nf-core documentation describes community pipeline standards for reproducible workflow configuration [<a href="#ref-13">13</a>]. These resources offer templates for documenting computational workflows. The Carpentries lessons provide foundational training in version control using Git [<a href="#ref-16">16</a>], which supports tracking changes to analysis scripts and parameters over time.
Common Failure Patterns and Their Root Causes
Several failure patterns recur across Multiome experiments, and recognizing these patterns speeds up troubleshooting.
Low cell recovery with high background signal often traces to the nuclei isolation step. Excessive debris in the nuclear suspension reduces capture efficiency and increases ambient contamination. The cardiac protocol addresses this with sequential filtration and sucrose cushion purification [<a href="#ref-4">4</a>], and fluorescence-activated nuclei sorting can further improve purity. If low recovery persists after optimizing isolation, check the loading density and the transposition efficiency.
Poor ATAC signal with adequate RNA signal typically indicates suboptimal transposition conditions. Over-tagmentation produces short fragments with reduced signal, while under-tagmentation yields large fragments that may not be efficiently sequenced. The prostate cancer study used Tn5 transposase-tagged nuclei from multiple cell lines [<a href="#ref-7">7</a>], and the transposition conditions should be optimized for each tissue type.
Poor RNA signal with adequate ATAC signal may indicate RNA degradation during nuclei isolation or inefficient reverse transcription. Check the RNA integrity before proceeding, and minimize the time between nuclei isolation and encapsulation.
Integration failures where the modalities cannot be reconciled often trace to batch effects or technical variation between samples. The aging brain study generated single-nucleus multiome profiles from 357 human brain samples [<a href="#ref-11">11</a>], and such large-scale studies require careful batch correction. If integration fails across samples, examine whether the failure is driven by batch-specific technical variation instead of biological differences.
When to Escalate to Expert Assistance
Certain situations warrant consultation with a bioinformatics specialist or the platform vendor. Escalate when quality metrics remain poor after systematic troubleshooting, when integration failures persist across multiple parameter combinations, or when the data show unexpected patterns that cannot be explained by the experimental design.
The NCBI provides official descriptions of databases, search systems, sequence resources, and analysis services [<a href="#ref-17">17</a>], and the EMBL-EBI Training program offers bioinformatics learning pathways and practical analysis education [<a href="#ref-15">15</a>]. These resources can help researchers identify appropriate tools and methods when standard approaches fail.
Before escalating, document the troubleshooting steps already attempted, including the parameters tested and the resulting quality metrics. This documentation helps the specialist diagnose the problem more efficiently and avoids repeating failed approaches.
Validation of Integration Results
Integration results should be validated before proceeding to downstream analyses. The gastric homeostasis study used weighted-nearest neighbors analysis for joint RNA and ATAC clustering and validated the results by identifying known regulators of stem-cell differentiation into mature cell types [<a href="#ref-5">5</a>]. This validation approach confirms that the integration recovers expected biological relationships.
Additional validation strategies include comparing cluster markers across modalities, examining whether known cell-type-specific transcription factors show concordant accessibility and expression, and testing whether the integration preserves expected lineage relationships. The sciCAN study confirmed that the integrated representation preserves biological relationships within the hematopoietic hierarchy when applied to 10x Multiome data [<a href="#ref-8">8</a>], demonstrating that integration quality can be assessed by checking known biological structure.
The diabetic cardiomyopathy study used immunofluorescent staining on paraffin-embedded tissues to verify findings from the integrated analysis [<a href="#ref-2">2</a>], and the memory T cell study used chromatin immunoprecipitation sequencing to confirm transcription factor binding [<a href="#ref-10">10</a>]. These orthogonal validation approaches provide confidence in the computational findings and should be considered when the biological conclusions have important implications.
Frequently Asked Questions
What is the difference between 10x Multiome and separate single-cell RNA and ATAC experiments?
The 10x Multiome platform measures gene expression and chromatin accessibility from the same nucleus, providing paired data for each cell. Separate experiments measure each modality from different cells, requiring computational integration to link the data. The paired nature of Multiome data enables direct analysis of regulatory relationships, such as linking open chromatin regions to the expression of nearby genes within the same cell [<a href="#ref-1">1</a>][<a href="#ref-3">3</a>].
How many cells should I aim to recover for a Multiome experiment?
The number of cells depends on the biological question and the heterogeneity of the tissue. Studies have analyzed from about 31,000 cells in a gastric homeostasis study [<a href="#ref-5">5</a>] to over 65,000 cells across multiple prostate cancer cell lines [<a href="#ref-7">7</a>]. For rare cell types or complex tissues, more cells provide better statistical power. The loading density should be adjusted to balance cell recovery against doublet rate.
What are the most important quality metrics for Multiome data?
The most important metrics include the number of cells recovered, median genes per cell for RNA, median unique fragments per cell for ATAC, fraction of reads in peaks, and fraction of reads mapped to the genome. The fraction of mitochondrial reads indicates cell health, and the transcription start site enrichment indicates ATAC signal quality. These metrics should be reviewed after Cell Ranger ARC processing and before downstream analysis.
How do I choose between Seurat WNN analysis and alternative integration methods?
Seurat WNN analysis is the standard approach for 10x Multiome data and works well for most datasets [<a href="#ref-5">5</a>]. Alternative methods such as sciCAN [<a href="#ref-8">8</a>] and scMVP [<a href="#ref-9">9</a>] may be useful when the standard workflow does not adequately integrate the modalities or when the data have specific characteristics such as high sparsity. The choice depends on the data characteristics and the biological questions.
What is the role of transcription factor motif analysis in Multiome data interpretation?
Transcription factor motif analysis identifies transcription factor binding sites in accessible chromatin regions, providing hypotheses about which transcription factors regulate gene expression in specific cell types. The bovine placenta study used ATAC-seq peaks to identify transcription factor binding sites and candidate gene regulatory networks [<a href="#ref-3">3</a>], and the gastric homeostasis study used SCENIC+ to uncover gene regulatory networks [<a href="#ref-5">5</a>]. Motif analysis is correlative and requires experimental validation.
How do I validate findings from Multiome data?
Validation methods include chromatin immunoprecipitation sequencing to confirm transcription factor binding [<a href="#ref-10">10</a>], immunofluorescent staining to verify protein expression [<a href="#ref-2">2</a>], and CRISPR/Cas9-mediated loss-of-function studies to test the functional role of transcription factors [<a href="#ref-18">18</a>]. The choice of validation method depends on the specific findings and the experimental system.
What are the common causes of poor Multiome data quality?
Poor data quality can result from suboptimal nuclei isolation, excessive debris or clumping, over- or under-tagmentation, insufficient sequencing depth, or incorrect cell calling. The cardiac nuclei isolation protocol emphasizes the importance of a clean nuclear suspension [<a href="#ref-4">4</a>], and the prostate cancer study used targeted sequencing to address data sparsity [<a href="#ref-7">7</a>]. Troubleshooting should begin with the nuclei isolation and proceed through each step of the workflow.
How should I document my Multiome experiment for reproducibility?
Document the experimental protocol, including tissue source, isolation method, transposition conditions, PCR parameters, and sequencing configuration. Document the computational environment, including software versions and parameters. Use version control for analysis scripts, as taught by The Carpentries lessons [<a href="#ref-16">16</a>], and follow community standards for reproducible workflows as described in the nf-core documentation [<a href="#ref-13">13</a>]. The Galaxy Training Network provides accessible workflow training that emphasizes reproducibility [<a href="#ref-12">12</a>].
Related Bioinformatics Guides
- Single-Cell Sequencing Workflow: From Sample Preparation to Data Analysis
- Proteomics Data Analysis in R: A Practical Workflow for Differential Expression and Visualization
- Metabolomics Data Analysis in R: A Practical Workflow
- Spatial Transcriptomics Workflow: From Sample Preparation to Data Analysis
- RNA Sequencing Data Analysis: From Raw Reads to Differential Expression
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
[1] [Chromatin and gene-regulatory dynamics of the developing human cerebral cortex at single-cell resolution.](https://pubmed.ncbi.nlm.nih.gov/34390642). Cell, 2021. [2] [Single-cell insights: pioneering an integrated atlas of chromatin accessibility and transcriptomic landscapes in diabetic cardiomyopathy.](https://pubmed.ncbi.nlm.nih.gov/38664790). Cardiovascular diabetology, 2024. [3] [Single cell multiome analysis of the bovine placenta identifies gene regulatory networks in trophoblast differentiation†.](https://pubmed.ncbi.nlm.nih.gov/39987557). Biology of reproduction, 2025. [4] [Protocol for isolation of nuclei from murine cardiac tissue for single-nucleus multiomic sequencing.](https://doi.org/10.1016/j.xpro.2026.104615). 2026. [5] [Single-nucleus multiome sequencing identifies candidate regulators of mouse gastric epithelial homeostasis.](https://pubmed.ncbi.nlm.nih.gov/42094467). bioRxiv : the preprint server for biology, 2026. [6] [Single-cell transcriptomic and chromatin accessibility atlas of peripheral blood mononuclear cells reveals immune cell heterogeneity and breed-specific characteristics in Duroc and Meishan pigs.](https://pubmed.ncbi.nlm.nih.gov/42032457). BMC genomics, 2026. [7] [Identify Regulatory eQTLs by Multiome Sequencing in Prostate Single Cells.](https://pubmed.ncbi.nlm.nih.gov/38948854). bioRxiv : the preprint server for biology, 2024. [8] [sciCAN: single-cell chromatin accessibility and gene expression data integration via cycle-consistent adversarial network.](https://pubmed.ncbi.nlm.nih.gov/36089620). NPJ systems biology and applications, 2022. [9] [A deep generative model for multi-view profiling of single-cell RNA-seq and ATAC-seq data.](https://pubmed.ncbi.nlm.nih.gov/35022082). Genome biology, 2022. [10] [Gene regulatory network determinants of rapid recall in human memory CD4<,sup>,+<,/sup>, T cells.](https://doi.org/10.1016/j.celrep.2026.117103). 2026. [11] [Single-nucleus multiome analysis in the human prefrontal cortex identifies gene expression and cis-regulatory elements associated with aging.](https://doi.org/10.1016/j.celrep.2026.117110). 2026. [12] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [13] [nf-core Documentation](https://nf-co.re/docs). nf-core. [14] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [15] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [16] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [17] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [18] [Single cell multiomics unravel the transcription networks controlling the different EMT tumor states](https://doi.org/10.21203/rs.3.rs-9426544/v1). 2026.This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.