Single-Cell TCR and BCR Sequencing: Library Preparation Strategies for Immune Repertoire Analysis

By Dr. Zubair Khalid, DVM, MS, PhD ·

Single-Cell TCR and BCR Sequencing: Library Preparation Strategies for Immune Repertoire Analysis

Key Takeaways

  • Paired TCR/BCR sequencing is critical for linking antigen specificity to cellular function and state. Preserving the natural pairing of alpha/beta T cell receptor (TCR) or heavy/light chain B cell receptor (BCR) chains within individual cells is essential for understanding clonal identity and antigen recognition, as demonstrated in studies of tumor-infiltrating lymphocytes and B cell receptor diversity correlating with immune checkpoint inhibitor sensitivity.
  • Library preparation strategies diverge between targeted amplicon and full-length approaches, each with distinct trade-offs. Targeted amplicon methods offer high depth on receptor loci for detailed repertoire composition analysis but sacrifice transcriptome data, while full-length single-cell RNA sequencing captures both receptor sequences and the broader transcriptome, enabling integration of receptor identity with cell state and functional programs.
  • Full-length single-cell RNA sequencing (scRNA-seq) is paramount for integrating receptor identity with cellular transcriptional state. This approach allows researchers to correlate specific TCR or BCR clonotypes with distinct cell phenotypes, differentiation trajectories, and functional programs, as exemplified by studies linking neoantigen-specific T cells to cytolytic programs and memory cell hallmarks.
  • 3'-directed scRNA-seq workflows can be adapted for receptor profiling using specialized methods like circVDJ-seq. This is crucial for studies employing single-nucleus RNA-seq, ATAC+RNA multi-omics, or spatial transcriptomics, enabling VDJ information recovery where standard 5'-end capture is not feasible and expanding the scope of immune repertoire analysis.
  • Platform selection (e.g., 10X Genomics, Parse Biosciences, microwell systems) significantly impacts gene detection breadth, multimodal capabilities, and cost. Comparative analyses reveal differences in captured gene sets and T cell subset marker detection, necessitating alignment of platform choice with specific research objectives, such as prioritizing non-coding gene capture or detailed effector function characterization.
  • Robust bioinformatics quality control, including assessment of cell viability, multiplet rates, and receptor capture efficiency, is essential for reliable repertoire analysis. Failure patterns such as low viability or high multiplet rates can lead to inaccurate receptor assignment and chimeric sequences, necessitating careful documentation and troubleshooting to ensure data integrity and reproducibility.

Immunology researchers studying clonal diversity and antigen specificity must select library preparation methods that preserve paired TCR or BCR information while capturing sufficient V(D)J coverage. The central decision is between targeted amplicon approaches that enrich for receptor transcripts and full-length or whole-transcriptome methods that capture receptor sequences alongside broader gene expression data. This article compares these strategies, defines the practical workflow decisions, and outlines quality controls, records, and interpretation limits that apply across single-cell immune repertoire studies.

Scope and Reader Context

This article serves biology students, researchers, laboratory professionals, and life-science practitioners who need to design or evaluate single-cell TCR and BCR sequencing experiments. The focus is library preparation strategy: how to choose between targeted amplicon and full-length approaches, how to preserve V(D)J coverage, how to maintain paired receptor information, and how to integrate receptor sequences with transcriptomic data. The discussion covers data inputs, workflow choices, controls, quality checks, reproducibility, interpretation limits, reporting, and practical decision criteria. The evidence base includes peer-reviewed studies of T cell and B cell repertoires in cancer, autoimmunity, infectious disease, and development, plus official documentation from bioinformatics training and workflow resources.

The Biological Rationale for Paired Receptor Sequencing

T cells and B cells recognize antigens through clonally distributed receptors generated by V(D)J recombination. The TCR alpha and beta chains, or the BCR heavy and light chains, together determine antigen specificity. Single-cell sequencing that preserves the natural pairing of these chains provides information that bulk repertoire sequencing cannot: the identity of the complete receptor expressed by each cell. This pairing information matters because the functional specificity of a lymphocyte depends on both chains, and the combination cannot be reliably inferred from population-level frequencies.

Studies of tumor-infiltrating lymphocytes illustrate why paired information matters. In hepatocellular carcinoma, deep single-cell RNA sequencing of T cells from peripheral blood, tumor, and adjacent normal tissue, coupled with assembled TCR sequences, enabled identification of 11 T cell subsets based on molecular and functional properties. Specific subsets such as exhausted CD8+ T cells and regulatory T cells were preferentially enriched and potentially clonally expanded in the tumor, and signature genes were identified for each subset [<a href="#ref-1">1</a>]. This type of analysis depends on knowing which TCR sequence belongs to which transcriptional state, which requires single-cell pairing.

The same logic applies to B cells. B cell receptor sequences define the antibody specificities present in a sample, and paired heavy and light chain information is required to reconstruct functional antibodies. Studies of B cell marker gene signatures in head and neck squamous cell carcinoma have shown that higher TCR and BCR diversity correlates with immune checkpoint inhibitor sensitivity, suggesting that receptor repertoire features carry prognostic information [<a href="#ref-2">2</a>]. Paired single-cell receptor sequencing provides the resolution needed to connect receptor identity to cell state and clinical outcome.

Core Principles of Library Preparation for Immune Repertoire Analysis

Preserving V(D)J Coverage

The variable regions of TCR and BCR genes are the targets of immune repertoire sequencing. These regions are assembled from V, D, and J gene segments, and the junctions between segments contain the complementarity-determining regions that contact antigen. Library preparation must capture these junctions with sufficient read depth and read length to identify the gene segments and resolve the junctional sequences.

Targeted amplicon approaches use PCR primers that anneal to conserved regions of the receptor genes, typically in the constant region and the leader or framework regions, to amplify the rearranged V(D)J sequence. The amplified product spans the junctional regions and can be sequenced at high depth. The tradeoff is that the primers must match the target genes, and primer mismatches can lead to amplification bias against certain alleles or gene segments.

Full-length approaches capture the entire receptor transcript, often as part of a whole-transcriptome single-cell RNA sequencing workflow. The receptor sequence is recovered from the full-length cDNA, and the V(D)J region is assembled from the sequencing reads. This approach preserves the natural pairing of the receptor chains because both chains are captured from the same cell, but the read depth per receptor may be lower than in a targeted approach, and the assembly of the V(D)J region from full-length reads requires robust bioinformatics.

Maintaining Pairing Information

The defining advantage of single-cell receptor sequencing is the preservation of chain pairing. In a single-cell workflow, each cell is isolated in a compartment, whether a droplet, a microwell, or a well of a plate, and the cDNA from that cell is barcoded. The barcode links all transcripts from that cell, including both chains of the TCR or BCR. After sequencing, the receptor chains are assembled and assigned to cells based on the barcode.

The choice of single-cell platform affects the scale and cost of paired receptor sequencing. Microwell-based systems capture thousands of cells in parallel and support multimodal readouts including mRNA expression, protein expression, immune repertoire, and antigen-specific T and B cell identification [<a href="#ref-3">3</a>]. Droplet-based systems such as those from 10X Genomics also capture paired receptor sequences at scale. The Parse Biosciences platform uses a different chemistry and captures a broader range of genes, including non-coding genes and pseudogenes, but shows differences in T cell subset marker detection compared to 10X [<a href="#ref-4">4</a>]. These platform differences matter for study design because they affect which genes are detected and how receptor sequences are recovered.

Connecting Receptor Identity to Cell State

Paired receptor sequencing becomes most informative when combined with transcriptomic data from the same cells. The transcriptional state of a T cell, such as exhaustion, memory, or effector differentiation, can be linked to its receptor sequence, enabling researchers to ask whether specific clonotypes are enriched in specific states. This integration requires library preparation that captures both the receptor and the broader transcriptome.

Studies of neoantigen-specific T cells in non-small cell lung cancer used coupled single-cell RNA sequencing and TCR sequencing to track T cell clones identified by their receptors. The analysis revealed that mutation-associated neoantigen-specific T cells expressed incompletely activated cytolytic programs and had hallmark transcriptional programs of tissue-resident memory cells, but with low levels of interleukin-7 receptor and reduced functional responsiveness to interleukin-7 compared to virus-specific cells [<a href="#ref-5">5</a>]. This level of insight requires both the receptor sequence and the transcriptome from the same cells.

At a Glance: Library Preparation Strategy Comparison

StrategyV(D)J CoveragePairing InformationTranscriptome DataBest Use Case
Targeted amplicon TCR/BCR enrichmentHigh depth on receptor loci, primer-dependent coveragePreserved if performed on single cellsLimited or absent unless combined with separate RNA-seqStudies focused on clonal diversity and repertoire composition where transcriptome data is secondary
Full-length single-cell RNA-seq with receptor assemblyModerate depth, depends on sequencing saturationPreserved through cell barcodesFull transcriptome capturedStudies requiring integration of receptor identity with cell state, differentiation, and functional programs
3'-directed workflows with VDJ recovery methodsVariable, depends on method such as circVDJ-seqPreserved through cell barcodesFull transcriptome captured from 3' endStudies using single-nucleus RNA-seq, ATAC + RNA multi-omics, or spatial transcriptomics where standard VDJ enrichment is unavailable
Bulk RNA-seq with computational repertoire reconstructionLow sensitivity for rare clonotypes, no pairingLost, chains reconstructed separatelyBulk transcriptomeExploratory studies where single-cell platforms are unavailable and pairing is not required

The choice between these strategies depends on the research question. If the goal is to characterize clonal diversity and identify dominant clonotypes, targeted amplicon sequencing provides high depth on the receptor loci. If the goal is to understand how receptor identity relates to cell state, full-length single-cell RNA sequencing with receptor assembly is required. If the study uses 3'-directed workflows such as single-nucleus RNA-seq or spatial transcriptomics, methods that recover VDJ information from these workflows, such as circVDJ-seq, offer a path to receptor profiling without specialized sequencing equipment [<a href="#ref-6">6</a>].

Practical Workflow: From Sample to Library

Sample Preparation and Cell Isolation

The first decision is the input material. Fresh cells are the standard input for single-cell RNA sequencing because the method requires intact cells with intact mRNA. Frozen cells can be used if the freezing protocol preserves viability and mRNA integrity. For single-nucleus RNA-seq, nuclei are isolated instead of whole cells, which enables profiling of frozen tissues but requires adaptation of the receptor recovery method.

Cell isolation should include a viability assessment. The BD Rhapsody system includes an imaging device for sample quality control and workflow quality control, including viability and multiplet assessment [<a href="#ref-3">3</a>]. Viability below acceptable thresholds increases the proportion of dying cells, which contribute ambient mRNA and can confound receptor assignment. Multiplets, where two or more cells are captured in the same compartment, create artificial receptor combinations that appear as chimeric sequences.

For studies of specific cell populations, fluorescence-activated cell sorting (FACS) can enrich for the cells of interest before single-cell capture. A protocol for identifying recirculating thymic regulatory T cells describes steps for thymic dissection, magnetic cell separation, FACS, single-cell sequencing, and downstream computational analysis [<a href="#ref-7">7</a>]. Sorting reduces the number of irrelevant cells captured and increases the yield of informative receptor sequences.

Reverse Transcription and Barcoding

The reverse transcription step converts mRNA to cDNA and incorporates the cell barcode and unique molecular identifier (UMI) into each cDNA molecule. The UMI distinguishes individual mRNA molecules, enabling digital counting of transcripts and correction of PCR amplification bias. The cell barcode links all transcripts to their cell of origin.

For targeted amplicon approaches, the reverse transcription may include a template-switching step that adds a known sequence to the 3' end of the cDNA, providing a priming site for subsequent PCR amplification. The V(D)J region is then amplified using primers specific to the receptor constant regions and the template-switch sequence. This approach enriches for receptor transcripts while preserving the cell barcode and UMI.

For full-length approaches, the reverse transcription produces full-length cDNA that is amplified and then fragmented for sequencing. The receptor sequences are recovered from the full-length reads during bioinformatics analysis. The tradeoff is that the sequencing reads are distributed across the entire transcriptome, so the depth per receptor locus is lower than in a targeted approach.

Amplification and Library Construction

The amplification step generates enough material for sequencing. PCR amplification introduces errors and bias, so the number of amplification cycles should be minimized while still producing sufficient library. The use of UMIs allows correction of PCR errors and amplification bias because each original mRNA molecule has a unique UMI, and reads sharing the same UMI and cell barcode can be collapsed into a single molecule count.

RNase H-dependent PCR is one approach for targeted TCR sequencing that provides high specificity and efficiency for TCR mRNA. This method uses RNase H to cleave the RNA strand in RNA-DNA hybrids, enabling specific priming and amplification of TCR sequences [<a href="#ref-8">8</a>]. The choice of amplification chemistry affects the specificity and efficiency of receptor enrichment.

Sequencing Considerations

The sequencing depth and read length required depend on the library preparation strategy. Targeted amplicon libraries can be sequenced at high depth because the library is enriched for receptor sequences. Full-length libraries require more sequencing to achieve adequate depth on the receptor loci because the reads are distributed across the transcriptome.

Read length matters for V(D)J assembly. The junctional regions of TCR and BCR genes contain the antigen-binding loops, and resolving these regions requires reads that span the junctions. Paired-end sequencing with sufficient read length enables assembly of the V(D)J region from overlapping reads.

Targeted Amplicon Approaches: Strengths and Limitations

Strengths of Targeted Enrichment

Targeted amplicon approaches provide high depth on the receptor loci, enabling detection of rare clonotypes and precise resolution of junctional sequences. The enrichment for receptor transcripts reduces the sequencing cost per receptor because the library is composed primarily of receptor sequences. This approach is well suited for studies focused on clonal diversity, repertoire composition, and clonal tracking.

The specificity of the amplification depends on the primer design. Primers that anneal to conserved regions of the receptor genes amplify all rearranged receptors, while primers that anneal to specific V gene families amplify subsets of the repertoire. The choice of primers affects the coverage and bias of the resulting repertoire.

Limitations of Targeted Enrichment

The primary limitation of targeted amplicon approaches is the loss of transcriptome information. If the goal is to connect receptor identity to cell state, targeted enrichment alone does not provide the transcriptional data needed. Some workflows combine targeted receptor enrichment with whole-transcriptome amplification, but this requires additional library preparation steps and increases cost.

Primer bias is another limitation. If the primers do not match all alleles or gene segments equally, the amplification will skew the repertoire toward the well-matched sequences. This bias can be detected by comparing the observed V gene usage to expected frequencies, but it cannot be fully corrected.

When to Choose Targeted Amplicon

Targeted amplicon approaches are appropriate when the research question is primarily about the receptor repertoire itself. Studies of clonal expansion, repertoire diversity, and V gene usage can be answered with targeted enrichment. Studies that require linking receptor identity to transcriptional state, differentiation trajectory, or functional program require full-length or combined approaches.

Full-Length and Whole-Transcriptome Approaches

Capturing Receptor and Transcriptome Together

Full-length single-cell RNA sequencing captures the entire transcriptome, including the TCR and BCR transcripts. The receptor sequences are assembled from the full-length reads during bioinformatics analysis. This approach preserves the natural pairing of receptor chains because both chains are captured from the same cell and linked by the cell barcode.

The advantage of this approach is the integration of receptor identity with cell state. Studies of the breast tumor microenvironment used paired single-cell RNA and TCR sequencing data from 27,000 T cells to reveal the combinatorial impact of TCR utilization on phenotypic diversity [<a href="#ref-9">9</a>]. The analysis showed continuous phenotypic expansions specific to the tumor microenvironment and supported a model of continuous activation in T cells. This type of insight requires both the receptor sequence and the transcriptome from the same cells.

Computational Challenges of Full-Length Assembly

The assembly of V(D)J regions from full-length reads is computationally demanding. The reads must be aligned to the receptor gene loci, the V, D, and J segments must be identified, and the junctional sequences must be resolved. Errors in assembly can create false clonotypes, so quality filtering is essential.

The depth of coverage on the receptor loci depends on the sequencing saturation. If the library is sequenced to a depth that provides adequate coverage of the transcriptome, the receptor loci may be covered at lower depth than in a targeted approach. This can reduce sensitivity for rare clonotypes.

3'-Directed Workflows and VDJ Recovery

Many single-cell and spatial transcriptomics workflows use 3'-directed barcoding of cDNA, which captures the 3' ends of transcripts instead of full-length transcripts. These workflows are efficient and cost-effective for gene expression profiling, but they do not capture the V(D)J regions, which are located at the 5' ends of the receptor transcripts.

circVDJ-seq addresses this limitation by enabling T cell clonotype detection from 3'-directed workflows. The method recovers VDJ information from single-cell and spatial transcriptomics workflows with 3'-barcoding of cDNA, including single-nucleus RNA sequencing, ATAC + RNA multi-omics, and spatial transcriptomics. Application of circVDJ-seq to freshly resected neuroblastomas and postmortem lymph nodes affected by pneumonia or COVID-19 revealed distinct immune microenvironments and T cell clonality patterns [<a href="#ref-6">6</a>]. This approach expands the range of studies that can include receptor profiling without requiring specialized sequencing equipment.

When to Choose Full-Length or 3'-Directed Approaches

Full-length approaches are appropriate when the research question requires integration of receptor identity with cell state. Studies of T cell differentiation, exhaustion, memory formation, and functional programs benefit from the transcriptome data. 3'-directed workflows with VDJ recovery methods are appropriate when the study uses single-nucleus RNA-seq, ATAC + RNA multi-omics, or spatial transcriptomics, where standard VDJ enrichment is not available.

Platform Selection and Comparative Performance

10X Genomics and Parse Biosciences

The choice of single-cell platform affects the quality and characteristics of the resulting data. A comparative analysis of single-cell TCR sequencing using 10X Genomics and Parse Biosciences platforms found that both platforms generated high-quality data and captured comparable TCR clonal landscapes, with strong concordance for dominant clones. However, more genes were captured with Parse, including a broader range of non-coding genes and pseudogenes. Although average gene expression was highly correlated across the platforms, feature selection yielded different sets of selected genes, and many T cell subset markers varied substantially between platforms. Genes enriched in Parse were typically longer and associated with naive and central-memory T cell genes, whereas those enriched in 10X were shorter and often associated with cytotoxic and effector genes [<a href="#ref-4">4</a>].

These differences have practical implications. If the study prioritizes a wider range of genes, including non-coding genes, Parse offers an advantage. If the study requires detailed characterization of effector function, 10X may be the apparent choice. The platform selection should be aligned with the study objectives.

Microwell-Based Systems

Microwell-based systems such as the BD Rhapsody capture cells in microwell cartridges and support multimodal readouts. The system enables capture of mRNA expression, protein expression, immune repertoire for TCR and BCR, and identification of antigen-specific T and B cells using dCODE Dextramer reagents. The imaging device provides sample quality control and workflow quality control, including viability and multiplet assessment [<a href="#ref-3">3</a>].

The multimodal capability of microwell-based systems is valuable for studies that require protein expression data alongside receptor sequences. The ability to identify antigen-specific cells using dextramer reagents adds a functional dimension to the receptor analysis.

Platform Selection Criteria

Platform selection should consider the following criteria:

  • Scale: How many cells need to be captured? Droplet-based and microwell-based systems capture thousands to tens of thousands of cells, while plate-based systems capture fewer cells with higher per-cell cost.
  • Multimodal requirements: Does the study require protein expression, antigen specificity, or other modalities beyond RNA and receptor sequences?
  • Gene detection breadth: Does the study require detection of non-coding genes or specific gene classes that may be captured differently across platforms?
  • Cost: The per-cell cost varies across platforms and affects the feasible scale of the study.
  • Reproducibility: The platform should produce consistent results across replicates and batches.

Bioinformatics Workflow for Receptor Assembly and Analysis

Quality Control of Single-Cell Data

The first step in the bioinformatics workflow is quality control of the single-cell data. This includes filtering cells based on the number of genes detected, the number of UMIs, and the proportion of mitochondrial reads. Cells with low gene counts may be dying or empty, while cells with high mitochondrial proportions may be stressed. The quality control thresholds should be set based on the data distribution and the expected biology of the sample.

The NCBI provides data resources for sequence analysis, including databases for raw sequencing data and processed results [<a href="#ref-10">10</a>]. The EMBL-EBI Training program offers learning pathways for bioinformatics data resources and practical analysis education [<a href="#ref-11">11</a>]. These resources support researchers in developing the computational skills needed for single-cell data analysis.

Receptor Assembly and Clonotype Calling

The assembly of receptor sequences from single-cell data requires specialized software. The reads are aligned to the receptor gene loci, and the V, D, and J segments are identified. The junctional sequences are resolved, and the receptor chains are paired based on the cell barcode.

The quality of the assembly depends on the read depth and the accuracy of the alignment. Low-depth cells may have incomplete receptor sequences, and alignment errors can create false clonotypes. The clonotype calling should include filters for read support and sequence quality.

Clonality and Diversity Metrics

The clonality and diversity of the repertoire are quantified using metrics such as the number of unique clonotypes, the frequency of the dominant clonotypes, and diversity indices that account for both richness and evenness. These metrics provide a summary of the repertoire composition and enable comparison across samples.

Studies of juvenile idiopathic arthritis used the computational method TRUST4 to construct TCR and BCR repertoires from bulk RNA-seq data and analyzed the clonality and diversity of the immune repertoire. The findings revealed significant differences in the frequency of clonotypes between the JIA and healthy control groups, and specific V and J genes were identified that could expand the understanding of JIA [<a href="#ref-12">12</a>]. This approach demonstrates the use of computational repertoire reconstruction from bulk data, although it does not preserve pairing information.

Integration of Receptor and Transcriptome Data

The integration of receptor sequences with transcriptome data enables analysis of the relationship between clonotype and cell state. Cells can be grouped by clonotype, and the transcriptional programs of different clonotypes can be compared. This analysis can reveal whether specific clonotypes are enriched in specific cell states, such as exhaustion or memory.

The TCR-RNA Integrating Model (TRIM) is a multimodal variational autoencoder framework that integrates RNA and TCR data and predicts T cell clonality and transcriptional states. TRIM learns a shared representation of the data conditioned on patient, tissue source, and treatment timepoint. Applied to datasets from patients with head and neck squamous cell carcinoma and colorectal cancer, TRIM accurately predicted intra-tumor T cell clonal expansion and transcriptional status based on T cells from blood or normal tissue before treatment [<a href="#ref-13">13</a>]. This type of integrative analysis requires library preparation that captures both receptor and transcriptome data.

Records and Measurements for Quality Assurance

Documentation of Library Preparation

The library preparation should be documented in sufficient detail to enable reproducibility. This includes the input cell count, viability, the reverse transcription conditions, the amplification cycles, the library quantification results, and the sequencing parameters. The documentation should also include the software versions and parameters used for receptor assembly and analysis.

The nf-core documentation provides standards for community pipeline usage, configuration, and reproducible workflow context [<a href="#ref-14">14</a>]. The Galaxy Training Network offers accessible workflow training and analysis tutorials [<a href="#ref-15">15</a>]. These resources support reproducible analysis practices.

Quality Metrics to Track

The following quality metrics should be tracked for each library:

  • Cell count and viability before capture
  • Number of cells captured and passing quality filters
  • Median genes and UMIs per cell
  • Proportion of reads mapping to receptor loci
  • Number of cells with paired receptor sequences
  • Number of unique clonotypes and clonality metrics
  • Proportion of reads with valid cell barcodes and UMIs

These metrics provide a record of the library quality and enable comparison across batches. Deviations from expected ranges should be investigated before proceeding with downstream analysis.

Batch Effects and Replicates

Batch effects can confound comparisons across samples processed in different batches. The library preparation should include replicates and appropriate experimental design to account for batch effects. The use of multiplexing, where multiple samples are pooled in a single run, can reduce batch effects but requires careful design of the sample barcodes.

The Carpentries lessons provide foundational training in computing, data, shell, Git, and programming [<a href="#ref-16">16</a>]. These skills support the management and analysis of large single-cell datasets.

Common Failure Patterns and Troubleshooting

Low Cell Viability

Low cell viability before capture reduces the number of informative cells and increases ambient mRNA. The ambient mRNA can be captured in empty droplets or wells and assigned to cells, creating false signals. The viability should be assessed before capture, and samples with low viability should be re-prepared or the viability threshold should be adjusted.

High Multiplet Rate

Multiplets create artificial receptor combinations that appear as chimeric sequences. The multiplet rate should be estimated from the data, and cells with evidence of multiplets should be filtered. The imaging quality control available in some systems can identify multiplets before sequencing [<a href="#ref-3">3</a>].

Primer Bias in Targeted Amplification

Primer bias in targeted amplification skews the repertoire toward well-matched sequences. This bias can be detected by comparing the observed V gene usage to expected frequencies. If the bias is severe, the primers should be redesigned or the analysis should account for the bias.

Low Receptor Capture Rate

The receptor capture rate is the proportion of cells with a successfully assembled receptor sequence. Low capture rates can result from low expression of receptor transcripts, incomplete reverse transcription, or insufficient sequencing depth. The capture rate should be monitored, and the library preparation should be optimized if the capture rate is low.

Chimeric Sequence Artifacts

Chimeric sequences can arise from PCR recombination during amplification. These artifacts create false clonotypes that are not present in the original sample. The use of UMIs enables correction of PCR errors and amplification bias, but chimeric sequences can still be problematic. The analysis should include filters for read support and sequence quality to remove chimeric sequences.

Limitations and Interpretation Boundaries

Sensitivity for Rare Clonotypes

The sensitivity for rare clonotypes depends on the sequencing depth and the number of cells captured. Rare clonotypes present at low frequency may not be detected, and the absence of a clonotype does not prove its absence from the sample. The sensitivity should be estimated from the data, and the interpretation should account for the detection limits.

Platform-Specific Biases

Different platforms capture different genes and may have different biases in receptor amplification. The comparative analysis of 10X Genomics and Parse Biosciences platforms found differences in gene capture and T cell subset marker detection [<a href="#ref-4">4</a>]. These platform-specific biases should be considered when comparing results across platforms or when selecting a platform for a specific study.

Computational Assembly Errors

The assembly of receptor sequences from sequencing reads is subject to errors. Alignment errors can create false clonotypes, and incomplete assembly can miss true clonotypes. The assembly quality should be assessed, and the interpretation should account for the error rate.

Descriptive Studies and Hypothesis Generation

Some single-cell studies are descriptive and generate hypotheses instead of testing causal mechanisms. A study of T cells in aneurysmal subarachnoid hemorrhage analyzed transcriptomic profiles and TCR repertoires in paired cerebrospinal fluid and peripheral blood from two patients. The study found that CSF-derived T cells showed stronger signaling pathway activation and increased clonal expansion than PB-derived T cells, but the authors noted that the observations are limited to these two individuals and cannot be generalized without large independent cohort validation [<a href="#ref-17">17</a>]. Descriptive studies provide directions for future mechanistic investigations but do not establish causal relationships.

Safety and Regulatory Context

Data Management and Privacy

Single-cell sequencing data from human samples may contain sensitive information. The data should be managed according to applicable regulations and institutional policies. The NCBI provides data resources for sequence data, and the submission of human data may require compliance with privacy and consent requirements [<a href="#ref-10">10</a>].

Reagent and Protocol Safety

The reagents used in library preparation, including enzymes, buffers, and primers, should be handled according to the manufacturer's safety data sheets. The protocols should be performed in appropriate laboratory conditions, and waste should be disposed of according to institutional guidelines.

Reproducibility Standards

Reproducibility is a core requirement for scientific research. The library preparation and analysis should be documented in sufficient detail to enable replication. The use of standardized workflows and pipelines supports reproducibility. The nf-core documentation provides standards for community pipeline usage and configuration [<a href="#ref-14">14</a>], and the Galaxy Training Network offers accessible workflow training [<a href="#ref-15">15</a>].

Professional Escalation Criteria

When to Consult a Bioinformatics Specialist

The bioinformatics analysis of single-cell receptor data requires specialized skills. If the analysis exceeds the expertise of the research team, a bioinformatics specialist should be consulted. The EMBL-EBI Training program offers learning pathways for bioinformatics data resources and practical analysis education [<a href="#ref-11">11</a>], and the Bioconductor project provides official package, workflow, installation, and reproducible genomic-analysis documentation [<a href="#ref-18">18</a>].

When to Seek Technical Support from Platform Vendors

Platform-specific issues, such as low capture rates or unexpected bias, may require technical support from the platform vendor. The BD Rhapsody system provides updated protocols, guides, and technical bulletins through the BD Scomix page and the BDB webpage [<a href="#ref-3">3</a>]. Technical support can help troubleshoot platform-specific issues.

When to Reconsider the Experimental Design

If the data quality is consistently poor across multiple libraries, the experimental design should be reconsidered. This may involve changing the input material, the cell isolation method, the library preparation strategy, or the sequencing parameters. The decision to redesign should be based on the quality metrics and the research question.

Decision Framework for Matching Library Preparation to Study Objectives

Selecting a library preparation strategy requires a structured decision process that connects the biological question to the technical capabilities of each method. A practical framework organizes the decision around five sequential criteria: pairing requirement, transcriptome need, input material, scale, and budget. Working through these criteria in order prevents the common error of choosing a platform before defining what the data must contain.

Step 1: Define the Pairing Requirement

The first decision is whether the study requires paired receptor chains from the same cell. Studies that track clonal expansion across tissues, such as comparing tumor-infiltrating lymphocytes with peripheral blood, need paired information to identify the same clone in different compartments. Studies of hepatocellular carcinoma used deep single-cell RNA sequencing with assembled TCR sequences to identify T cell subsets and delineate their developmental trajectory, which required knowing which TCR chains belonged to which cell [<a href="#ref-1">1</a>]. If the research question only requires V gene usage frequencies or clonotype diversity at the population level, bulk approaches with computational reconstruction may suffice, but the loss of pairing information must be accepted.

Step 2: Assess the Transcriptome Requirement

The second criterion is whether receptor identity must be linked to transcriptional state. Studies that ask how clonotype relates to exhaustion, memory, or effector function require full-length or combined approaches that capture both receptor and transcriptome. A study of breast tumors used paired single-cell RNA and TCR sequencing data from 27,000 T cells to reveal the combinatorial impact of TCR utilization on phenotypic diversity [<a href="#ref-9">9</a>]. If the study only needs receptor sequences without cell state information, targeted amplicon approaches provide higher depth on the receptor loci at lower cost.

Step 3: Evaluate Input Material Constraints

The input material determines which library preparation strategies are feasible. Fresh cells are required for standard single-cell RNA sequencing workflows. Frozen cells can be used if the freezing protocol preserves viability and mRNA integrity, but the quality should be verified before committing to a full experiment. Single-nucleus RNA-seq enables profiling of frozen tissues but requires adaptation of the receptor recovery method. The BD Rhapsody system includes an imaging device for sample quality control and workflow quality control, including viability and multiplet assessment [<a href="#ref-3">3</a>]. If the input material is limited or of uncertain quality, the library preparation strategy should include additional quality checkpoints.

Step 4: Determine Scale and Throughput

The number of cells needed for the study affects platform selection. Droplet-based and microwell-based systems capture thousands to tens of thousands of cells in parallel. Plate-based systems capture fewer cells with higher per-cell cost. The scale requirement should be determined by the expected frequency of the cell populations of interest. Rare populations require more cells to achieve adequate representation. A study of primary Sjogren syndrome applied single-cell RNA sequencing to 57,288 peripheral blood mononuclear cells from five patients and five healthy controls to identify disease-specific immune cell subsets [<a href="#ref-19">19</a>]. The scale of this study was driven by the need to detect expanded subpopulations within a heterogeneous mixture.

Step 5: Apply Budget Constraints

The final criterion is the budget available for the study. Targeted amplicon approaches are generally more cost-efficient per cell when transcriptome data is not required. Full-length approaches require more sequencing to achieve adequate depth on the receptor loci because reads are distributed across the transcriptome. The budget should be allocated based on the number of cells, the sequencing depth, and the number of samples. A comparative analysis of 10X Genomics and Parse Biosciences platforms found that both generated high-quality data with comparable TCR clonal landscapes, but the platforms differ in gene capture and cost structure [<a href="#ref-4">4</a>]. The budget constraint should be applied after the biological requirements are defined, not before.

Record System for Library Preparation Decisions

A standardized record system supports reproducibility and enables troubleshooting when problems arise. The record should capture the decision criteria, the chosen strategy, and the quality metrics at each step.

Decision Log

The decision log documents the rationale for the chosen library preparation strategy. For each study, record the pairing requirement, the transcriptome requirement, the input material, the scale, and the budget. This log enables other researchers to understand why a particular approach was chosen and to evaluate whether the choice was appropriate for the research question.

Library Preparation Record

The library preparation record captures the technical details of each library. Include the input cell count, viability percentage, reverse transcription conditions, amplification cycles, library quantification results, and sequencing parameters. The record should also include the software versions and parameters used for receptor assembly and analysis. The nf-core documentation provides standards for community pipeline usage, configuration, and reproducible workflow context [<a href="#ref-14">14</a>]. The Galaxy Training Network offers accessible workflow training and analysis tutorials that support reproducible analysis practices [<a href="#ref-15">15</a>].

Quality Metrics Dashboard

The quality metrics dashboard tracks the key indicators of library quality across batches. The following metrics should be recorded for each library:

  • Cell count and viability before capture
  • Number of cells captured and passing quality filters
  • Median genes and UMIs per cell
  • Proportion of reads mapping to receptor loci
  • Number of cells with paired receptor sequences
  • Number of unique clonotypes and clonality metrics
  • Proportion of reads with valid cell barcodes and UMIs

These metrics provide a record of the library quality and enable comparison across batches. Deviations from expected ranges should be investigated before proceeding with downstream analysis.

Troubleshooting Method for Low Receptor Capture Rates

Low receptor capture rates are a common failure pattern in single-cell receptor sequencing. The capture rate is the proportion of cells with a successfully assembled receptor sequence. A systematic troubleshooting method isolates the cause and guides corrective action.

Step 1: Verify Input Cell Quality

The first check is the quality of the input cells. Low viability before capture reduces the number of informative cells and increases ambient mRNA. The ambient mRNA can be captured in empty droplets or wells and assigned to cells, creating false signals. The viability should be assessed before capture, and samples with low viability should be re-prepared or the viability threshold should be adjusted.

Step 2: Check Reverse Transcription Efficiency

The reverse transcription step converts mRNA to cDNA and incorporates the cell barcode and UMI. Inefficient reverse transcription reduces the amount of cDNA available for receptor amplification. The efficiency can be assessed by comparing the number of genes detected per cell to expected values for the cell type. If the gene detection is low across all cells, the reverse transcription conditions should be optimized.

Step 3: Evaluate Amplification Specificity

The amplification step enriches for receptor transcripts. If the primers do not match the target genes efficiently, the amplification will produce low yields of receptor sequences. RNase H-dependent PCR is one approach for targeted TCR sequencing that provides high specificity and efficiency for TCR mRNA [<a href="#ref-8">8</a>]. The amplification specificity can be assessed by examining the proportion of reads mapping to receptor loci. If the proportion is low, the primers or amplification conditions should be reviewed.

Step 4: Assess Sequencing Depth

The sequencing depth affects the recovery of receptor sequences. Low depth on the receptor loci reduces the sensitivity for rare clonotypes and can result in incomplete receptor assembly. The depth should be sufficient to provide adequate coverage of the receptor loci given the library composition. If the depth is insufficient, additional sequencing should be performed.

Step 5: Review Bioinformatics Parameters

The bioinformatics analysis parameters affect the receptor assembly and clonotype calling. The alignment parameters, the quality filters, and the clonotype calling thresholds should be reviewed. The Bioconductor project provides official package, workflow, installation, and reproducible genomic-analysis documentation [<a href="#ref-18">18</a>]. The EMBL-EBI Training program offers learning pathways for bioinformatics data resources and practical analysis education [<a href="#ref-11">11</a>]. These resources support the development of appropriate analysis parameters.

Platform Comparison Decision Matrix

The platform comparison decision matrix translates the decision framework into a practical tool for selecting between major platforms. The matrix compares platforms across the criteria that matter for library preparation decisions.

Criterion10X GenomicsParse BiosciencesBD Rhapsody
Pairing informationPreserved through cell barcodesPreserved through cell barcodesPreserved through cell barcodes
Transcriptome captureFull transcriptome, shorter gene biasFull transcriptome, broader gene capture including non-coding genesFull transcriptome with multimodal protein and antigen readouts
Input materialFresh cells, frozen cells with optimizationFresh cells, frozen cells with optimizationFresh cells with imaging quality control
ScaleThousands to tens of thousands of cellsThousands to tens of thousands of cellsThousands of cells with imaging quality control
Multimodal capabilityRNA and TCR or BCRRNA and TCR or BCRRNA, protein, TCR, BCR, and antigen-specific detection
Gene detection characteristicsShorter genes associated with cytotoxic and effector markersLonger genes associated with naive and central-memory markersPlatform-specific capture chemistry

The comparative analysis of 10X Genomics and Parse Biosciences platforms found that genes enriched in Parse were typically longer and associated with naive and central-memory T cell genes, whereas those enriched in 10X were shorter and often associated with cytotoxic and effector genes [<a href="#ref-4">4</a>]. These differences should be considered when selecting a platform for studies that depend on specific gene classes.

Common Failure Patterns and Corrective Actions

Pattern 1: Consistent Low Viability Across Samples

Low viability across multiple samples suggests a systematic issue with the sample preparation protocol. The cell isolation method, the storage conditions, or the reagents may be the cause. The protocol should be reviewed, and the viability should be assessed at each step to identify the point of loss.

Pattern 2: High Multiplet Rate

Multiplets create artificial receptor combinations that appear as chimeric sequences. The multiplet rate should be estimated from the data, and cells with evidence of multiplets should be filtered. The imaging quality control available in some systems can identify multiplets before sequencing [<a href="#ref-3">3</a>]. If the multiplet rate is consistently high, the cell concentration loaded into the capture system should be reduced.

Pattern 3: Primer Bias in Targeted Amplification

Primer bias in targeted amplification skews the repertoire toward well-matched sequences. This bias can be detected by comparing the observed V gene usage to expected frequencies. If the bias is severe, the primers should be redesigned or the analysis should account for the bias.

Pattern 4: Batch Effects Across Samples

Batch effects can confound comparisons across samples processed in different batches. The library preparation should include replicates and appropriate experimental design to account for batch effects. The use of multiplexing, where multiple samples are pooled in a single run, can reduce batch effects but requires careful design of the sample barcodes.

Pattern 5: Chimeric Sequence Artifacts

Chimeric sequences can arise from PCR recombination during amplification. These artifacts create false clonotypes that are not present in the original sample. The use of UMIs enables correction of PCR errors and amplification bias, but chimeric sequences can still be problematic. The analysis should include filters for read support and sequence quality to remove chimeric sequences.

Professional Escalation Criteria

When to Consult a Bioinformatics Specialist

The bioinformatics analysis of single-cell receptor data requires specialized skills. If the analysis exceeds the expertise of the research team, a bioinformatics specialist should be consulted. The EMBL-EBI Training program offers learning pathways for bioinformatics data resources and practical analysis education [<a href="#ref-11">11</a>], and the Bioconductor project provides official package, workflow, installation, and reproducible genomic-analysis documentation [<a href="#ref-18">18</a>].

When to Seek Technical Support from Platform Vendors

Platform-specific issues, such as low capture rates or unexpected bias, may require technical support from the platform vendor. The BD Rhapsody system provides updated protocols, guides, and technical bulletins through the BD Scomix page and the BDB webpage [<a href="#ref-3">3</a>]. Technical support can help troubleshoot platform-specific issues.

When to Reconsider the Experimental Design

If the data quality is consistently poor across multiple libraries, the experimental design should be reconsidered. This may involve changing the input material, the cell isolation method, the library preparation strategy, or the sequencing parameters. The decision to redesign should be based on the quality metrics and the research question. A study of T cells in aneurysmal subarachnoid hemorrhage analyzed transcriptomic profiles and TCR repertoires in paired cerebrospinal fluid and peripheral blood from two patients, but the authors noted that the observations are limited to these two individuals and cannot be generalized without large independent cohort validation [<a href="#ref-17">17</a>]. This example illustrates the importance of matching the experimental design to the scope of the conclusions that can be drawn.

Frequently Asked Questions

What is the difference between targeted amplicon and full-length approaches for TCR and BCR sequencing?

Targeted amplicon approaches use PCR primers to enrich for receptor transcripts, providing high depth on the receptor loci but limited transcriptome information. Full-length approaches capture the entire transcriptome, including the receptor transcripts, enabling integration of receptor identity with cell state. The choice depends on whether the research question requires transcriptome data beyond the receptor sequences.

Why is paired TCR or BCR information important?

The TCR alpha and beta chains, or the BCR heavy and light chains, together determine antigen specificity. Paired information preserves the natural combination of chains expressed by each cell, which is required to reconstruct functional receptors and to link receptor identity to cell state. Bulk repertoire sequencing loses this pairing information.

How do I choose between 10X Genomics and Parse Biosciences platforms?

The choice depends on the study objectives. Both platforms generate high-quality data and capture comparable TCR clonal landscapes. Parse captures more genes, including non-coding genes and pseudogenes, while 10X may be better for detailed characterization of effector function. The platform selection should be aligned with the study objectives [<a href="#ref-4">4</a>].

What is circVDJ-seq and when should I use it?

circVDJ-seq is a method for T cell clonotype detection from 3'-directed workflows such as single-nucleus RNA sequencing, ATAC + RNA multi-omics, and spatial transcriptomics. It enables receptor profiling from workflows where standard VDJ enrichment is not available, without requiring specialized sequencing equipment [<a href="#ref-6">6</a>].

How do I assess the quality of my single-cell receptor data?

Quality metrics include cell count and viability before capture, number of cells passing quality filters, median genes and UMIs per cell, proportion of reads mapping to receptor loci, number of cells with paired receptor sequences, and clonality metrics. These metrics should be tracked for each library and compared across batches.

What are the main limitations of single-cell TCR and BCR sequencing?

The main limitations include sensitivity for rare clonotypes, platform-specific biases, computational assembly errors, and the descriptive nature of many studies. The interpretation should account for these limitations, and the findings should be validated in independent cohorts where possible.

How do I integrate receptor sequences with transcriptome data?

Integration requires library preparation that captures both receptor and transcriptome data. The receptor sequences are assembled and assigned to cells based on the cell barcode, and the transcriptional programs of different clonotypes are compared. Integrative analysis methods such as TRIM can predict T cell clonality and transcriptional states from multimodal data [<a href="#ref-13">13</a>].

What training resources are available for single-cell bioinformatics?

The EMBL-EBI Training program offers learning pathways for bioinformatics data resources and practical analysis education [<a href="#ref-11">11</a>]. The Galaxy Training Network provides accessible workflow training and analysis tutorials [<a href="#ref-15">15</a>]. The Carpentries lessons provide foundational training in computing, data, shell, Git, and programming [<a href="#ref-16">16</a>]. The Bioconductor project provides official package and workflow documentation [<a href="#ref-18">18</a>].

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [Landscape of Infiltrating T Cells in Liver Cancer Revealed by Single-Cell Sequencing.](https://pubmed.ncbi.nlm.nih.gov/28622514). Cell, 2017. [2] [Comprehensive analysis of single-cell and bulk RNA-sequencing data identifies B cell marker genes signature that predicts prognosis and analysis of immune checkpoints expression in head and neck squamous cell carcinoma](https://doi.org/10.1016/j.heliyon.2023.e22656). Heliyon, 2023. [3] [BD Rhapsody™ Single-Cell Analysis System Workflow: From Sample to Multimodal Single-Cell Sequencing Data.](https://pubmed.ncbi.nlm.nih.gov/36495444). Methods in molecular biology (Clifton, N.J.), 2023. [4] [Immune profiling of T cells: A comparative analysis of single-cell TCR sequencing using 10X Genomics and Parse Biosciences platforms](https://doi.org/10.1016/j.bbrep.2026.102592). Biochemistry and Biophysics Reports, 2026. [5] [Transcriptional programs of neoantigen-specific TIL in anti-PD-1-treated lung cancers.](https://pubmed.ncbi.nlm.nih.gov/34290408). Nature, 2021. [6] [circVDJ-seq for T cell clonotype detection in single-cell and spatial multi-omics.](https://doi.org/10.1186/s13073-026-01691-1). 2026. [7] [Protocol for identifying recirculating thymic regulatory T cells and characterizing the role of Eos in their function using scRNA-seq and TCR-seq.](https://doi.org/10.1016/j.xpro.2026.104620). 2026. [8] [RNase H-dependent PCR-enabled T-cell receptor sequencing for highly specific and efficient targeted sequencing of T-cell receptor mRNA for single-cell and repertoire analysis](https://doi.org/10.1038/s41596-019-0195-x). Nature Protocols, 2019. [9] [Single-Cell Map of Diverse Immune Phenotypes in the Breast Tumor Microenvironment.](https://pubmed.ncbi.nlm.nih.gov/29961579). Cell, 2018. [10] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [11] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [12] [Comprehensive analysis of juvenile idiopathic arthritis patients’ immune characteristics based on bulk and single-cell sequencing data](https://doi.org/10.3389/fmolb.2024.1359235). Frontiers in Molecular Biosciences, 2024. [13] [Multimodal framework for the joint analysis of single-cell RNA and T cell receptor sequencing data predicts T cell response to cancer immunotherapy](https://doi.org/10.1038/s41467-026-70505-0). Nature Communications, 2026. [14] [nf-core Documentation](https://nf-co.re/docs). nf-core. [15] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [16] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [17] [Integrated analysis of scRNA-seq and TCR-seq on T cells in aneurysmal subarachnoid hemorrhage: A descriptive, hypothesis-generating case study.](https://doi.org/10.1097/md.0000000000049817). 2026. [18] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [19] [Single-Cell RNA Sequencing Reveals the Expansion of Cytotoxic CD4(+) T Lymphocytes and a Landscape of Immune Cells in Primary Sjögren's Syndrome.](https://pubmed.ncbi.nlm.nih.gov/33603736). Frontiers in immunology, 2020.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.