# Phasing Strategies for Complex Genomes: A Comparison of Hi-C, Strand-seq, and Linked-Read Approaches


## Key Takeaways

- **Hi-C** leverages proximity ligation of chromatin to generate a statistical phasing signal based on interaction frequencies between homologous chromosomes, offering dual utility for both scaffolding and phasing from a single library preparation. Its accuracy is dependent on sufficient heterozygous variant density and sequencing depth, with limitations in regions of low heterozygosity or high repetitiveness.

- **Strand-seq** utilizes single-cell DNA replication to directly observe template strand segregation, providing a deterministic global phase signal that is independent of variant density and outperforms Hi-C in phasing accuracy, particularly for complex or low-heterozygosity genomes. However, it requires more complex single-cell laboratory workflows and scales in cost with genome size due to increased cell requirements.

- **Linked-read** approaches employ microfluidic partitioning to barcode DNA molecules, enabling phasing based on the co-occurrence of variants within shared barcodes, and are cost-effective for standard short-read sequencing. Their primary limitation is the phasing distance, which is constrained by the length of the input high-molecular-weight DNA molecules, making them less suitable for chromosome-scale phasing of large genomes.

- **Genome size and heterozygosity** are critical factors in technology selection: linked-reads are best for smaller genomes, while Strand-seq offers superior chromosome-scale phasing for large genomes and low-heterozygosity scenarios. Hi-C provides a balanced approach for medium to large genomes, especially when scaffolding is also a priority.

- **Practical workflow considerations** highlight distinct requirements: Hi-C needs intact nuclei, Strand-seq demands single-cell suspension and processing, and linked-reads necessitate high-molecular-weight purified DNA. Each technology has specific quality control metrics and common failure patterns, such as low valid junction rates in Hi-C, insufficient cell numbers in Strand-seq, or short DNA molecules in linked-reads.

---

Researchers planning a de novo genome assembly project must decide which phasing technology to use before generating data. This decision affects total project cost, sequencing throughput requirements, and the accuracy of haplotype-resolved assemblies. Hi-C, Strand-seq, and linked-read approaches each use different molecular biology to connect sequence reads to their parental chromosome of origin, and each has distinct tradeoffs that become more pronounced as genome size and heterozygosity increase. This article compares these three phasing strategies across cost, throughput, phasing accuracy, and practical workflow considerations, with recommendations for different genome sizes and heterozygosity levels.

## Scope and Decision Context for Phasing Technology Selection

Phasing refers to the assignment of genetic variants to their respective homologous chromosomes. For diploid organisms, this means determining which alleles are inherited together on the maternal chromosome and which are on the paternal chromosome. Haplotype information is crucial for biomedical and population genetics research because it enables the study of allele-specific expression, compound heterozygosity, and evolutionary history [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. For polyploid species, phasing becomes more complex because more than two homologous copies of each chromosome must be distinguished [<a href="#ref-3">3</a>].

The choice of phasing technology is connected to the assembly strategy. Long-read sequencing platforms such as Pacific Biosciences and Oxford Nanopore produce reads that can span repetitive regions and provide contiguous assembly graphs, but these reads alone do not reliably phase haplotypes across large genomic distances [<a href="#ref-4">4</a>]. Phasing technologies add long-range information that connects variants across distances that exceed individual read lengths. The three approaches compared here differ in how they generate this long-range connectivity.

Hi-C uses proximity ligation to capture physical interactions between chromatin regions that are close in three-dimensional nuclear space. Strand-seq uses single-cell DNA replication to identify the template strand of each chromosome. Linked-read approaches use microfluidic partitioning to tag short reads that originate from the same high-molecular-weight DNA molecule. Each method produces a different type of phasing signal, and the quality of that signal depends on genome characteristics that vary between species and even between individuals.

## Core Principles of Hi-C Phasing

Hi-C sequencing captures chromatin conformation by crosslinking DNA within intact nuclei, digesting the crosslinked DNA, ligating the fragments that are physically close in three-dimensional space, and then sequencing the ligated junctions. The resulting read pairs connect genomic loci that are near each other in the nucleus, which includes loci on the same chromosome that are separated by large linear distances. This long-range connectivity is the basis for both scaffolding and phasing.

For phasing, Hi-C exploits the observation that homologous chromosomes occupy distinct territories within the nucleus. Interactions between loci on the same homolog are more frequent than interactions between loci on different homologs. By analyzing the pattern of interaction frequencies across a set of heterozygous variants, it is possible to assign variants to one of two haplotypes based on which interactions are enriched.

The practical advantage of Hi-C is that it provides both scaffolding and phasing information from a single library preparation. The same data that orders and orients contigs into chromosome-scale scaffolds can be used to phase variants across those scaffolds. This dual utility reduces the number of separate experiments required for a chromosome-scale haplotype-resolved assembly.

The main limitation of Hi-C phasing is that the phasing signal is statistical instead of deterministic. The interaction frequency differences between homologous chromosomes are subtle, and the accuracy of phasing depends on sequencing depth and the density of heterozygous variants. In regions with low heterozygosity, there may be insufficient variant density to establish a reliable phase signal. In repetitive regions, the mapping of Hi-C reads can be ambiguous, which reduces the effective coverage available for phasing.

## Core Principles of Strand-seq Phasing

Strand-seq is a single-cell sequencing technique that preserves the directionality of DNA template strands. During DNA replication, each daughter cell receives one newly synthesized strand and one template strand. By sequencing single cells at low coverage and identifying the direction of read alignment, it is possible to determine which homolog was inherited from which parent for each chromosome region.

The phasing signal in Strand-seq comes from the segregation pattern of template strands across many single cells. Each cell contributes a sparse sampling of one homolog or the other at each genomic region. By combining information across hundreds of cells, it is possible to build a global phase that spans entire chromosomes. This global phase signal is fundamentally different from the local statistical signal of Hi-C because it directly observes the physical segregation of homologous chromosomes.

Graphasing is a workflow that synthesizes the global phase signal of Strand-seq with assembly graph topology to produce chromosome-scale de novo haplotypes for diploid genomes [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. The workflow integrates with any assembly workflow that outputs an assembly graph and has a haplotype assembly mode. Graphasing performs comparably to trio-phasing in contiguity, phasing accuracy, and assembly quality, and it outperforms Hi-C in phasing accuracy [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. In human assemblies, Graphasing generates over 18 chromosome-spanning haplotypes [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

The practical cost of Strand-seq is that it requires single-cell sorting and library preparation, which adds laboratory complexity and time compared to bulk sequencing approaches. The number of cells required depends on genome size and the desired phasing accuracy. Larger genomes require more cells to achieve the same coverage of template strand information across all chromosomes.

## Core Principles of Linked-Read Phasing

Linked-read approaches use microfluidic devices to partition high-molecular-weight DNA into thousands of nanoliter-scale reactions. Each partition contains a small amount of DNA, typically less than one haploid genome equivalent. The DNA in each partition is fragmented, barcoded with a unique molecular identifier, and then sequenced. Reads that share the same barcode originated from the same high-molecular-weight DNA molecule, which provides long-range connectivity without requiring long reads.

The phasing signal in linked-reads comes from the co-barcoding of variants that are present on the same DNA molecule. If two heterozygous variants are within the length of a single high-molecular-weight DNA molecule, they will frequently appear in the same barcode pool. By analyzing the co-occurrence of variants across barcodes, it is possible to phase variants that are separated by distances up to the length of the input DNA molecules.

The practical advantage of linked-reads is that they use standard short-read sequencing, which is widely available and relatively inexpensive per base. The library preparation requires specialized instrumentation for the microfluidic partitioning step, but the sequencing itself can be performed on any short-read platform. This makes linked-reads accessible to laboratories that do not have access to long-read or single-cell sequencing infrastructure.

The main limitation of linked-read phasing is that the phasing distance is limited by the length of the input DNA molecules. High-molecular-weight DNA extraction is critical, and the quality of the DNA preparation directly affects the achievable phasing distance. In practice, the effective phasing distance is often shorter than the theoretical maximum because of DNA fragmentation during extraction and library preparation.

## At a Glance: Technology Comparison for Phasing Decisions

The following table summarizes the key differences between Hi-C, Strand-seq, and linked-read approaches for phasing decisions. These comparisons reflect typical performance characteristics and should be evaluated in the context of a specific genome project.

| Feature | Hi-C | Strand-seq | Linked-Reads |
|---------|------|------------|--------------|
| Molecular basis | Proximity ligation of chromatin in intact nuclei | Template strand segregation in single cells | Microfluidic barcoding of high-molecular-weight DNA |
| Phasing signal type | Statistical interaction frequency differences | Direct observation of homolog segregation | Co-barcoding of variants on shared DNA molecules |
| Primary instrumentation | Standard sequencing plus crosslinking and ligation reagents | Single-cell sorting and library preparation | Microfluidic partitioning instrument |
| Phasing accuracy | Lower than Strand-seq in comparative studies | Comparable to trio-phasing, higher than Hi-C | Variable depending on DNA molecule length |
| Scaffolding utility | Provides chromosome-scale scaffolding from same data | Does not provide scaffolding information | Provides local contiguity information |
| Input DNA requirement | Intact nuclei, not purified DNA | Single-cell suspension | High-molecular-weight purified DNA |
| Genome size suitability | All sizes, but requires sufficient variant density | All sizes, but cell number scales with genome size | Best for smaller genomes due to molecule length limits |
| Heterozygosity requirement | Requires sufficient heterozygous variant density | Works with lower heterozygosity due to direct signal | Requires sufficient heterozygous variant density |
| Cost structure | Moderate library cost, standard sequencing | Higher library cost due to single-cell processing | Moderate library cost, standard sequencing |

## Cost Comparison Across Phasing Technologies

Total project cost for phasing includes library preparation reagents, sequencing, and computational analysis. The relative contribution of each component differs substantially across the three technologies.

Hi-C library preparation requires crosslinking reagents, restriction enzymes, ligases, and biotin-based pull-down reagents. The cost per library is moderate, and multiple libraries can be prepared in parallel. Sequencing depth requirements for Hi-C phasing are typically lower than for the initial genome assembly, because the phasing signal is derived from interaction frequency patterns instead of from base-level coverage. The computational cost of Hi-C data processing includes read alignment, interaction matrix construction, and phasing analysis, which requires substantial memory for large genomes.

Strand-seq library preparation is more expensive per cell because it requires single-cell sorting, whole-genome amplification, and library construction for each cell individually. The total cost scales with the number of cells sequenced, which depends on genome size and the desired phasing accuracy. Sequencing depth per cell is low, typically less than 0.1-fold coverage, but the cumulative cost across hundreds of cells can be substantial. The computational cost of Strand-seq data processing includes read alignment, template strand classification, and the Graphasing workflow, which integrates assembly graph topology with the single-cell phasing signal [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

Linked-read library preparation requires the microfluidic partitioning instrument and associated reagents. The cost per library is moderate, and the sequencing cost is similar to standard short-read sequencing. The computational cost of linked-read data processing includes read alignment, barcode deconvolution, and phasing analysis, which is generally less memory-intensive than Hi-C analysis.

For a typical diploid mammalian genome, the total cost of Strand-seq phasing is often higher than Hi-C or linked-reads because of the single-cell processing requirement. However, the higher phasing accuracy of Strand-seq may justify the additional cost for projects where haplotype resolution is the primary scientific objective [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. For projects where scaffolding is also needed, Hi-C provides two functions for a single library cost, which can make it the most cost-effective option overall.

## Throughput and Scalability Considerations

Throughput refers to the number of samples that can be processed in a given time period and the amount of sequencing data generated per sample. The three phasing technologies differ in their throughput characteristics at both the laboratory and sequencing stages.

Hi-C has moderate laboratory throughput. The crosslinking and proximity ligation steps require several days of bench work, but multiple samples can be processed in parallel. The sequencing throughput is efficient because a single Hi-C library provides both scaffolding and phasing information. For projects with many samples, Hi-C can be batched effectively, and the per-sample sequencing cost decreases with multiplexing.

Strand-seq has lower laboratory throughput because each cell must be sorted and processed individually. The number of cells required per sample depends on genome size, and processing hundreds of cells for each sample creates a bottleneck. The sequencing throughput is low per cell but high in aggregate across the full cell set. For projects with many samples, the single-cell processing requirement makes Strand-seq less scalable than Hi-C or linked-reads.

Linked-reads have high laboratory throughput because the microfluidic partitioning step is automated and processes one sample at a time with minimal hands-on time. The sequencing throughput is similar to standard short-read sequencing. For projects with many samples, linked-reads can be processed rapidly, but the phasing distance limitation may reduce the utility of the data for large genomes.

The scalability of each technology also depends on the availability of specialized instrumentation. Hi-C requires only standard molecular biology equipment. Strand-seq requires a cell sorter or a microfluidic single-cell processing instrument. Linked-reads require the specific microfluidic partitioning instrument. Laboratories that do not have access to these instruments must factor in the cost and time of sending samples to a core facility or commercial provider.

## Phasing Accuracy Across Genome Sizes

Phasing accuracy is the proportion of variants that are assigned to the correct haplotype. Accuracy is influenced by the density of heterozygous variants, the length of the phasing signal, and the error rate of the underlying sequencing data.

For small genomes such as bacteria and fungi, all three technologies can achieve high phasing accuracy because the entire genome can be covered by a small number of phasing signals. Linked-reads are particularly effective for small genomes because the high-molecular-weight DNA molecules can span a substantial fraction of the genome. Hi-C also works well for small genomes because the interaction frequency signal is strong when the genome is compact.

For medium genomes such as insects and small plants, the choice of phasing technology depends on heterozygosity. Species with high heterozygosity provide more variant density for Hi-C and linked-read phasing, which improves accuracy. Species with low heterozygosity may benefit from Strand-seq because the direct observation of template strand segregation does not depend on variant density [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

For large genomes such as mammals and large plants, the phasing signal must span hundreds of megabases to produce chromosome-scale haplotypes. Hi-C provides this long-range signal through chromatin interaction frequencies, but the accuracy is limited by the statistical nature of the signal. Strand-seq provides a direct global phase signal that spans entire chromosomes, and Graphasing has been shown to produce over 18 chromosome-spanning haplotypes in human assemblies [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. Linked-reads are limited by the length of the input DNA molecules, which typically span only tens of kilobases to a few hundred kilobases, making chromosome-scale phasing difficult without additional scaffolding information.

For polyploid genomes, the complexity of phasing increases because more than two homologous copies must be distinguished [<a href="#ref-3">3</a>]. Hi-C and linked-reads have difficulty resolving multiple haplotypes because their phasing signals are designed for diploid comparisons. Strand-seq can potentially distinguish multiple homologs because the template strand segregation pattern is unique to each homolog, but the computational analysis becomes more complex. The review of polyploid assembly strategies notes that haplotyping by phasing and third-generation technologies are having an impact on polyploid plant genome assembly [<a href="#ref-3">3</a>].

## Heterozygosity Requirements and Limitations

Heterozygosity is the proportion of genomic positions that differ between the two homologous chromosomes. This parameter directly affects the feasibility and accuracy of phasing for Hi-C and linked-reads, which rely on heterozygous variants as markers for haplotype assignment.

Hi-C phasing requires a sufficient density of heterozygous variants across the genome. The interaction frequency signal is computed at variant positions, and regions with low variant density provide little information for haplotype assignment. For species with low heterozygosity, such as some inbred laboratory strains and domesticated crops, Hi-C phasing may produce fragmented haplotypes with switch errors at regions where variants are sparse.

Linked-read phasing has the same variant density requirement as Hi-C, but the phasing signal is local instead of global. Variants must be present on the same high-molecular-weight DNA molecule to be phased together. In regions with low variant density, the barcode co-occurrence signal is weak, and the phasing accuracy decreases.

Strand-seq does not depend on variant density for the initial phase signal because the template strand segregation is observed directly at the DNA level. However, the assignment of variants to haplotypes still requires heterozygous variants as markers. The difference is that Strand-seq provides a global framework into which variants can be placed, so even sparse variants can be assigned to the correct haplotype if the framework is accurate [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

For species with very low heterozygosity, such as highly inbred lines, none of the three technologies can produce meaningful haplotypes because there are too few variants to distinguish the homologs. In these cases, the only option is to introduce genetic variation through crosses or to accept a collapsed assembly that does not distinguish haplotypes.

## Practical Workflow for Hi-C Phasing

The Hi-C phasing workflow begins with cell collection and crosslinking. Cells are treated with formaldehyde to crosslink proteins to DNA and to crosslink proteins to each other, preserving the three-dimensional structure of the nucleus. The crosslinked chromatin is then digested with a restriction enzyme, and the resulting fragments are ligated to each other. The ligated junctions are purified using biotin labeling and streptavidin pull-down, and the resulting library is sequenced with paired-end reads.

The bioinformatics workflow for Hi-C phasing includes read alignment to a reference genome or to the assembly graph, filtering of spurious ligation products, and construction of an interaction matrix. The interaction matrix is then used for two purposes: scaffolding contigs into chromosome-scale scaffolds and phasing variants across those scaffolds. Several software tools are available for these steps, and the choice of tools depends on the assembly workflow being used.

Quality control for Hi-C data includes checking the fraction of reads that map to valid ligation junctions, the distribution of interaction distances, and the consistency of the interaction matrix across biological replicates. Low mapping rates or abnormal interaction distance distributions indicate problems with library preparation or sequencing. The Galaxy Training Network provides accessible workflow training for sequence analysis that can be adapted for Hi-C data processing [<a href="#ref-5">5</a>].

The main limitation of the Hi-C workflow is that the phasing accuracy depends on the depth of sequencing and the quality of the interaction matrix. Low sequencing depth produces noisy interaction frequencies, which increases switch errors. High sequencing depth improves accuracy but increases cost. The optimal depth depends on genome size and heterozygosity, and it should be determined empirically for each project.

## Practical Workflow for Strand-seq Phasing

The Strand-seq workflow begins with the preparation of a single-cell suspension from the organism of interest. Cells are sorted into individual wells of a plate, and each cell undergoes whole-genome amplification. The amplified DNA from each cell is then used to construct a sequencing library with a unique barcode. The libraries are pooled and sequenced at low coverage per cell.

The bioinformatics workflow for Strand-seq phasing includes read alignment, identification of template strand direction for each read, and classification of each cell as having inherited the maternal or paternal homolog at each genomic region. The single-cell classifications are then combined across cells to build a global phase. The Graphasing workflow integrates this global phase signal with assembly graph topology to produce chromosome-scale haplotypes [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

Quality control for Strand-seq data includes checking the number of cells that pass quality filters, the coverage per cell, and the consistency of template strand assignments across cells. Cells with low coverage or ambiguous template strand assignments should be excluded from the analysis. The number of cells required depends on genome size and the desired phasing accuracy, and it should be determined empirically for each project.

The main limitation of the Strand-seq workflow is the laboratory complexity of single-cell processing. The cell sorting, whole-genome amplification, and library construction steps require specialized equipment and expertise. The cost per cell is higher than bulk sequencing approaches, and the total cost scales with the number of cells required. For projects with limited budgets, the higher accuracy of Strand-seq phasing must be weighed against the additional cost [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

## Practical Workflow for Linked-Read Phasing

The linked-read workflow begins with the extraction of high-molecular-weight DNA. The quality of this DNA is the most important factor in the success of linked-read phasing, because the phasing distance is directly limited by the length of the input DNA molecules. DNA extraction methods that minimize shearing, such as agarose plug preparation or gentle lysis protocols, are recommended.

The high-molecular-weight DNA is then loaded into the microfluidic partitioning instrument, which partitions the DNA into thousands of nanoliter-scale reactions. Each partition receives a unique barcode, and the DNA within each partition is fragmented and tagged with that barcode. The barcoded fragments are then pooled and sequenced with standard short-read sequencing.

The bioinformatics workflow for linked-read phasing includes read alignment, barcode deconvolution, and phasing analysis. The phasing analysis identifies variants that share barcodes and uses the co-occurrence pattern to assign variants to haplotypes. The quality of the phasing depends on the length of the input DNA molecules and the depth of sequencing.

Quality control for linked-read data includes checking the distribution of molecule lengths, the number of barcodes per molecule, and the coverage per barcode. Short molecules or low coverage per barcode reduce the phasing signal and increase switch errors. The main limitation of the linked-read workflow is that the phasing distance is limited to the length of the input DNA molecules, which is typically tens to hundreds of kilobases. This limitation makes chromosome-scale phasing difficult without additional scaffolding information.

## Integration with Assembly Workflows

The choice of phasing technology must be integrated with the overall assembly workflow. Long-read assembly produces contigs and assembly graphs that can be used as the foundation for phasing [<a href="#ref-4">4</a>]. The phasing technology then adds long-range information that connects variants across contigs and scaffolds.

Hi-C phasing integrates naturally with long-read assembly because the same Hi-C data used for scaffolding can be used for phasing. The assembly graph provides the local structure, and the Hi-C interaction matrix provides the long-range connectivity. This integration is straightforward and well supported by existing bioinformatics tools.

Strand-seq phasing integrates with assembly workflows through the Graphasing workflow, which synthesizes the global phase signal of Strand-seq with assembly graph topology [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. Graphasing readily integrates with any assembly workflow that outputs an assembly graph and has a haplotype assembly mode [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. This flexibility allows researchers to use their preferred long-read assembler and then add Strand-seq phasing as a separate step.

Linked-read phasing integrates with assembly workflows by providing local phasing information that can be used to resolve haplotypes within contigs. The linked-read data does not provide chromosome-scale connectivity, so it must be combined with other scaffolding information to produce chromosome-scale haplotypes.

The nf-core documentation provides standards for reproducible bioinformatics pipelines that can be applied to phasing workflows [<a href="#ref-6">6</a>]. The Bioconductor project provides packages for genomic analysis that can be used for phasing data processing and visualization [<a href="#ref-7">7</a>]. These resources support the implementation of reproducible phasing workflows.

## Records and Measurements for Phasing Projects

Maintaining detailed records of phasing project parameters is essential for quality control and reproducibility. The following measurements should be recorded for each phasing experiment.

For Hi-C phasing, record the number of cells used for crosslinking, the restriction enzyme used for digestion, the number of sequencing reads generated, the mapping rate, the fraction of reads that map to valid ligation junctions, and the interaction distance distribution. These measurements provide a basis for evaluating library quality and for comparing results across experiments.

For Strand-seq phasing, record the number of cells sorted, the number of cells that pass quality filters, the coverage per cell, the template strand classification rate, and the number of cells that contribute to the final phase. These measurements provide a basis for evaluating the quality of the single-cell data and for determining whether additional cells are needed.

For linked-read phasing, record the input DNA molecule length distribution, the number of barcodes generated, the coverage per barcode, and the phasing distance achieved. These measurements provide a basis for evaluating the quality of the high-molecular-weight DNA preparation and for determining whether the phasing distance is sufficient for the project goals.

The NCBI provides data resources for storing and accessing genomic data, including raw sequencing reads and assembled genomes [<a href="#ref-8">8</a>]. Depositing phasing data in public databases supports reproducibility and enables other researchers to reuse the data. The EMBL-EBI Training provides learning pathways for bioinformatics data analysis that can help researchers develop the skills needed for phasing data processing [<a href="#ref-9">9</a>].

## Common Failure Patterns in Phasing Experiments

Several failure patterns recur across phasing experiments, and recognizing these patterns early can save time and money.

Low mapping rates in Hi-C data often indicate problems with library preparation, such as incomplete crosslinking or excessive ligation of non-adjacent fragments. The fraction of reads that map to valid ligation junctions should be monitored, and libraries with low valid junction fractions should be repeated.

Insufficient cell numbers in Strand-seq experiments produce fragmented phases with switch errors. The number of cells required depends on genome size, and projects with large genomes may require more cells than initially planned. The quality of the phase should be evaluated after an initial batch of cells, and additional cells should be processed if the phase is fragmented.

Short DNA molecules in linked-read experiments produce short phasing distances that are insufficient for chromosome-scale phasing. The DNA extraction protocol should be optimized to maximize molecule length, and the molecule length distribution should be measured before proceeding with library preparation.

Low heterozygosity in the target genome reduces the variant density available for phasing. This limitation affects Hi-C and linked-reads more than Strand-seq, and it should be assessed before choosing a phasing technology. If heterozygosity is very low, none of the three technologies will produce meaningful haplotypes.

## Limitations and Interpretation Boundaries

The three phasing technologies have limitations that affect the interpretation of results. Understanding these limitations is essential for avoiding overinterpretation of phasing data.

Hi-C phasing produces statistical haplotypes that are accurate at the population level but may contain switch errors at individual variant positions. The accuracy of the phase should be evaluated using independent validation data, such as parental genotypes or pedigree information, when available.

Strand-seq phasing produces direct observations of template strand segregation, but the interpretation depends on the assumption that the template strand pattern is consistent across cells. Cells with aneuploidy or other chromosomal abnormalities may produce inconsistent patterns that complicate the analysis.

Linked-read phasing produces local haplotypes that are accurate within the phasing distance but may not be connected across larger genomic distances. The phase should be interpreted as local unless additional scaffolding information provides chromosome-scale connectivity.

The Graphasing workflow performs comparably to trio-phasing in contiguity, phasing accuracy, and assembly quality, and it outperforms Hi-C in phasing accuracy [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. However, these results were demonstrated in human assemblies, and the performance may differ for other species with different genome characteristics [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

## Safety and Regulatory Context for Phasing Data

Phasing data from human samples contains sensitive genetic information that is subject to privacy regulations. Researchers working with human phasing data must ensure that their data handling procedures comply with applicable regulations, including informed consent requirements and data protection standards. The NCBI provides guidance on data submission and access for human genomic data [<a href="#ref-8">8</a>].

For non-human species, phasing data may be subject to export controls or material transfer agreements, depending on the species and the jurisdiction. Researchers should verify that their data handling procedures comply with all applicable regulations before beginning a phasing project.

The computational analysis of phasing data requires access to high-performance computing resources. Researchers should ensure that their computing environment meets the security requirements for the data being processed, particularly for human data.

## Professional Escalation Criteria for Phasing Problems

Researchers should escalate phasing problems to a specialist when they encounter issues that cannot be resolved with standard troubleshooting. The following criteria indicate when professional escalation is appropriate.

If the Hi-C interaction matrix shows abnormal patterns that persist across multiple library preparations, the problem may be in the cell collection or crosslinking steps. A specialist in chromatin biology should be consulted to evaluate the protocol.

If the Strand-seq template strand classification rate is consistently low across multiple batches of cells, the problem may be in the single-cell sorting or whole-genome amplification steps. A specialist in single-cell genomics should be consulted to evaluate the protocol.

If the linked-read phasing distance is consistently shorter than expected, the problem may be in the high-molecular-weight DNA extraction. A specialist in DNA extraction should be consulted to evaluate the protocol.

If the phasing accuracy is lower than expected for the genome size and heterozygosity level, the problem may be in the bioinformatics analysis. A specialist in genome assembly and phasing should be consulted to evaluate the analysis workflow.

## Decision Framework for Matching Phasing Technology to Project Constraints

Selecting among Hi-C, Strand-seq, and linked-reads requires a structured evaluation that goes beyond the technical characteristics of each method. A practical decision framework should incorporate project-specific constraints including genome size, heterozygosity, available instrumentation, budget limits, and the intended downstream use of the phased assembly. The framework below provides a stepwise approach for matching phasing technology to project constraints, with explicit criteria for when each technology is the most appropriate choice.

### Step 1: Assess Genome Size and Complexity Class

The first decision point is the genome size and ploidy level of the target organism. Genome size determines the physical distance that the phasing signal must span to produce chromosome-scale haplotypes. For genomes smaller than 500 megabases, linked-reads can provide sufficient phasing distance because high-molecular-weight DNA molecules can cover a meaningful fraction of the genome. For genomes between 500 megabases and 3 gigabases, both Hi-C and Strand-seq are viable options, and the choice depends on heterozygosity and accuracy requirements. For genomes larger than 3 gigabases, Strand-seq provides the most reliable chromosome-scale phasing because the direct template strand signal does not degrade with genome size [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

Ploidy level is the second consideration within this step. Diploid genomes are compatible with all three technologies, but polyploid genomes require careful evaluation. The review of polyploid plant genome assembly strategies notes that haplotyping by phasing and third-generation technologies are having an impact on polyploid plant genome assembly [<a href="#ref-3">3</a>]. For polyploid genomes, Strand-seq has a potential advantage because the template strand segregation pattern is unique to each homolog, but the computational analysis becomes more complex. Hi-C and linked-reads have difficulty resolving multiple haplotypes because their phasing signals are designed for diploid comparisons [<a href="#ref-3">3</a>].

### Step 2: Measure Heterozygosity Before Committing

Heterozygosity should be estimated from preliminary sequencing data before selecting a phasing technology. A small amount of short-read or long-read sequencing, typically 15 to 30 fold coverage, can provide an accurate estimate of the heterozygous variant density. This estimate directly determines whether Hi-C or linked-reads will have sufficient variant markers for reliable phasing.

For species with heterozygosity above 0.5 percent, Hi-C and linked-reads will generally have adequate variant density for phasing. For species with heterozygosity between 0.1 and 0.5 percent, Hi-C phasing may produce fragmented haplotypes with switch errors in regions where variants are sparse. For species with heterozygosity below 0.1 percent, Strand-seq is the preferred choice because the direct observation of template strand segregation does not depend on variant density for the initial phase signal [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. The assignment of variants to haplotypes still requires heterozygous markers, but Strand-seq provides a global framework into which even sparse variants can be placed [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

### Step 3: Evaluate Available Instrumentation and Expertise

The decision framework must account for the laboratory infrastructure available to the research team. Hi-C requires only standard molecular biology equipment including a thermal cycler, centrifuge, and standard sequencing access. Strand-seq requires a cell sorter or microfluidic single-cell processing instrument, whole-genome amplification reagents, and expertise in single-cell genomics. Linked-reads require the specific microfluidic partitioning instrument and high-molecular-weight DNA extraction expertise.

Teams without access to specialized instrumentation should factor in the cost and time of sending samples to a core facility or commercial provider. The choice between sending samples externally and investing in in-house capacity depends on the number of projects planned and the long-term value of the instrumentation. The Carpentries lessons provide foundational training in computing and data skills that can help research teams build the bioinformatics capacity needed for phasing data processing [<a href="#ref-10">10</a>].

### Step 4: Determine the Primary Output Requirement

The intended downstream use of the phased assembly determines which technology provides the best value. If the project requires both chromosome-scale scaffolding and phasing from a single data type, Hi-C is the most efficient choice because the same library provides both functions. If the project requires the highest possible phasing accuracy and parental samples are unavailable, Strand-seq with Graphasing is the preferred choice because it performs comparably to trio-phasing in contiguity, phasing accuracy, and assembly quality [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

If the project requires local phasing information for variant resolution within genes or small genomic regions, linked-reads provide sufficient accuracy at lower cost. The phasing distance limitation of linked-reads is acceptable when the biological question does not require chromosome-scale haplotypes.

### Step 5: Calculate Total Cost Including Failure Risk

The total cost calculation should include also the direct cost of library preparation and sequencing but also the expected cost of failed experiments and repeated library preparations. Hi-C libraries have a moderate failure rate due to the multiple enzymatic steps and the sensitivity of the crosslinking reaction to cell quality. Strand-seq has a higher per-sample cost but provides direct quality metrics that allow early identification of failed cells. Linked-reads have a lower failure rate but are sensitive to DNA quality, and poor high-molecular-weight DNA extraction can produce libraries with insufficient phasing distance.

The cost of bioinformatics analysis should also be included in the total calculation. Hi-C data processing requires substantial memory for interaction matrix construction in large genomes. Strand-seq data processing requires alignment and classification of single-cell data followed by the Graphasing workflow [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. Linked-read data processing is generally less memory-intensive than Hi-C analysis.

### Step 6: Apply the Decision Matrix

The following decision matrix summarizes the recommended technology choices based on the combination of genome size and heterozygosity. These recommendations reflect the performance characteristics described in the comparative literature and should be validated with pilot data for each new species.

| Genome Size | High Heterozygosity | Low Heterozygosity |
|-------------|---------------------|---------------------|
| Small (under 500 Mb) | Linked-reads or Hi-C | Strand-seq |
| Medium (500 Mb to 3 Gb) | Hi-C | Strand-seq |
| Large (over 3 Gb) | Strand-seq | Strand-seq |

For small genomes with high heterozygosity, linked-reads provide the most cost-effective solution because the phasing distance covers a substantial fraction of the genome. Hi-C is also effective but adds the complexity of chromatin preparation. For small genomes with low heterozygosity, Strand-seq is preferred because the direct template strand signal does not depend on variant density [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

For medium genomes with high heterozygosity, Hi-C provides the best balance of cost and utility because the same data provides scaffolding and phasing. For medium genomes with low heterozygosity, Strand-seq is preferred because Hi-C will struggle with sparse variant density.

For large genomes, Strand-seq is the preferred choice regardless of heterozygosity because the direct global phase signal spans entire chromosomes [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. Hi-C can provide chromosome-scale scaffolding for large genomes, but the phasing accuracy is limited by the statistical nature of the interaction frequency signal. Linked-reads are not recommended for large genomes because the phasing distance is limited to the length of the input DNA molecules.

### Pilot Testing and Validation Protocol

Before committing to a full-scale phasing project, a pilot test should be conducted with a small number of samples to validate the chosen technology. The pilot should include library preparation, sequencing, and bioinformatics analysis to confirm that the phasing signal is sufficient for the target genome.

For Hi-C pilot testing, prepare one library and sequence to the planned depth. Evaluate the fraction of reads that map to valid ligation junctions and the interaction distance distribution. If the valid junction fraction is below 50 percent, the library preparation protocol should be optimized before proceeding.

For Strand-seq pilot testing, process an initial batch of 50 to 100 cells and evaluate the template strand classification rate. If the classification rate is below 80 percent, the single-cell sorting or whole-genome amplification steps should be optimized. The quality of the phase should be evaluated after this initial batch, and additional cells should be processed if the phase is fragmented.

For linked-read pilot testing, measure the input DNA molecule length distribution before library preparation. If the median molecule length is below 50 kilobases, the DNA extraction protocol should be optimized to reduce shearing. After sequencing, evaluate the phasing distance achieved and confirm that it is sufficient for the project goals.

The Galaxy Training Network provides accessible workflow training for sequence analysis that can be adapted for phasing data processing [<a href="#ref-5">5</a>]. The nf-core documentation provides standards for reproducible bioinformatics pipelines that can be applied to phasing workflows [<a href="#ref-6">6</a>]. The Bioconductor project provides packages for genomic analysis that can be used for phasing data processing and visualization [<a href="#ref-7">7</a>]. These resources support the implementation of reproducible phasing workflows and can help research teams build the computational skills needed for successful phasing projects.

The NCBI provides data resources for storing and accessing genomic data, including raw sequencing reads and assembled genomes [<a href="#ref-8">8</a>]. Depositing pilot data in public databases supports reproducibility and enables other researchers to evaluate the performance of the chosen phasing technology for similar genomes. The EMBL-EBI Training provides learning pathways for bioinformatics data analysis that can help researchers develop the skills needed for phasing data processing [<a href="#ref-9">9</a>].

## Frequently Asked Questions

### What is the difference between phasing and scaffolding?

Phasing assigns genetic variants to their respective homologous chromosomes, determining which alleles are inherited together. Scaffolding orders and orients contigs into chromosome-scale structures. Hi-C provides both scaffolding and phasing information from a single library, while Strand-seq provides only phasing information and linked-reads provide local phasing information that can assist with scaffolding.

### How many cells are needed for Strand-seq phasing of a mammalian genome?

The number of cells required depends on genome size and the desired phasing accuracy. Larger genomes require more cells to achieve the same coverage of template strand information across all chromosomes. The quality of the phase should be evaluated after an initial batch of cells, and additional cells should be processed if the phase is fragmented.

### Can Hi-C phasing be used for polyploid genomes?

Hi-C phasing is designed for diploid comparisons and has difficulty resolving multiple haplotypes in polyploid genomes. The interaction frequency signal becomes more complex when more than two homologous copies must be distinguished. Polyploid phasing remains an active area of research, and the review of polyploid assembly strategies notes the impact of haplotyping by phasing and third-generation technologies [<a href="#ref-3">3</a>].

### What is the minimum heterozygosity required for linked-read phasing?

There is no fixed minimum heterozygosity threshold, but the phasing accuracy decreases as variant density decreases. In regions with low variant density, the barcode co-occurrence signal is weak, and the phasing accuracy decreases. For species with very low heterozygosity, none of the three technologies will produce meaningful haplotypes because there are too few variants to distinguish the homologs.

### How does Graphasing compare to trio-phasing?

Graphasing performs comparably to trio-phasing in contiguity, phasing accuracy, and assembly quality [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. The advantage of Graphasing is that it does not require parental data, which is often difficult to acquire for non-model organisms. Graphasing integrates the global phase signal of Strand-seq with assembly graph topology to produce chromosome-scale de novo haplotypes [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

### What is the cost difference between Hi-C and Strand-seq phasing?

Strand-seq is generally more expensive than Hi-C because of the single-cell processing requirement. The cost per cell is higher than bulk sequencing approaches, and the total cost scales with the number of cells required. However, the higher phasing accuracy of Strand-seq may justify the additional cost for projects where haplotype resolution is the primary scientific objective [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].

### Can linked-reads be used for chromosome-scale phasing?

Linked-reads are limited by the length of the input DNA molecules, which typically span only tens to hundreds of kilobases. This limitation makes chromosome-scale phasing difficult without additional scaffolding information. For chromosome-scale phasing, Hi-C or Strand-seq are more appropriate choices.

### What assembly graph formats are compatible with Graphasing?

Graphasing integrates with any assembly workflow that outputs an assembly graph and has a haplotype assembly mode [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>]. The Canu assembler provides graph-based assembly outputs in graphical fragment assembly format for analysis or integration with complementary phasing and scaffolding techniques [<a href="#ref-4">4</a>]. Researchers should verify that their chosen assembler produces the required graph format before planning a Strand-seq phasing experiment.

## Related Bioinformatics Guides

- [Long-Read Sequencing for De Novo Assembly of Complex Genomes: Case Studies and Best Practices](/knowledge/bioinformatics/long-read-sequencing-for-de-novo-assembly-of-complex-genomes-case-studies-and-best-practices)
- [Metagenomics Assembly: Strategies for Reconstructing Microbial Genomes](/knowledge/bioinformatics/metagenomics-assembly-strategies-for-reconstructing-microbial-genomes)
- [Long-Read Genome Assembly and Polishing Strategies](/knowledge/bioinformatics/long-read-genome-assembly-and-polishing-strategies)
- [Metagenomic Assembly and Binning: A Practical Workflow for Recovering Genomes from Complex Microbial Communities](/knowledge/bioinformatics/metagenomic-assembly-and-binning-a-practical-workflow-for-recovering-genomes-from-complex-microb)
- [RNA-Seq vs qPCR: Validation and Comparison](/knowledge/bioinformatics/rna-seq-vs-qpcr-validation-and-comparison)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)

## References and Further Reading

<a id="ref-1"></a>[<a href="#ref-1">1</a>] [Phasing Diploid Genome Assembly Graphs with Single-Cell Strand Sequencing.](https://pubmed.ncbi.nlm.nih.gov/38529499). bioRxiv : the preprint server for biology, 2024.

<a id="ref-2"></a>[<a href="#ref-2">2</a>] [Graphasing: phasing diploid genome assembly graphs with single-cell strand sequencing.](https://pubmed.ncbi.nlm.nih.gov/39390579). Genome biology, 2024.

<a id="ref-3"></a>[<a href="#ref-3">3</a>] [Current Strategies of Polyploid Plant Genome Sequence Assembly.](https://pubmed.ncbi.nlm.nih.gov/30519250). Frontiers in plant science, 2018.

<a id="ref-4"></a>[<a href="#ref-4">4</a>] [Canu: scalable and accurate long-read assembly via adaptive k-mer weighting and repeat separation.](https://pubmed.ncbi.nlm.nih.gov/28298431). Genome research, 2017.

<a id="ref-5"></a>[<a href="#ref-5">5</a>] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.

<a id="ref-6"></a>[<a href="#ref-6">6</a>] [nf-core Documentation](https://nf-co.re/docs). nf-core.

<a id="ref-7"></a>[<a href="#ref-7">7</a>] [Bioconductor](https://bioconductor.org/). Bioconductor Project.

<a id="ref-8"></a>[<a href="#ref-8">8</a>] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.

<a id="ref-9"></a>[<a href="#ref-9">9</a>] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.

<a id="ref-10"></a>[<a href="#ref-10">10</a>] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.