# FASTA, FASTQ, and GFA: A Field Guide to File Formats in Genome Assembly Projects


## Key Takeaways

- **FASTA** is the standard for linear sequence representation (reference genomes, contigs) and is universally accepted by alignment and annotation tools, but it permanently discards per-base quality scores and cannot encode structural relationships.
- **FASTQ** is essential for raw sequencing data, preserving per-base Phred quality scores critical for read trimming, error correction, variant calling, and assembly algorithms that use confidence metrics.
- **GFA** is specifically designed for assembly graphs, enabling representation of branching structures, alternative haplotypes, and structural variants that linear FASTA formats collapse, making it vital for detailed structural analysis.
- Converting **FASTQ to FASTA** results in irreversible loss of quality information, while converting **GFA to FASTA** requires explicit path selection, potentially obscuring structural complexity and alternative genomic arrangements.
- Understanding **quality score encoding** (Phred+33 vs. Phred+64) is paramount when processing FASTQ files, as misinterpretation leads to systematically incorrect confidence values impacting all downstream analyses.
- **Data provenance** (sequencing platform, assembler version, parameters) is critical for correctly interpreting file formats, especially quality score encoding and header conventions, ensuring reproducibility and accurate downstream analysis.

---

Genome assembly projects depend on three primary file formats that serve distinct roles: FASTA for reference sequences and assembled contigs, FASTQ for raw sequencing reads with quality information, and GFA for graph-based representations of assembly structures. Each format encodes different information, and using the wrong format for a given task leads to data loss, misinterpretation, or corrupted downstream analyses. This field guide explains the structure, intended use cases, and practical conversion considerations for each format, with attention to quality score handling and graph representation.

## The Role of File Formats in Assembly Workflows

A genome assembly workflow moves data through several stages, each requiring a specific file format. Raw sequencing instruments produce FASTQ files that contain both nucleotide sequences and per-base quality scores. Assembly algorithms consume these FASTQ files and produce FASTA files containing contiguous consensus sequences. When assembly graphs are involved, GFA files capture the branching structure that linear FASTA representations cannot express.

Understanding these formats matters because each one carries different information. FASTA files store sequence data only. FASTQ files store sequence data plus quality information. GFA files store sequence data plus structural relationships between sequence segments. Converting between formats without accounting for these differences causes silent data loss. For example, converting FASTQ to FASTA discards quality scores permanently, and converting GFA to FASTA collapses graph branches into linear paths, losing alternative structural arrangements.

Researchers working with public genomic data encounter these formats constantly. The [National Center for Biotechnology Information](https://www.ncbi.nlm.nih.gov/) maintains sequence databases that accept and distribute data in FASTA and FASTQ formats, and the scale of available data continues to grow as sequencing technologies improve. Understanding file formats is a prerequisite for using these resources effectively.

## FASTA Format: Sequence-Only Representation

FASTA is the simplest and most widely used sequence format in bioinformatics. It represents biological sequences using single-letter nucleotide or amino acid codes preceded by a header line that begins with the greater-than symbol.

### Structure and Syntax

A FASTA file consists of alternating header lines and sequence lines. The header line starts with the greater-than character followed by an identifier and optional descriptive text. Sequence lines contain the actual nucleotide or amino acid characters, typically wrapped at a fixed width for readability.

A minimal FASTA record looks like this:

```
>contig_1 length=1500
ACGTACGTACGTACGTACGTACGTACGTACGTACGTACGT
ACGTACGTACGTACGTACGTACGTACGTACGTACGTACGT
```

The header identifier is the first word after the greater-than symbol. Everything after the first whitespace is descriptive and varies by source. Different databases use different header conventions, so parsing scripts must handle this variability.

### Common Use Cases

FASTA files serve as the standard format for reference genomes, assembled contigs, protein sequences, and primer sequences. Most alignment tools, annotation pipelines, and comparative genomics applications accept FASTA input. The format is also the standard output for genome assemblers that produce linear contigs.

The [SyntenyPair Explorer](https://doi.org/10.64898/2026.07.23.740353) tool accepts FASTA genome assemblies as input for computing assembly summary statistics and for visualizing structural comparisons between related genomes. This demonstrates the continued utility of FASTA as an input format for downstream analysis tools even as assembly graphs become more common.

### Limitations

FASTA files cannot store quality information, read identifiers beyond the header, or structural relationships between sequences. When a FASTA file is created from FASTQ data, the quality scores are lost. When a FASTA file is created from a GFA graph, the alternative paths and branching structure are collapsed into a single linear representation.

These limitations are acceptable when the goal is to represent a finished consensus sequence. They become problematic when researchers need to evaluate assembly quality, identify sequencing errors, or understand structural variation that manifests as alternative graph paths.

## FASTQ Format: Sequence Plus Quality

FASTQ extends FASTA by adding a quality score for every nucleotide position. This format is the standard output of high-throughput sequencing instruments and the required input for most read alignment and assembly tools.

### Structure and Syntax

A FASTQ record contains four lines per read. The first line begins with the at sign and contains the read identifier and optional description. The second line contains the nucleotide sequence. The third line begins with the plus sign and may repeat the identifier or contain only the plus sign. The fourth line contains quality characters, one for each nucleotide in the sequence.

A minimal FASTQ record looks like this:

```
@read_001 length=150
ACGTACGTACGTACGTACGTACGTACGTACGTACGTACGT
+
IIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIIII
```

The quality line uses ASCII characters to encode Phred quality scores. The mapping between ASCII characters and numeric quality scores depends on the encoding scheme used by the sequencing platform. Older Illumina platforms used a Phred+64 encoding, while current platforms and most other technologies use Phred+33. Misinterpreting the encoding scheme leads to incorrect quality scores and can affect downstream variant calling and error correction.

### Quality Score Handling

Phred quality scores represent the probability that a base call is incorrect. A Phred score of Q20 corresponds to a 1 in 100 error probability, and Q30 corresponds to a 1 in 1000 error probability. Higher scores indicate higher confidence.

Quality scores serve several purposes in assembly workflows. Read alignment tools use them to weight mismatches and gaps. Error correction tools use them to identify and correct sequencing errors. Assembly tools use them to resolve conflicts between reads. Trimming tools use them to remove low-quality bases from read ends.

The [European Bioinformatics Institute](https://www.ebi.ac.uk/training) provides training materials on working with sequencing data, including quality assessment and quality control procedures. These resources emphasize that quality score inspection should be a routine step in any assembly project, not an afterthought.

### Common Use Cases

FASTQ files are the primary input for de novo assembly, read mapping, variant calling, and metagenomic analysis. The format preserves the information needed to distinguish high-confidence base calls from low-confidence ones, which is essential for accurate assembly.

Metagenomic analysis pipelines such as [MGS2AMR](https://doi.org/10.1186/s40168-023-01674-z) process FASTQ data from clinical specimens to detect antibiotic resistance genes and identify their organism of origin. These pipelines rely on quality information during assembly and read mapping to achieve accurate results.

### Limitations

FASTQ files are larger than FASTA files because they store quality characters for every base. A FASTQ file is typically three to four times larger than the corresponding FASTA file. Storage and transfer costs must be considered when planning sequencing projects.

FASTQ files do not store alignment information, variant calls, or assembly graphs. They represent reads as independent entities without relationships to each other or to a reference. The format is an input format for analysis, not an output format for results.

## GFA Format: Graph Representation of Assemblies

GFA is the Graphical Fragment Assembly format, designed to represent assembly graphs. Unlike FASTA and FASTQ, which represent linear sequences, GFA can represent branching structures, alternative paths, and relationships between sequence segments.

### Structure and Syntax

A GFA file contains segment records, link records, and optional path records. Segment records define sequence segments and begin with the letter S. Link records define connections between segments and begin with the letter L. Path records define traversals through the graph and begin with the letter P.

A minimal GFA file looks like this:

```
S	1	ACGTACGTACGTACGTACGT
S	2	TTTTAAAACCCCGGGG
L	1	+	2	+	0M
P	path_1	1+,2+	*
```

Segment records contain an identifier, the sequence, and optional tags. Link records contain the source segment, source orientation, target segment, target orientation, overlap information, and optional tags. Path records contain a path identifier, a list of segment traversals, and optional tags.

### Graph Representation and Assembly Structure

Assembly graphs capture the complexity that linear FASTA representations cannot express. Repeats, heterozygous regions, and structural variants create branching structures in the graph. Each branch represents an alternative arrangement of sequence segments.

Long-read assemblies of organelle genomes demonstrate this complexity. Research on the [Golden Wattle chloroplast and mitochondrial genomes](https://doi.org/10.46471/gigabyte.36) found that different assembly algorithms produced contrasting arrangements of genomic segments, with mapped reads spanning alternate paths. A linear FASTA representation would hide this structural diversity, while a GFA representation preserves it.

The [MGS2AMR pipeline](https://doi.org/10.1186/s40168-023-01674-z) uses GFA files as input for its gene-centric analysis of metagenomic data. The pipeline implements algorithms that optimize and annotate assembly paths within the raw GFA graph, improving sensitivity for detecting antibiotic resistance genes and aiding species annotation. This demonstrates that GFA files are also intermediate artifacts but can serve as direct input for downstream analysis tools.

### Common Use Cases

GFA files are produced by assembly tools that generate graphs, including many long-read assemblers. They are used for inspecting assembly structure, identifying misassemblies, resolving haplotypes, and understanding structural variation.

GFA files also serve as input for tools that analyze graph structure. The algorithms used in [MGS2AMR](https://doi.org/10.1186/s40168-023-01674-z) operate directly on GFA graphs to identify optimal paths through seed segments and to annotate the origin of detected genes. This approach achieves results comparable to whole-genome sequencing of isolates for antimicrobial resistance prediction.

### Limitations

GFA files are more complex than FASTA or FASTQ files and require specialized tools for visualization and analysis. Not all downstream tools accept GFA input, so conversion to FASTA is often necessary for compatibility.

Converting GFA to FASTA requires choosing a path through the graph. Different paths represent different structural arrangements, and the choice affects the resulting linear sequence. Researchers must understand what the conversion does and document the path selection criteria.

## At a Glance: Format Comparison Table

| Format | Primary Content | Typical Use | Quality Scores | Structural Information | Common Output Of | Common Input To |
|--------|----------------|-------------|----------------|----------------------|-----------------|-----------------|
| FASTA | Nucleotide or amino acid sequences | Reference genomes, assembled contigs, protein sequences | No | No | Genome assemblers, sequence databases | Alignment tools, annotation pipelines, comparative genomics tools |
| FASTQ | Nucleotide sequences with per-base quality scores | Raw sequencing reads, read mapping, variant calling | Yes | No | Sequencing instruments | Assembly tools, read aligners, quality trimming tools |
| GFA | Sequence segments with graph connections | Assembly graphs, structural variation analysis, haplotype resolution | No | Yes | Graph-based assemblers | Graph analysis tools, assembly inspection tools, specialized pipelines |

## Practical Workflow: Moving Between Formats

Most assembly projects require moving data between formats at specific stages. Understanding when and how to convert is essential for avoiding data loss and misinterpretation.

### Step 1: Assess Input Data Format

Before starting an assembly, verify that input data is in the expected format. Raw sequencing data should be in FASTQ format with quality scores. If data is in FASTA format, determine whether quality information is available elsewhere or has been permanently lost.

Check the quality score encoding. The first quality character in a read can indicate the encoding scheme. Phred+33 encoding produces quality characters in the ASCII range 33 to 73, while Phred+64 encoding produces characters in the range 64 to 104. Using the wrong encoding interpretation corrupts all downstream quality-based analyses.

### Step 2: Run Assembly and Capture Output Formats

Assembly tools produce different output formats depending on the tool and its configuration. Some tools produce only FASTA contigs. Others produce both FASTA contigs and GFA graphs. Configure the assembly to produce all available output formats when graph information may be needed later.

Record the assembly tool version, parameters, and output formats in the project log. This information is essential for reproducing the assembly and for interpreting the output files.

### Step 3: Inspect Assembly Graphs Before Converting

When a GFA file is available, inspect the graph structure before converting to FASTA. Identify the number of segments, the number of links, and the presence of alternative paths. This inspection reveals whether the assembly contains structural complexity that a linear representation would hide.

Tools for graph visualization and analysis are available through bioinformatics training resources. The [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible workflow training that includes assembly and graph analysis tutorials. These resources help researchers develop the skills needed to interpret assembly graphs correctly.

### Step 4: Convert Formats With Explicit Documentation

When converting GFA to FASTA, document the path selection method. Different assemblers use different conventions for choosing the primary path through the graph. The choice affects the resulting linear sequence and any downstream analyses.

When converting FASTQ to FASTA, recognize that quality information is permanently lost. Only perform this conversion when quality information is not needed for downstream analysis.

### Step 5: Validate Converted Files

After conversion, validate the output files. Check that sequence lengths are consistent with expectations, that headers contain the expected identifiers, and that no sequences were truncated or merged. Compare the number of records in the output to the number expected from the input.

Validation is particularly important when converting between formats with different record structures. A FASTA file created from a FASTQ file should contain the same number of records as the original, with the same identifiers and sequences.

## Options and Tradeoffs in Format Selection

The choice of file format depends on the analysis goal, the available tools, and the need to preserve specific types of information.

### FASTA for Finished Products

FASTA is appropriate for representing finished consensus sequences, reference genomes, and sequences intended for public databases. The format is universally accepted and simple to parse. The lack of quality and structural information is acceptable when the sequence represents a validated consensus.

Public sequence databases such as those maintained by the [National Center for Biotechnology Information](https://www.ncbi.nlm.nih.gov/) accept FASTA-formatted sequences for submission and distribution. Researchers depositing assembled genomes typically provide FASTA files as the primary sequence representation.

### FASTQ for Raw Data and Quality-Sensitive Analysis

FASTQ is required for any analysis that depends on quality information. Read trimming, error correction, variant calling, and assembly all benefit from or require quality scores. Raw sequencing data should always be stored in FASTQ format.

Storage costs for FASTQ files are significant, but the information they contain is irreplaceable. Once quality scores are discarded, they cannot be recovered. Researchers should retain FASTQ files for the duration of the project and consider long-term archival options.

### GFA for Structural Analysis

GFA is appropriate when assembly structure matters. Haplotype resolution, structural variant detection, and repeat analysis all benefit from graph representations. GFA files also enable the development of specialized analysis tools that operate on graph structure.

The [MGS2AMR pipeline](https://doi.org/10.1186/s40168-023-01674-z) demonstrates the value of GFA-based analysis. By operating directly on the assembly graph, the pipeline achieves sensitive detection of antibiotic resistance genes and correct assignment of their organism of origin. This approach would not be possible with linear FASTA sequences alone.

## Observations and Measurements in Assembly Projects

Systematic observation and measurement are essential for successful assembly projects. The following measurements should be recorded for every assembly.

### Read Quality Metrics

Before assembly, measure read quality using standard metrics. Per-base quality distributions, total read counts, read length distributions, and GC content provide a baseline for evaluating assembly performance. These metrics are computed from FASTQ files and should be recorded in the project log.

The [European Bioinformatics Institute](https://www.ebi.ac.uk/training) provides training on quality assessment for sequencing data. These resources emphasize that quality assessment should be performed before assembly, not after, because poor-quality input data produces poor-quality assemblies regardless of the assembly algorithm used.

### Assembly Statistics

After assembly, measure assembly statistics from the FASTA output. Total assembly size, number of contigs, N50, L50, and maximum contig length provide basic descriptors of assembly quality. These statistics are computed from FASTA files and are widely reported in assembly publications.

For graph-based assemblies, additional statistics describe the graph structure. Number of segments, number of links, number of alternative paths, and graph complexity metrics provide information that linear statistics cannot capture. These statistics are computed from GFA files.

### Coverage and Completeness

Coverage measurements describe how thoroughly the reads cover the assembled sequence. Average coverage, coverage distribution, and regions of zero coverage provide information about assembly completeness and potential misassembly.

Completeness assessment uses conserved gene sets to estimate what fraction of the expected genome content is present in the assembly. These assessments are typically performed on FASTA files using specialized tools.

## Records and Documentation Requirements

Maintaining accurate records is essential for reproducible assembly projects. The following records should be maintained for every project.

### Input Data Records

Record the source of each FASTQ file, including the sequencing platform, library preparation method, sequencing date, and any preprocessing steps applied. Record the total number of reads, total bases, and quality metrics for each file.

### Processing Records

Record every processing step applied to the data, including the tool name, version, parameters, and date of execution. Record the input and output file formats for each step. This information allows the assembly to be reproduced and audited.

### Output Records

Record the assembly statistics, graph statistics, and completeness assessment results. Record the file formats of all output files and their locations. Record any conversions performed and the parameters used for each conversion.

The [nf-core documentation](https://nf-co.re/docs) emphasizes the importance of reproducible workflow standards in bioinformatics. Community pipelines provide standardized processing steps and documentation practices that support reproducibility across projects and institutions.

## Quality Controls and Validation

Quality controls should be applied at multiple stages of the assembly workflow to detect problems early and prevent wasted effort.

### Input Quality Control

Validate FASTQ files before assembly. Check that all records have four lines, that sequence and quality lines have equal length, and that quality characters are within the expected ASCII range for the encoding scheme. Check for adapter contamination and low-quality regions.

### Assembly Quality Control

Validate assembly output after completion. Check that FASTA files contain valid sequence characters and that headers are properly formatted. Check that GFA files contain valid segment, link, and path records. Verify that the assembly size is consistent with expectations based on the organism and sequencing depth.

### Completeness and Correctness Assessment

Assess assembly completeness using conserved gene sets. Assess assembly correctness by mapping reads back to the assembly and checking for consistent coverage and alignment patterns. These assessments provide evidence that the assembly is suitable for downstream analysis.

The [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible tutorials on assembly quality assessment and validation. These resources help researchers apply appropriate quality controls without requiring extensive command-line experience.

## Common Failure Patterns in Format Handling

Several recurring problems affect researchers working with assembly file formats. Recognizing these patterns helps prevent data loss and misinterpretation.

### Quality Score Encoding Misinterpretation

Misinterpreting the quality score encoding is one of the most common errors in FASTQ handling. Using Phred+33 decoding on Phred+64 encoded data produces quality scores that are systematically too high, while the reverse produces scores that are too low. This error propagates through all downstream analyses that use quality information.

Prevention requires knowing the sequencing platform and version used to generate the data. When this information is unavailable, the quality character distribution can provide clues about the encoding scheme.

### Silent Quality Score Loss

Converting FASTQ to FASTA discards quality scores without warning. Researchers who perform this conversion for one analysis and later need quality information for another analysis find that the information is gone. This problem is preventable by retaining the original FASTQ files.

### GFA to FASTA Path Selection Bias

Converting GFA to FASTA requires selecting a path through the assembly graph. Different path selection methods produce different linear sequences, and the choice can affect downstream analyses. Researchers who do not document the path selection method produce results that cannot be reproduced or compared across studies.

### Header Parsing Errors

FASTA and FASTQ headers follow different conventions across databases and tools. Parsing scripts that assume a specific header structure fail when encountering headers from a different source. Robust parsing requires handling variable header formats.

### Truncated or Corrupted Files

File transfer interruptions and storage failures can truncate or corrupt sequence files. FASTA files with incomplete records, FASTQ files with mismatched sequence and quality lines, and GFA files with broken links all cause downstream analysis failures. File integrity checks should be performed after transfer and before analysis.

## Limitations of Format-Based Analysis

File formats impose constraints on what information can be represented and analyzed. Understanding these limitations is essential for interpreting results correctly.

### FASTA Cannot Represent Uncertainty

FASTA represents each position as a single nucleotide character. Ambiguous bases can be represented using IUPAC codes, but the format cannot represent alternative sequences, phased haplotypes, or structural variants. Researchers who need to represent these features must use graph-based formats.

### FASTQ Quality Scores Are Platform-Dependent

Quality scores are calibrated differently across sequencing platforms and even across runs on the same platform. Comparing quality scores across platforms requires calibration or normalization. Researchers should be cautious when using quality scores from different sources in the same analysis.

### GFA Does Not Include Quality Information

GFA files represent sequence segments and their connections but do not include quality scores for individual bases. Researchers who need quality information for graph-based analyses must obtain it from the original FASTQ files or from additional annotations.

### Assembly Graphs Are Simplified Representations

Assembly graphs simplify the true complexity of the underlying sequencing data. The graph structure depends on the assembly algorithm, its parameters, and the input data. Different assemblers produce different graphs from the same input data, as demonstrated by the [organelle genome assemblies of the Golden Wattle](https://doi.org/10.46471/gigabyte.36).

## Safety and Regulatory Context

Genome assembly projects involving human data are subject to additional considerations beyond file format handling. The consensus position in genomics since the Human Genome Project has been that data should be shared widely to achieve the greatest societal benefit. However, this position relies on imprecise definitions of broad data sharing, and implementation varies among landmark genomic studies.

Researchers working with human genomic data must understand the governance requirements for data deposition and sharing. The [framework for genomic data sharing](https://doi.org/10.1038/s41588-024-02049-2) includes considerations about the limits of general research use, the governance of public data deposition from extant samples, and the question of whether participants should be encouraged to share their participant status publicly.

These considerations affect file format choices because data sharing often requires converting data into specific formats for deposition in public databases. Researchers should understand the format requirements of the relevant data repositories before planning data deposition.

## Professional Escalation Criteria

Certain situations warrant escalation to more experienced colleagues or specialized support services. The following criteria indicate when professional escalation is appropriate.

### Unresolved Quality Score Encoding Issues

If quality score encoding cannot be determined from available documentation and the quality character distribution is ambiguous, escalate to a bioinformatics specialist. Incorrect quality score interpretation can invalidate all downstream analyses.

### Assembly Graph Complexity Beyond Tool Capabilities

If the assembly graph contains complexity that available tools cannot visualize or analyze, escalate to specialists with experience in graph-based assembly analysis. Attempting to analyze complex graphs with inadequate tools produces misleading results.

### Data Sharing Governance Questions

If questions arise about the permissibility of depositing or sharing genomic data, escalate to institutional data governance officers or ethics committees. Data sharing decisions have legal and ethical implications that require institutional guidance.

### Reproducibility Failures

If an assembly cannot be reproduced from the recorded parameters and input data, escalate to the tool developers or community support channels. Reproducibility failures indicate either incomplete records or tool behavior that requires expert interpretation.

## A Practical Decision Framework for Selecting Assembly File Formats

Choosing between FASTA, FASTQ, and GFA is not a one-time decision made at the start of a project. It is a recurring judgment call that researchers face at every stage of the assembly workflow, from raw data receipt through final data deposition. A structured decision framework helps standardize these choices, reduces the risk of silent data loss, and produces records that other researchers can audit. This section provides a decision framework organized around five questions, a record system for tracking format decisions, and troubleshooting methods for common format-related failures.

### The Five-Question Format Selection Framework

Before converting or creating any sequence file, work through five questions in order. Each question narrows the set of acceptable formats and produces a documented rationale that becomes part of the project record.

**Question 1: What information does the downstream tool require as input?**

Every analysis tool has explicit input format requirements. Read aligners and assemblers typically require FASTQ because they use quality scores during alignment and error correction. Annotation tools and comparative genomics applications typically accept FASTA. Graph analysis tools require GFA. Check the tool documentation before preparing input files. The [nf-core documentation](https://nf-co.re/docs) provides pipeline-specific input requirements that clarify which formats are accepted at each workflow stage.

If the downstream tool accepts multiple formats, choose the format that preserves the most information. For example, if an assembler accepts both FASTA and FASTQ input, provide FASTQ so quality information is available for error correction.

**Question 2: Does the analysis depend on per-base quality information?**

Quality scores matter for read trimming, error correction, variant calling, and any analysis where base call confidence affects the result. If the answer is yes, the file must be FASTQ or a format derived from FASTQ that retains quality information. Converting to FASTA at this stage would discard information that the analysis needs.

If the analysis does not depend on quality information, FASTA may be acceptable. Examples include computing GC content, searching for specific sequence motifs, or comparing assembled contigs against a reference database.

**Question 3: Does the analysis depend on structural relationships between sequences?**

Structural relationships matter for haplotype resolution, structural variant detection, repeat analysis, and understanding alternative arrangements of genomic segments. If the answer is yes, the file should be GFA or a format that preserves graph structure. Linear FASTA representations collapse alternative paths and hide structural complexity.

Research on organelle genomes demonstrates why this distinction matters. Assemblies of the [Golden Wattle chloroplast and mitochondrial genomes](https://doi.org/10.46471/gigabyte.36) produced contrasting arrangements of genomic segments depending on the assembly algorithm used. Mapped reads spanned alternate paths, providing evidence that the structural diversity was real and not an artifact. A linear FASTA representation would have hidden this information entirely.

**Question 4: Is this file an intermediate artifact or a final product?**

Intermediate files support ongoing analysis and should preserve maximum information. Final products represent validated results intended for sharing or deposition. FASTA is appropriate for final consensus sequences submitted to public databases. FASTQ is appropriate for raw data archival. GFA is appropriate for structural analysis results.

The distinction matters because conversions are often one-way. Converting FASTQ to FASTA discards quality scores permanently. Converting GFA to FASTA selects one path through the graph and discards alternative paths. Once these conversions are performed, the original information cannot be recovered unless the source files were retained.

**Question 5: What do the project records say about the source and provenance of this data?**

Data provenance determines which interpretations are valid. Quality score encoding depends on the sequencing platform and software version. Header conventions vary across databases and tools. Graph structure depends on the assembler and its parameters. Without provenance records, researchers cannot correctly interpret the files they are working with.

The [European Bioinformatics Institute](https://www.ebi.ac.uk/training) provides training on data management practices that emphasize provenance tracking as a core component of reproducible research. These practices apply directly to file format decisions because format interpretation depends on knowing the data source.

### Applying the Framework to Common Scenarios

The framework resolves common format selection problems through a consistent decision path.

**Scenario 1: Preparing input for a read assembler**

Question 1 identifies that the assembler requires FASTQ. Question 2 confirms that quality information is needed for error correction. Questions 3 and 4 do not apply because the input is raw reads, not assembled structure. Question 5 confirms that the FASTQ files have documented provenance including platform and encoding scheme. The decision is FASTQ with verified quality score encoding.

**Scenario 2: Submitting an assembled genome to a public database**

Question 1 identifies that the database accepts FASTA for sequence records. Question 2 determines that quality information is not part of the submission. Question 3 determines that structural information is not part of the standard submission format. Question 4 identifies this as a final product. Question 5 confirms that the FASTA file was derived from validated assembly output. The decision is FASTA with complete header documentation.

**Scenario 3: Investigating structural variation in a new assembly**

Question 1 identifies that graph analysis tools require GFA. Question 2 determines that quality information is not directly used in graph analysis. Question 3 confirms that structural relationships are the focus of the investigation. Question 4 identifies this as an intermediate analysis artifact. Question 5 confirms that the GFA file was produced by the assembler with documented parameters. The decision is GFA with graph statistics recorded.

**Scenario 4: Archiving raw sequencing data for future use**

Question 1 identifies that future tools may require FASTQ. Question 2 confirms that quality information is essential for future analyses. Question 3 determines that structural relationships are not present in raw reads. Question 4 identifies this as raw data requiring maximum information preservation. Question 5 confirms that the FASTQ files have complete provenance records. The decision is FASTQ with integrity checks performed and documented.

### A Record System for Format Decisions

Format decisions should be recorded systematically so that other researchers can understand why specific formats were chosen and what information was preserved or discarded. The following record structure captures the essential information for each format decision.

**Format Decision Record Template**

For every file created or converted, record the following fields:

| Field | Description | Example |
|-------|-------------|---------|
| File name | Full path and filename | reads_trimmed.fastq.gz |
| Source file | Original file and format | reads_raw.fastq.gz |
| Tool used | Software and version | fastp 0.23.4 |
| Parameters | Command-line options | --trim_front1 10 --qualified_quality_phred 20 |
| Date | Date of creation | 2025-03-15 |
| Format chosen | FASTA, FASTQ, or GFA | FASTQ |
| Information preserved | Quality, structure, or sequence only | Quality scores retained |
| Information discarded | What was lost in conversion | None |
| Rationale | Why this format was chosen | Required for assembler input |
| Validation performed | Checks applied to output | Record count matches input |

This record structure supports reproducibility because it captures both the decision and the reasoning behind it. The [Galaxy Training Network](https://training.galaxyproject.org/) emphasizes that reproducible workflows require documenting beyond the commands but also the decisions that led to those commands.

**Maintaining a Format Decision Log**

Create a single log file for the project that records every format decision in chronological order. This log serves as an audit trail that answers questions about data provenance and format choices. Update the log whenever a file is created, converted, or deleted.

The log should be stored with the project data and version controlled. The [Carpentries lessons](https://carpentries.org/lessons) provide foundational training on version control with Git, which is appropriate for maintaining format decision logs alongside analysis scripts and configuration files.

### Troubleshooting Format-Related Failures

When downstream analyses fail or produce unexpected results, format-related issues are a common cause. The following troubleshooting method systematically eliminates format problems before investigating other causes.

**Step 1: Verify File Integrity**

Check that the file is complete and uncorrupted. For FASTA files, verify that every header line begins with the greater-than symbol and that sequence lines contain only valid IUPAC characters. For FASTQ files, verify that every record has four lines and that sequence and quality lines have equal length. For GFA files, verify that segment, link, and path records are properly formatted and that links reference existing segments.

File integrity checks should be performed after transfer and before analysis. Truncated files from interrupted transfers are a common source of mysterious analysis failures.

**Step 2: Verify Format Interpretation**

Confirm that the tool is interpreting the file format correctly. For FASTQ files, verify the quality score encoding. A quick check of the quality character distribution reveals whether the encoding is Phred+33 or Phred+64. Phred+33 produces characters in the ASCII range 33 to 73, while Phred+64 produces characters in the range 64 to 104.

For GFA files, verify that the tool supports the GFA version used in the file. GFA version differences affect record syntax and tag definitions. Tools that expect one version may fail or misinterpret files in another version.

**Step 3: Verify Record Counts and Identifiers**

Compare the number of records in the file to the expected count. For FASTQ files, the expected count comes from the sequencing instrument output or previous processing steps. For FASTA files, the expected count comes from the assembly output or database records. For GFA files, the expected counts come from the assembler statistics.

Check that identifiers match expectations. Header identifiers should be consistent with the naming convention used in the project. Mismatched identifiers indicate that files were mixed or that a processing step renamed records unexpectedly.

**Step 4: Verify Information Preservation**

Confirm that the file contains the information the analysis requires. If the analysis needs quality scores, verify that the file is FASTQ and that quality characters are present. If the analysis needs structural relationships, verify that the file is GFA and that link records are present.

This step catches silent data loss from earlier conversions. A FASTA file created from FASTQ data will not contain quality information, and a FASTA file created from GFA data will not contain structural relationships. If the analysis requires this information, the source files must be located and the conversion redone with appropriate format choices.

**Step 5: Escalate Persistent Failures**

If format-related checks pass but the analysis still fails, escalate to tool developers or community support channels. Provide the format decision records, file integrity check results, and a minimal example that reproduces the failure. The [Bioconductor project](https://bioconductor.org/) provides support channels for its packages, and the [nf-core community](https://nf-co.re/docs) provides support for its pipelines.

### Common Failure Patterns and Their Resolution

The following failure patterns recur across assembly projects. Recognizing them speeds up troubleshooting and prevents wasted effort.

**Pattern 1: Assembler rejects FASTQ input**

The assembler reports that the input file is not valid FASTQ. Check that the file has four lines per record, that sequence and quality lines have equal length, and that quality characters are within the expected ASCII range. Check that the file is not compressed when the tool expects uncompressed input, or that the compression format is supported.

**Pattern 2: Quality scores appear inflated or deflated**

Downstream analyses report quality scores that are systematically too high or too low. This pattern indicates a quality score encoding mismatch. Verify the encoding scheme against the sequencing platform documentation. If the encoding cannot be determined, examine the quality character distribution to identify the likely scheme.

**Pattern 3: Assembly graph shows unexpected complexity**

The GFA file contains more segments or links than expected based on the organism and sequencing depth. This pattern may indicate genuine structural complexity or may indicate assembly errors. Compare the graph structure across multiple assemblers. The [Golden Wattle organelle study](https://doi.org/10.46471/gigabyte.36) demonstrated that different assemblers produced contrasting arrangements of genomic segments, with mapped reads providing evidence about which arrangements were supported by the data.

**Pattern 4: Downstream tool reports missing sequences**

A tool that expects FASTA input reports that sequences are missing. Check that the FASTA file contains all expected records and that headers match the expected identifiers. Check whether the file was truncated during transfer or storage. Check whether a conversion step dropped records that did not meet filtering criteria.

**Pattern 5: Graph analysis produces inconsistent results**

Different runs of the same graph analysis produce different results from the same GFA file. This pattern indicates that the analysis tool is not deterministic or that the GFA file is being modified between runs. Verify that the GFA file has not been altered and that the tool version and parameters are consistent.

### Integrating Format Decisions into Reproducible Workflows

Format decisions should be embedded in reproducible workflow definitions instead of made ad hoc during analysis. Workflow frameworks provide structured ways to define input and output formats for each step.

The [nf-core documentation](https://nf-co.re/docs) describes how community pipelines define expected input and output formats for each process. These definitions make format requirements explicit and prevent accidental format mismatches. Pipelines also include validation steps that check file formats before processing begins.

The [Galaxy Training Network](https://training.galaxyproject.org/) provides tutorials on building reproducible analysis workflows that include format validation and documentation. These workflows make format decisions visible and auditable, supporting the reproducibility goals that are central to modern genomics research.

The [Bioconductor project](https://bioconductor.org/) provides R packages for genomic analysis that include functions for reading and validating FASTA, FASTQ, and GFA files. These packages enforce format standards and provide clear error messages when files do not conform to expected structures.

### Measuring the Impact of Format Decisions

Format decisions affect storage requirements, analysis time, and the types of analyses that can be performed. Measuring these impacts helps researchers make informed choices and plan for resource requirements.

**Storage Impact**

FASTQ files are typically three to four times larger than the corresponding FASTA files because they store quality characters for every base. GFA files vary in size depending on graph complexity. Record the file sizes for each format in the project log to track storage requirements.

**Analysis Time Impact**

Some analyses run faster on FASTA input because they skip quality score processing. Other analyses require FASTQ input and cannot run on FASTA at all. Record the runtime for each analysis step to understand the time implications of format choices.

**Analysis Capability Impact**

Format choices determine which analyses are possible. Quality-based analyses require FASTQ. Structural analyses require GFA. Sequence-only analyses can use FASTA. Record which analyses were performed and which were not possible due to format constraints.

These measurements provide evidence for future format decisions. When planning a new project, review the measurements from previous projects to estimate storage, time, and capability requirements for different format choices.

### Professional Escalation Criteria for Format Problems

Most format problems can be resolved through the troubleshooting method described above. Certain situations warrant escalation to specialists with deeper expertise.

**Escalate when quality score encoding cannot be determined**

If the sequencing platform documentation is unavailable and the quality character distribution is ambiguous, escalate to a bioinformatics specialist. Incorrect quality score interpretation can invalidate all downstream analyses that use quality information.

**Escalate when GFA files from different assemblers cannot be reconciled**

If different assemblers produce GFA files with fundamentally different structures and mapped reads do not clearly support one arrangement over another, escalate to assembly specialists. Reconciling conflicting graph structures requires expertise in assembly algorithms and graph theory.

**Escalate when format conversion produces unexpected results**

If converting between formats produces files that fail validation or produce unexpected record counts, escalate to tool developers. Conversion tools may have bugs or may implement format specifications differently than expected.

**Escalate when data sharing requirements conflict with format choices**

If a data repository requires a specific format that would require discarding information, escalate to institutional data governance officers. The [framework for genomic data sharing](https://doi.org/10.1038/s41588-024-02049-2) describes governance considerations that affect data deposition decisions, including format requirements and information preservation obligations.

## Frequently Asked Questions

### What is the difference between FASTA and FASTQ files?

FASTA files contain sequence data only, with a header line followed by sequence lines. FASTQ files contain sequence data plus a quality score for every nucleotide position, organized as four lines per record. FASTQ files are larger but preserve the information needed for quality-based analyses such as read trimming and error correction.

### When should I convert FASTQ to FASTA?

Convert FASTQ to FASTA only when quality information is not needed for downstream analysis and when the original FASTQ files are retained for future use. Common reasons include submitting sequences to databases that require FASTA format and running tools that only accept FASTA input. The conversion permanently discards quality scores, so it should not be performed casually.

### What information does a GFA file contain that FASTA does not?

GFA files contain sequence segments plus the connections between them, representing the assembly graph structure. This includes alternative paths, branching structures, and relationships between segments that linear FASTA representations cannot express. GFA files are essential for understanding structural variation, resolving haplotypes, and analyzing repeats.

### How do I choose the correct quality score encoding for FASTQ files?

Determine the sequencing platform and software version used to generate the data. Current Illumina platforms and most other technologies use Phred+33 encoding, while older Illumina platforms used Phred+64. When documentation is unavailable, examine the distribution of quality characters. Phred+33 produces characters in the ASCII range 33 to 73, while Phred+64 produces characters in the range 64 to 104.

### Can I convert GFA to FASTA without losing information?

Converting GFA to FASTA always loses structural information because FASTA cannot represent graph branches or alternative paths. The conversion produces a linear sequence that represents one path through the graph. Document the path selection method and retain the original GFA file to preserve the full structural information.

### Why do different assemblers produce different GFA files from the same data?

Assembly algorithms use different strategies for resolving repeats, handling sequencing errors, and constructing the graph. These differences produce different graph structures from the same input data. The [organelle genome assemblies of the Golden Wattle](https://doi.org/10.46471/gigabyte.36) demonstrated that different assembly algorithms produced contrasting arrangements of genomic segments, with mapped reads supporting alternate paths.

### What quality metrics should I record for my assembly project?

Record read quality metrics from FASTQ files before assembly, including per-base quality distributions, read counts, read length distributions, and GC content. Record assembly statistics from FASTA output after assembly, including total size, contig count, N50, and maximum contig length. For graph-based assemblies, record graph statistics from GFA files, including segment counts, link counts, and alternative path information.

### How should I document file format conversions in my project records?

Record the tool name, version, parameters, input file, output file, and date for every conversion. For GFA to FASTA conversions, document the path selection method. For FASTQ to FASTA conversions, note that quality information was discarded. This documentation supports reproducibility and helps other researchers understand the provenance of the data.

## Related Bioinformatics Guides

- [Genomic Data Analysis Tools: A Comparative Guide for Researchers](/knowledge/bioinformatics/genomic-data-analysis-tools-a-comparative-guide-for-researchers)
- [Evaluating Genome Assembly Quality: Metrics and Tools](/knowledge/bioinformatics/evaluating-genome-assembly-quality-metrics-and-tools)
- [Metagenome Co-Assembly: Strategies for Multi-Sample Data](/knowledge/bioinformatics/metagenome-co-assembly-strategies-for-multi-sample-data)
- [Persistent Identifiers for Research Data: A Guide to Selection and Use](/knowledge/bioinformatics/persistent-identifiers-for-research-data-a-guide-to-selection-and-use)
- [De Novo Genome Assembly with Long Reads: A Practical Workflow](/knowledge/bioinformatics/de-novo-genome-assembly-with-long-reads-a-practical-workflow)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Overcoming challenges associated with broad sharing of human genomic data.](https://doi.org/10.1038/s41588-024-02049-2). 2025.
- [MGS2AMR: a gene-centric mining of metagenomic sequencing data for pathogens and their antimicrobial resistance profile.](https://doi.org/10.1186/s40168-023-01674-z). 2023.
- [Average Nucleotide Identity and Digital DNA-DNA Hybridization Analysis Following PromethION Nanopore-Based Whole Genome Sequencing Allows for Accurate Prokaryotic Typing.](https://doi.org/10.3390/diagnostics14161800). 2024.
- [Long-read assemblies reveal structural diversity in genomes of organelles - an example with <i>Acacia pycnantha</i>.](https://doi.org/10.46471/gigabyte.36). 2021.
- [SyntenyPair Explorer: an installation-free, browser-based tool for interactive pairwise genome synteny visualization](https://doi.org/10.64898/2026.07.23.740353). bioRxiv, 2026.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.