Quiz 2
Registry Synced

Learning Objectives

400 words
2 min read

Reading compass

Now · 1. The Assembly Problem

Learning Objectives

  • Understand the genome assembly problem
  • Compare OLC and De Bruijn approaches
  • Evaluate assembly quality with N50

1. The Assembly Problem

Reconstruct the original genome from short sequencing reads. Challenges: repeats longer than reads create ambiguities, sequencing errors, non-uniform coverage, large genomes (human: 3 billion bp), heterozygosity.

2. Assembly Approaches

Overlap-Layout-Consensus (OLC):
  1. Find all read overlaps (pairwise O(n^2))
  2. Build overlap graph, find path
  3. Derive consensus sequence from path Best for long reads (Sanger, PacBio). Used by Canu, Celera Assembler. De Bruijn Graph:
  4. Break reads into k-mers (substrings of length k)
  5. Build graph: k-mers as edges, (k-1)-mers as nodes
  6. Find Eulerian path (visit each edge exactly once)
  7. Traverse path to reconstruct sequence Best for short reads (Illumina). Used by Velvet, SPAdes.

3. Assembly Quality Metrics

  • N50: 50% of assembled bases are in contigs >= N50 length. Higher = better.
  • L50: Number of contigs covering 50% of assembly. Lower = better.
  • BUSCO: Benchmarking Universal Single-Copy Orthologs - measures completeness.
Q1: Why is genome assembly difficult?
Repeats longer than read length create ambiguous assembly. Sequencing errors break contiguity. Non-uniform coverage creates gaps. Large genomes require massive computation. Q2: What is the difference between OLC and De Bruijn assembly?
OLC: overlap reads O(n^2), best for long reads. De Bruijn: break reads into k-mers, graph-based, best for short reads, more efficient for large datasets. Q3: What is N50?
Length-weighted median: 50% of assembled bases are in contigs >= N50. Higher N50 = more contiguous assembly. Example: N50=100kb means half the assembly is in contigs >= 100kb. Q4: What is a contig vs a scaffold?
Contig: contiguous sequence from reads (no gaps). Scaffold: contigs joined by linking information with gaps (Ns) between contigs. Scaffolds are longer but have unknown bases. Q5: How to choose k-mer size in De Bruijn graphs?
Smaller k: better for low coverage, more connected graph, more repeats. Larger k: more specific, fewer repeats, more fragmented. Typical range: 21-127 depending on read length and coverage. Q6: What are BUSCOs?
Benchmarking Universal Single-Copy Orthologs - genes expected to be present as single copies in most species. Assembly completeness = % of BUSCOs found. >90% is good for eukaryotes. Join Discord PreviousMotif Finding & Regulatory GenomicsNextSequence Analysis & Genomics
Document outline

Keep your place and jump directly to a heading.

Table of Contents
System Normal // Awaiting Context

Intelligence Hub

Navigate the knowledge graph to generate context. The Hub adapts dynamically to surface backlinks, related notes, and metadata insights.