Learning Objectives
400 words
2 min read
Visual companion
Python
Type and operator map
Python Week 1: the first filter for runtime behavior
View
Revision summary
What this note is really saying
Short form
# Learning Objectives - Understand the genome assembly problem - Compare OLC and De Bruijn approaches - Evaluate assembly quality with N50 ## 1. The Assembly Problem Reconstruct the original genome from short sequencing reads.

Learning Objectives
- Understand the genome assembly problem
- Compare OLC and De Bruijn approaches
- Evaluate assembly quality with N50
1. The Assembly Problem
Reconstruct the original genome from short sequencing reads. Challenges: repeats longer than reads create ambiguities, sequencing errors, non-uniform coverage, large genomes (human: 3 billion bp), heterozygosity.
2. Assembly Approaches
Overlap-Layout-Consensus (OLC):
- Find all read overlaps (pairwise O(n^2))
- Build overlap graph, find path
- Derive consensus sequence from path Best for long reads (Sanger, PacBio). Used by Canu, Celera Assembler. De Bruijn Graph:
- Break reads into k-mers (substrings of length k)
- Build graph: k-mers as edges, (k-1)-mers as nodes
- Find Eulerian path (visit each edge exactly once)
- Traverse path to reconstruct sequence Best for short reads (Illumina). Used by Velvet, SPAdes.
3. Assembly Quality Metrics
- N50: 50% of assembled bases are in contigs >= N50 length. Higher = better.
- L50: Number of contigs covering 50% of assembly. Lower = better.
- BUSCO: Benchmarking Universal Single-Copy Orthologs - measures completeness.
Q1: Why is genome assembly difficult?Repeats longer than read length create ambiguous assembly. Sequencing errors break contiguity. Non-uniform coverage creates gaps. Large genomes require massive computation. Q2: What is the difference between OLC and De Bruijn assembly?OLC: overlap reads O(n^2), best for long reads. De Bruijn: break reads into k-mers, graph-based, best for short reads, more efficient for large datasets. Q3: What is N50?Length-weighted median: 50% of assembled bases are in contigs >= N50. Higher N50 = more contiguous assembly. Example: N50=100kb means half the assembly is in contigs >= 100kb. Q4: What is a contig vs a scaffold?Contig: contiguous sequence from reads (no gaps). Scaffold: contigs joined by linking information with gaps (Ns) between contigs. Scaffolds are longer but have unknown bases. Q5: How to choose k-mer size in De Bruijn graphs?Smaller k: better for low coverage, more connected graph, more repeats. Larger k: more specific, fewer repeats, more fragmented. Typical range: 21-127 depending on read length and coverage. Q6: What are BUSCOs?Benchmarking Universal Single-Copy Orthologs - genes expected to be present as single copies in most species. Assembly completeness = % of BUSCOs found. >90% is good for eukaryotes. Join Discord PreviousMotif Finding & Regulatory GenomicsNextSequence Analysis & Genomics