🔄 Seq2Seq & Attention Mechanisms
127 words
1 min read
Visual companion
Python
Type and operator map
Python Week 1: the first filter for runtime behavior
View
Revision summary
What this note is really saying
Short form
# 🔄 Seq2Seq & Attention Mechanisms ## 1. 🎯 Learning Objectives - Explain encoder-decoder for sequence tasks - Compute attention weights for alignment - Describe attention variants: additive, multiplicative, self-attention ## 2.

🔄 Seq2Seq & Attention Mechanisms
1. 🎯 Learning Objectives
- Explain encoder-decoder for sequence tasks
- Compute attention weights for alignment
- Describe attention variants: additive, multiplicative, self-attention
2. 📖 Core Content
3.1 Encoder-Decoder
Encoder RNN: reads input sequence, produces context vector c. Decoder RNN: generates output sequence from c. Problem: Information bottleneck — c must capture entire input sequence.
3.2 Attention Mechanism
Compute attention weights: αₜᵢ = softmax(score(hₜ, sᵢ)) Context vector: cₜ = Σᵢ αₜᵢ hᵢ Score functions:
- Additive (Bahdanau): vᵃ tanh(W₁hₜ + W₂sᵢ)
- Dot product (Luong): hₜᵀ sᵢ
- Scaled dot: hₜᵀ sᵢ / √d
3.3 Benefits
- No information bottleneck
- Alignment learned end-to-end
- Interpretability (visualize attention weights) Join Discord PreviousTransformers & AttentionNextTransformer Deep Dive