BSDA5006 — Deep Learning for Computer Vision (DL-CV)
247 words
1 min read
Visual companion
Python
Type and operator map
Python Week 1: the first filter for runtime behavior
View
Revision summary
What this note is really saying
Short form
# BSDA5006 — Deep Learning for Computer Vision (DL-CV) > **Course**: Degree Elective (4 credits) **Topics**: CNNs, Advanced architectures, Object detection, Segmentation, GANs, Vision Transformers **Files**: 10 comprehensive topic files ## Course Index File Topic Key Concepts 01-cnn-fundamentals CNN Architecture Con...

BSDA5006 — Deep Learning for Computer Vision (DL-CV)
Course: Degree Elective (4 credits) Topics: CNNs, Advanced architectures, Object detection, Segmentation, GANs, Vision Transformers Files: 10 comprehensive topic files
Course Index
| File | Topic | Key Concepts |
|---|---|---|
| 01-cnn-fundamentals | CNN Architecture | Convolution, pooling, stride, padding, feature maps, receptive field |
| 02-advanced-cnns | Advanced CNNs | VGG, Inception, ResNet, DenseNet, MobileNet, EfficientNet |
| 03-transfer-learning | Transfer Learning | Pre-training, fine-tuning, feature extraction, domain adaptation |
| 04-object-detection | Object Detection | R-CNN, Fast R-CNN, YOLO, SSD, anchor boxes, IoU, NMS, mAP |
| 05-segmentation | Segmentation | Semantic, instance, panoptic, FCN, U-Net, Mask R-CNN |
| 06-gans | GANs | Generator, discriminator, DCGAN, CGAN, WGAN, CycleGAN |
| 07-vision-transformers | Vision Transformers | ViT, patch embedding, positional encoding, hybrid architectures |
| 08-data-augmentation | Data Augmentation | Transformations, MixUp, CutMix, RandAugment, AutoAugment |
| 09-image-processing | Image Processing | Gradients, filtering, edge detection, frequency domain, histograms |
| 10-diffusion-models | Diffusion Models | DDPM, noise scheduling, U-Net for denoising, sampling |
Key Formulas
| Formula | Description |
|---|---|
| O=(W−K+2P)/S+1 | Output size after convolution |
| Receptive Field=1+∑(Ki−1) | RF computation for stacked conv layers |
| IoU=TP+FP+FNTP | Intersection over Union |
| mAP=N1∑AP | Mean Average Precision |
| LGAN=minGmaxD[logD(x)+log(1−D(G(z)))] | GAN loss |
| Patch Embedding(I)=Linear(Concat(p1,p2,...,pn)) | ViT patch embedding |
Exam Weightage
| Topic | Quiz 1 | Quiz 2 | End Term |
|---|---|---|---|
| Image Processing | ★★★★★ | ★★★ | ★★ |
| CNN Fundamentals | ★★★★★ | ★★★★ | ★★★ |
| Advanced CNNs | ★★ | ★★★★★ | ★★★★ |
| Transfer Learning | ★★ | ★★★★★ | ★★★ |
| Object Detection | ★ | ★★★ | ★★★★★ |
| Segmentation | ★ | ★★★ | ★★★★★ |
| GANs | ★ | ★★★ | ★★★★ |
| Vision Transformers | ★ | ★★ | ★★★★ |
Cross-Course Links
- BSDA5004 (LLMs): Transformer attention is shared with ViTs
- BSDA5002 (GenAI Foundations): GANs and diffusion models covered in-depth
- BSDA5013 (DL Practice): Hands-on CV implementation projects Join Discord PreviousDiffusion Models for CV