Quiz 2

BSDA5006 — Deep Learning for Computer Vision (DL-CV)

247 words
1 min read
Python Week 1: the first filter for runtime behavior
Visual companion
Python
Type and operator map

Python Week 1: the first filter for runtime behavior

View
Revision summary

What this note is really saying

Short form

# BSDA5006 — Deep Learning for Computer Vision (DL-CV) > **Course**: Degree Elective (4 credits) **Topics**: CNNs, Advanced architectures, Object detection, Segmentation, GANs, Vision Transformers **Files**: 10 comprehensive topic files ## Course Index File Topic Key Concepts 01-cnn-fundamentals CNN Architecture Con...

BSDA5006 — Deep Learning for Computer Vision (DL-CV)

Course: Degree Elective (4 credits) Topics: CNNs, Advanced architectures, Object detection, Segmentation, GANs, Vision Transformers Files: 10 comprehensive topic files

Course Index

FileTopicKey Concepts
01-cnn-fundamentalsCNN ArchitectureConvolution, pooling, stride, padding, feature maps, receptive field
02-advanced-cnnsAdvanced CNNsVGG, Inception, ResNet, DenseNet, MobileNet, EfficientNet
03-transfer-learningTransfer LearningPre-training, fine-tuning, feature extraction, domain adaptation
04-object-detectionObject DetectionR-CNN, Fast R-CNN, YOLO, SSD, anchor boxes, IoU, NMS, mAP
05-segmentationSegmentationSemantic, instance, panoptic, FCN, U-Net, Mask R-CNN
06-gansGANsGenerator, discriminator, DCGAN, CGAN, WGAN, CycleGAN
07-vision-transformersVision TransformersViT, patch embedding, positional encoding, hybrid architectures
08-data-augmentationData AugmentationTransformations, MixUp, CutMix, RandAugment, AutoAugment
09-image-processingImage ProcessingGradients, filtering, edge detection, frequency domain, histograms
10-diffusion-modelsDiffusion ModelsDDPM, noise scheduling, U-Net for denoising, sampling

Key Formulas

FormulaDescription
O=(WK+2P)/S+1O = (W - K + 2P)/S + 1Output size after convolution
Receptive Field=1+(Ki1)\text{Receptive Field} = 1 + \sum (K_i - 1)RF computation for stacked conv layers
IoU=TPTP+FP+FN\text{IoU} = \frac{TP}{TP + FP + FN}Intersection over Union
mAP=1NAPmAP = \frac{1}{N}\sum APMean Average Precision
LGAN=minGmaxD[logD(x)+log(1D(G(z)))]\mathcal{L}_{GAN} = \min_G\max_D [\log D(x) + \log(1-D(G(z)))]GAN loss
Patch Embedding(I)=Linear(Concat(p1,p2,...,pn))\text{Patch Embedding}(I) = \text{Linear}(\text{Concat}(p_1, p_2, ..., p_n))ViT patch embedding

Exam Weightage

TopicQuiz 1Quiz 2End Term
Image Processing★★★★★★★★★★
CNN Fundamentals★★★★★★★★★★★★
Advanced CNNs★★★★★★★★★★★
Transfer Learning★★★★★★★★★★
Object Detection★★★★★★★★
Segmentation★★★★★★★★
GANs★★★★★★★
Vision Transformers★★★★★★
  • BSDA5004 (LLMs): Transformer attention is shared with ViTs
  • BSDA5002 (GenAI Foundations): GANs and diffusion models covered in-depth
  • BSDA5013 (DL Practice): Hands-on CV implementation projects Join Discord PreviousDiffusion Models for CV
Document outline

Keep your place and jump directly to a heading.

Table of Contents
System Normal // Awaiting Context

Intelligence Hub

Navigate the knowledge graph to generate context. The Hub adapts dynamically to surface backlinks, related notes, and metadata insights.