Learning Objectives
483 words
2 min read
Visual companion
Python
Type and operator map
Python Week 1: the first filter for runtime behavior
View
Revision summary
What this note is really saying
Short form
# Learning Objectives - Understand clustering for market segmentation - Apply k-means and hierarchical clustering - Determine optimal number of clusters - Distance metrics (Euclidean distance) - Week 6: Factor analysis concepts ## 1. What is Cluster Analysis?

Learning Objectives
- Understand clustering for market segmentation
- Apply k-means and hierarchical clustering
- Determine optimal number of clusters
- Distance metrics (Euclidean distance)
- Week 6: Factor analysis concepts
1. What is Cluster Analysis?
Intuition: Group similar customers together based on their characteristics (demographics, behavior, attitudes). Customers in same cluster are more similar to each other than to customers in other clusters.
Applications: Market segmentation, customer profiling, product positioning, test market selection, classification of respondents.
2. K-Means Clustering
Algorithm:
- Choose K (number of clusters)
- Randomly initialize K centroids (cluster centers)
- Assign each data point to nearest centroid (Euclidean distance)
- Recalculate centroids as mean of points in cluster
- Repeat steps 3-4 until convergence (no changes in assignments) Advantages: Fast, works well for large datasets, easy to understand Disadvantages: Need to specify K, sensitive to initial centroids, sensitive to outliers, assumes spherical clusters
3. Hierarchical Clustering
Agglomerative (bottom-up): Each point starts as its own cluster. Merge the two closest clusters at each step. Continue until one cluster remains.
Divisive (top-down): All points start as one cluster. Recursively split into smaller clusters.
Linkage Criteria:
- Single linkage: distance between closest points (produces chains)
- Complete linkage: distance between farthest points (produces compact clusters)
- Average linkage: average distance between all pairs
- Ward's method: minimize within-cluster variance (most common) Dendrogram: Tree diagram showing hierarchical structure. Cut at appropriate height to determine clusters.
4. Determining Number of Clusters
- Elbow Method: Plot within-cluster sum of squares vs K. Look for bend.
- Silhouette Score: Measures cohesion (within cluster) vs separation (between clusters). Range -1 to 1. Higher is better.
- Davies-Bouldin Index: Average similarity between each cluster and its most similar one. Lower is better.
- Interpretability: Do clusters make business sense?
Q1: What is the purpose of cluster analysis in marketing?Segment customers into homogeneous groups for targeted marketing: different products, prices, promotions, communication channels for each segment. Q2: How does k-means clustering work step by step?
- Choose K. 2) Initialize K centroids. 3) Assign points to nearest centroid. 4) Update centroids to mean of points. 5) Repeat steps 3-4 until convergence. Q3: What is a dendrogram and how is it used?
Tree diagram showing hierarchical clustering steps. Height of merge = similarity. To determine clusters, cut horizontally at desired similarity level. Number of vertical lines intersected = number of clusters. Q4: How to choose optimal K in k-means?Elbow method (WCSS vs K), silhouette score (higher better), Davies-Bouldin index (lower better), plus business considerations: clusters should be interpretable, actionable, and of meaningful size. Q5: Compare k-means and hierarchical clustering.K-means: fast, large datasets, need to pre-specify K. Hierarchical: slower, smaller datasets (< 1000), no need to specify K, produces dendrogram, more informative. Join Discord PreviousFactor AnalysisNextConjoint Analysis