Data Science
K-Means Clustering
Each round takes O(n * k) time to compare every point against every center. This repeats for a fixed number of rounds until the centers settle.
The idea, in plain English
Imagine sorting a pile of mixed candy into a few bowls by color, but you do not know the colors ahead of time. You start by guessing where a few 'bowl centers' are. Each piece of candy goes into the bowl with the closest center. Then you move each center to the middle of the candy that landed in its bowl. You repeat this until the bowls stop changing.
How it works
- 1Pick k starting center points (here we fix them ahead of time so the result is the same every run).
- 2Assign every data point to whichever center is closest to it. This creates k clusters.
- 3Move each center to the average position of the points now assigned to it. Repeat the assign-and-move steps until the centers stop moving.
When you'd use it
Use this to group similar things when you do not have labels for them, such as customer segmentation, grouping similar images, or finding natural clusters in sensor readings.
Common beginner mistakes
- Do not pick random starting centers without fixing them. Different runs can land on different, and sometimes worse, groupings.
- Do not choose the wrong number of clusters (k). Too few lumps unrelated things together. Too many splits a real group apart.
Try it — edit and run
Click the code to edit · press ⌘/Ctrl+↵ to run
Editable code. Tab and Shift+Tab indent. Press Escape, then Tab, to move focus out of the editor.
Points: 1 2 3 10 11 12
Initial centers: 1 10
Final centers: 2.00 11.00
Cluster 1: 1 2 3
Cluster 2: 10 11 12Not sure this is the right topic? See the learning paths → or where this leads →