Skip to content

Data Science

K-Nearest Neighbors

This takes O(n) time per prediction, since you check the distance to every existing point. There is no separate training step, but predictions get slower as the dataset grows.

The idea, in plain English

K-nearest neighbors works like peer pressure: 'you are like the people standing closest to you.' To label a new point, measure its distance to every point you already know the label for. Look at the k closest ones. Let them vote. Whichever label is most common among those neighbors becomes your prediction.

How it works

  1. 1Measure the distance from the new point to every labeled point you have.
  2. 2Sort those distances and keep the k closest labeled points. These are the 'nearest neighbors'.
  3. 3Count up the labels among those k neighbors and predict whichever label appears most often.

When you'd use it

Use this for simple classification tasks with a reasonably small, well-labeled dataset. Examples include recommending a product category based on similar past customers, or classifying a flower species from its measurements.

Common beginner mistakes

  • Do not pick an even k in a two-class problem. This can produce ties in the vote.
  • Do not skip scaling your features first. A feature measured in the thousands (like income) can dominate the distance calculation over one measured in single digits (like age).

Try it — edit and run

Click the code to edit · press ⌘/Ctrl+↵ to run

Editable code. Tab and Shift+Tab indent. Press Escape, then Tab, to move focus out of the editor.

Expected output — hit Run to try it
Query point: 2,2
Predicted label: A

Not sure this is the right topic? See the learning paths → or where this leads →