Skip to content

Generative AI

Cosine Similarity

O(d) time and space, where d is the number of dimensions in each vector.

The idea, in plain English

Imagine a map where words with similar meanings sit close together. 'King' and 'queen' are near neighbors, while 'bread' sits in a different part of the map. Each word's spot on that map is called its embedding: a list of numbers that stands in for its meaning. Cosine similarity is a score for how close two spots are. But instead of measuring straight-line distance, it compares the angle between them, seen from the map's center. The score always lands between -1 and 1. A score of 1.00 means the two point in the exact same direction, so they are very similar. A score of 0.00 means they are unrelated. A negative score means they point in close to opposite directions.

How it works

  1. 1Take the dot product of the two vectors. Multiply the matching positions together, then add up the results.
  2. 2Compute each vector's magnitude, or length. Square every number in it, add the squares together, then take the square root.
  3. 3Divide the dot product by the product of the two magnitudes. The result is always between -1 and 1. A bigger number means the two are more similar.

When you'd use it

Use this when comparing embeddings, which are number representations of meaning. It helps you find similar words, similar documents, or which stored chunk of text best matches a search query.

Common beginner mistakes

  • Don't compare raw dot products without dividing by the magnitudes. That unfairly favors longer vectors over vectors that are genuinely more similar.
  • Don't compare vectors with different lengths, or dimensions. They must come from the same embedding space, or the comparison means nothing.

Try it — edit and run

Click the code to edit · press ⌘/Ctrl+↵ to run

Editable code. Tab and Shift+Tab indent. Press Escape, then Tab, to move focus out of the editor.

Expected output — hit Run to try it
king vs queen similarity: 0.96
king vs bread similarity: 0.60

Not sure this is the right topic? See the learning paths → or where this leads →