Skip to content

Generative AI

Bag-of-Words Vectors

O(s · w) time to build one vector, where s is the sentence length and w is the vocabulary size · O(w) space per vector.

The idea, in plain English

Imagine you dump every word from a sentence into a bag. Then you count how many of each word landed inside. Word order no longer matters — only the totals do. This is a bag-of-words vector: a list of numbers, with one count per word. Two sentences that use a lot of the same words end up with very similar lists of numbers.

How it works

  1. 1Collect every unique word across all the sentences you care about. This is your shared vocabulary.
  2. 2Sort the vocabulary so the word order stays fixed and predictable, for example alphabetically.
  3. 3For each sentence, build a vector that is the same length as the vocabulary. Slot i counts how many times vocabulary word i appears in that sentence.

When you'd use it

This is a simple, classic way to turn text into numbers for comparison or basic search. It is the ancestor of the embedding vectors that modern models use (see Cosine Similarity and Vector Search / RAG Retrieval).

Common beginner mistakes

  • Don't forget that word order is completely thrown away. 'Dog bites man' and 'man bites dog' produce the exact same bag-of-words vector.
  • Don't build a different vocabulary for each sentence. Vectors are only comparable if you build them against the same shared vocabulary.

Try it — edit and run

Click the code to edit · press ⌘/Ctrl+↵ to run

Editable code. Tab and Shift+Tab indent. Press Escape, then Tab, to move focus out of the editor.

Expected output — hit Run to try it
Vocabulary: cat dog log mat on sat the
Sentence 1: the cat sat on the mat
Sentence 1 vector: 1 0 0 1 1 1 2
Sentence 2: the dog sat on the log
Sentence 2 vector: 0 1 1 0 1 1 2

Not sure this is the right topic? See the learning paths → or where this leads →