Data Science
Naive Bayes (tiny classifier)
Training takes O(n) time, where n is the total number of words across all training documents — just counting. Classifying a new message with m words takes O(m) time, since each word needs one lookup.
The idea, in plain English
Imagine a mail sorter who has read thousands of letters. They learned that words like 'win' and 'free' show up a lot in junk mail, while words like 'lunch' and 'meeting' show up a lot in real messages. When a new letter arrives, the sorter does not read it deeply. They just ask, for each word in it, 'how often did I see this word in spam versus real mail?' and multiply those odds together. Naive Bayes is that mail sorter turned into math. It is called 'naive' because it assumes every word's presence is independent of every other word. This assumption is wrong, but it makes the math simple, and the method still works well in practice.
How it works
- 1Count how often every word appears in each class of training document (here: 'spam' and 'ham'), plus how many documents belong to each class.
- 2For a new message, start with the 'prior' — how common each class is overall. Then multiply in the probability of seeing each of the message's words in that class.
- 3Add 1 to every word count before dividing. This is called Laplace smoothing, and it stops a word the model has never seen from zeroing out the whole calculation.
- 4Whichever class ends up with the higher score after this multiplication is the prediction. Normalize the scores so they add up to 100% and read them as a probability.
When you'd use it
This is the classic go-to for text classification with small-to-medium data — spam filtering, sentiment tagging, or routing support tickets by topic. Use it anywhere 'which words appear' is a strong enough signal on its own, without needing word order or grammar.
Common beginner mistakes
- Do not skip Laplace smoothing. Without it, a single word the model has never seen in a class multiplies that class's score down to exactly zero, no matter how well every other word matched.
- Do not treat Naive Bayes' 'independence' assumption as literally true. Words in real language depend on each other, but the model works well anyway because it only needs to get the ranking between classes right, not the exact probabilities.
Try it — edit and run
Click the code to edit · press ⌘/Ctrl+↵ to run
Editable code. Tab and Shift+Tab indent. Press Escape, then Tab, to move focus out of the editor.
Test message: win a free lunch
Vocabulary size: 20
P(spam): 0.86
P(ham): 0.14
Prediction: spamNot sure this is the right topic? See the learning paths → or where this leads →