Generative AI
Self-Attention
O(n² · d) time and O(n²) space for a sequence of n tokens with vectors of d dimensions. Every token computes a score against every other token, which is why very long inputs get expensive fast.
The idea, in plain English
Imagine reading the sentence: 'The trophy didn't fit in the suitcase because it was too big.' To figure out what 'it' refers to, you glance back at the earlier words and decide the trophy is more relevant than the suitcase. Self-attention is a language model doing exactly that, for every single word. It looks at all the other words, and itself, in the sentence, and decides how much to 'focus' on each one. Each word gets a weight between 0 and 1. Words that seem more relevant get a bigger weight. Blending all the words together using those weights gives the model a new, context-aware understanding of the word it is focusing from.
How it works
- 1Give each token, or word, a vector: a small list of numbers that stands in for its meaning. This is a toy embedding.
- 2For the token you are computing attention from, take the dot product of its vector with every token's vector, including its own. Each raw number you get is the attention score for that pair.
- 3Run all of that token's scores through softmax (see Softmax & Temperature) to turn them into attention weights that add up to 1. A bigger score becomes a bigger weight.
- 4Multiply every token's vector by its attention weight, then add all the results together. This weighted sum is the context vector: a new representation of the original token, blended with whatever else in the sentence it decided to focus on.
When you'd use it
This is the mechanism at the heart of Transformers, the architecture behind modern language models. It is how a model figures out which other words in a sentence matter for understanding or generating any given word, instead of only ever looking at fixed nearby positions.
Common beginner mistakes
- Don't assume a word only pays attention to its immediate neighbors. Self-attention scores every word against every other word in the sequence, no matter how far apart they are.
- Don't skip the softmax step. Raw dot-product scores are not weights yet. They need to become a normalized set of numbers that add up to 1 before you can use them to blend vectors together.
- Don't assume this toy version works exactly like real Transformers. Real self-attention uses three different learned projections of each token, called query, key, and value, instead of reusing the same raw vector for all three roles. This lets the model learn different notions of 'relevance', rather than just raw similarity.
Try it — edit and run
Click the code to edit · press ⌘/Ctrl+↵ to run
Editable code. Tab and Shift+Tab indent. Press Escape, then Tab, to move focus out of the editor.
Tokens: the cat sat
Attention from "the":
-> the: score=2 weight=0.58
-> cat: score=1 weight=0.21
-> sat: score=1 weight=0.21
context vector: [0.79, 0.42, 0.79]
Attention from "cat":
-> the: score=1 weight=0.21
-> cat: score=2 weight=0.58
-> sat: score=1 weight=0.21
context vector: [0.42, 0.79, 0.79]
Attention from "sat":
-> the: score=1 weight=0.21
-> cat: score=1 weight=0.21
-> sat: score=2 weight=0.58
context vector: [0.79, 0.79, 0.42]Not sure this is the right topic? See the learning paths → or where this leads →