Your next step
Attention across a sequence
Self and cross attention, order signals, masking, many heads.
Start learningWhat you’ll learn
Self-attention uses one sequence
Identify where queries, keys and values originate.
About 4 minutes · Open activity
Each position gets its own update
Recognize that attention is repeated across positions.
About 4 minutes · Open activity
Representations become contextual
Explain a token's changed representation after incorporating other positions.
About 4 minutes · Open activity
Positions need an order signal
Explain why a Transformer needs position information.
About 4 minutes · Open activity
A causal mask prevents looking ahead
Identify legal context for next-token prediction.
About 4 minutes · Open activity
Cross-attention connects two sequences
Identify attention from an output position to source representations.
About 4 minutes · Open activity
Multiple heads offer several learned views
Explain why heads can combine different relationships.
About 4 minutes · Open activity
A heatmap is only one view
Interpret an attention picture with appropriate limits.
About 5 minutes · Open activity
Compare three attention diagrams
Classify self versus cross attention, block an illegal future link, and state what a heatmap cannot establish.
About 6 minutes · Open activity