Sep 02, 20269 min read
What attention heads actually attend to
- Parameters
- 370M
- Tokens
- 20B
- Hardware
- 4x A100
- Training time
- 3 days
I sorted every head in tern-370m by the pattern it produces on held-out text. Most fall into a handful of simple types.
Findings
About a third of heads attend mostly to the previous token. A smaller group copies earlier tokens, and roughly one in ten could be removed with no measurable change in loss.