All writing
Sep 02, 20269 min read

What attention heads actually attend to

T = 2048
Parameters
370M
Tokens
20B
Hardware
4x A100
Training time
3 days

I sorted every head in tern-370m by the pattern it produces on held-out text. Most fall into a handful of simple types.

Findings

About a third of heads attend mostly to the previous token. A smaller group copies earlier tokens, and roughly one in ten could be removed with no measurable change in loss.

Next: Scaling laws below 100M parameters