Aug 11, 202611 min read
Scaling laws below 100M parameters
- Parameters
- 1M to 60M
- Tokens
- 0.5B to 6B
- Hardware
- 1x A100
- Training time
- 40 runs
Scaling laws are usually fit on large models. I trained forty small ones to see whether the same curves hold at the low end.
Where it bends
Below roughly 5M parameters the fit drifts from a clean power law. Embedding size starts to dominate, and the usual compute-optimal ratios stop applying.