All writing
Aug 11, 202611 min read

Scaling laws below 100M parameters

x10
Parameters
1M to 60M
Tokens
0.5B to 6B
Hardware
1x A100
Training time
40 runs

Scaling laws are usually fit on large models. I trained forty small ones to see whether the same curves hold at the low end.

Where it bends

Below roughly 5M parameters the fit drifts from a clean power law. Embedding size starts to dominate, and the usual compute-optimal ratios stop applying.

Next: Sparse experts, dense lessons