Jun 19, 202616 min read
Sparse experts, dense lessons
- Parameters
- 7.4B total
- Active
- 1.2B
- Hardware
- 8x H100
- Training time
- 12 days
lattice-8x1b routes each token to 2 of 8 experts. Getting the router to use all of them took three attempts.
Routing collapse
In my first run one expert received over half the tokens by step 5k. An auxiliary balancing loss with a small coefficient fixed it, and a little router noise early on kept it fixed.