All writing
Jun 19, 202616 min read

Sparse experts, dense lessons

8 experts
Parameters
7.4B total
Active
1.2B
Hardware
8x H100
Training time
12 days

lattice-8x1b routes each token to 2 of 8 experts. Getting the router to use all of them took three attempts.

Routing collapse

In my first run one expert received over half the tokens by step 5k. An auxiliary balancing loss with a small coefficient fixed it, and a little router noise early on kept it fixed.

Next: Training a 1.3B language model on a single node