Sep 28, 202614 min read
Training a 1.3B language model on a single node
- Parameters
- 1.3B
- Tokens
- 48B
- Hardware
- 8x H100
- Training time
- 9 days
halcyon-1.3b is a decoder-only transformer trained on 48B tokens on one 8-GPU node. This post covers the setup, the data mix, and the three loss spikes that cost me most of a week.
Setup
24 layers, width 2048, 16 heads, rotary embeddings, no biases. Everything ran in bf16 with a cosine schedule.
lr_peak = 3e-4
batch_tokens = 1_048_576
warmup_steps = 2000
Loss spikes
All three spikes appeared in the first 10k steps. Lowering the Adam epsilon and clipping gradients at 0.5 removed them, at no visible cost to final loss.