All writing
Sep 28, 202614 min read

Training a 1.3B language model on a single node

L = 24
Parameters
1.3B
Tokens
48B
Hardware
8x H100
Training time
9 days

halcyon-1.3b is a decoder-only transformer trained on 48B tokens on one 8-GPU node. This post covers the setup, the data mix, and the three loss spikes that cost me most of a week.

Setup

24 layers, width 2048, 16 heads, rotary embeddings, no biases. Everything ran in bf16 with a cosine schedule.

lr_peak       = 3e-4
batch_tokens  = 1_048_576
warmup_steps  = 2000

Loss spikes

All three spikes appeared in the first 10k steps. Lowering the Adam epsilon and clipping gradients at 0.5 removed them, at no visible cost to final loss.

Next: What attention heads actually attend to