Machine learning researcher focused on training custom models and developing efficient architectures for practical machine learning systems.
Data mix, learning rate schedule, and the three loss spikes that cost me a week.
A look at tern-370m: which heads track position, which copy tokens, and which do nothing.
Fitting loss curves on tiny models, and where the power law starts to bend.
Routing collapse, load balancing, and what I changed for lattice-8x1b.
| Name | Parameters | Training tokens | Released | Weights |
|---|---|---|---|---|
| halcyon-1.3b | 1.3B | 48B | Sep 2026 | Hugging Face |
| tern-370m | 370M | 20B | Aug 2026 | Hugging Face |
| lattice-8x1b | 7.4B total, 1.2B active | 120B | Jun 2026 | Hugging Face |
| mica-60m | 60M | 6B | Mar 2026 | Hugging Face |
I'm a machine learning researcher who trains custom models and studies the architectural choices that make them more efficient, capable, and practical to use.
My work sits at the intersection of model development and research: designing experiments, training models from the ground up, and learning from both successful and unsuccessful runs.
I’m particularly interested in efficient architectures that reduce the computational cost of machine learning while preserving the capabilities that matter in real-world applications.