TEMPLAR
Run
Now Running
TEST_RUN_003
--model=llama-8b--arch=dense
Pre-training with SparseLoCo and low-bandwidth pipelining. Read the paper →
Training Progress
Is the model learning, and how fast
Training Loss
Tokens Processed
MFU — compute vs effective
Network & Topology
Stage chains, replica sync and round health