mew
An LLM training stack written from scratch, from tokenizer to 4D parallelism and GRPO
mew is a GPT-style language model and training stack that I am building from scratch, one component at a time. The goal is a stack that covers 4D-parallel pretraining (DP × TP × PP × CP, plus expert parallelism for MoE) and post-training (SFT and GRPO), where every piece is understood well enough to predict how it should perform before measuring it.
Principles
- Predict, then measure. Each component starts with a cost model (memory, communication volume, expected step time). The gap between prediction and measurement, and its explanation, is the deliverable.
- Correctness before speed. Every feature lands with a parity test against a single-GPU reference run.
- Compare against the industrial version. Each stage is benchmarked side by side with the PyTorch built-in or TorchTitan on the same hardware.
- Publish as I go. Each phase ends with a write-up on the blog.
What’s in mew Today
- Tokenization: a byte-pair encoding (BPE) tokenizer, trained and applied with multiprocessing.
- Model: Transformer blocks with RoPE and grouped-query attention, written on top of basic PyTorch ops.
- Kernels: a Triton FlashAttention (forward and backward) with bf16 AMP, checked against the eager reference.
- Training: a custom AdamW and LR schedule, memmap data loaders, Hydra configs and W&B tracking.
- Distributed (on the
feature/data_parallelbranch): aDistributedDataParallelwrapper with async per-parameter all-reduce andno_sync()for gradient accumulation. - Performance tooling: FLOPs-per-token accounting, MFU and tokens/s logging, CUDA memory snapshots, NVTX/Nsight Systems profiling, and an interactive memory report.
Roadmap
| Phase | Focus | Status | Write-up |
|---|---|---|---|
| 0 | Harness: launch, reference runs, comm logging | in progress | Building a Training Stack from Scratch |
| 1 | Fast DDP: bucketing, comm/compute overlap | planned | Why Naive DDP Is Slow |
| 2 | ZeRO and FSDP | planned | A Whiteboard Memory Model |
| 3 | Tensor and sequence parallelism | planned | Why TP Stays Inside NVLink |
| 4 | Pipeline parallelism | planned | Choosing a 3D Layout |
| 5 | Context parallelism (ring attention) | planned | Long-Context Training |
| 6 | MoE and expert parallelism | planned | Where MoE Training Time Goes |
| 7 | Checkpointing, fault tolerance, capstone | planned | Capstone vs TorchTitan |
| 8 | SFT | planned | What Packing Buys You |
| 9 | GRPO | planned | Where RL Time Goes |
The full plan, including the compute budget and reading list, lives in docs/ROADMAP.md.