sk Skill · whitecircle
optimize
Recommend throughput/memory levers to raise tokens/s/GPU or cut peak memory for a given Halo training config — and call out the levers that do NOT help at fine-grained MoE shapes (low-precision fp8/fp4 compute, native DeepGEMM, torch.compile on EP MoE, sub-bf16 master weights) so the user does not chase dead-ends. Use when the user asks how to speed up training, fit a longer seq / bigger batch, get more tok/s/GPU,…
Open on skills.sh ↗read 2026-09-17
- installs 8w
- 0
- 30-day movement
- starts with the next reading
- Related entries
- 7
- Connections
- 4
Python
- Host repository
- whitecircle/halo
- Allowed tools
- Read, Grep, Glob, Bash
- Host stars
- 20
- Host language
- Python