BigHugger
sk Skill · whitecircle

optimize

Recommend throughput/memory levers to raise tokens/s/GPU or cut peak memory for a given Halo training config — and call out the levers that do NOT help at fine-grained MoE shapes (low-precision fp8/fp4 compute, native DeepGEMM, torch.compile on EP MoE, sub-bf16 master weights) so the user does not chase dead-ends. Use when the user asks how to speed up training, fit a longer seq / bigger batch, get more tok/s/GPU,…

installs 8w
0
30-day movement
starts with the next reading
Related entries
7
Connections
4
Python
Host repository
whitecircle/halo
Allowed tools
Read, Grep, Glob, Bash
Host stars
20
Host language
Python