BigHugger
sk Skill · omer-metin

distributed-training

Use when training models across multiple GPUs or nodes, handling large models that don't fit in memory, or optimizing training throughput — covers DDP, FSDP, DeepSpeed ZeRO, model/data parallelism, and gradient checkpointingUse when ", " mentioned.

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
0
Python
Host repository
omer-metin/skills-for-antigravity
Host stars
152
Host language
Python