BigHugger
sk Skill · lazyFrogLOL

llm-inference-batching-scheduler

Guidance for implementing batching schedulers for LLM inference systems with compilation-based accelerators. This skill applies when optimizing request batching to minimize cost while meeting latency thresholds, particularly when dealing with shape compilation costs, padding overhead, and multi-bucket request distributions. Use this skill for tasks involving batch planning, shape selection, generation-length…

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
0
Python
Host repository
lazyFrogLOL/Harness_Engineering
Host stars
128
Host language
Python