sk Skill · ericrisco
vllm
Use when self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint, splitting a model across GPUs with tensor or pipeline parallelism, loading quantized weights, serving one or many LoRA adapters, and debugging KV-cache OOM from memory-utilisation and context-length flags. NOT renting or provisioning the GPU box (that is runpod or modal), NOT single-user…
Open on skills.sh ↗read 2026-09-17
- installs 8w
- 0
- 30-day movement
- starts with the next reading
- Related entries
- 2
- Connections
- 0
pythonbashtensor-parallelinference-serverJavaScriptllm-servingquantizationvllm
- Host repository
- ericrisco/rsc-harness
- Host stars
- 84
- Host language
- JavaScript