BigHugger
sk Skill · ericrisco

vllm

Use when self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint, splitting a model across GPUs with tensor or pipeline parallelism, loading quantized weights, serving one or many LoRA adapters, and debugging KV-cache OOM from memory-utilisation and context-length flags. NOT renting or provisioning the GPU box (that is runpod or modal), NOT single-user…

installs 8w
0
30-day movement
starts with the next reading
Related entries
2
Connections
0
pythonbashtensor-parallelinference-serverJavaScriptllm-servingquantizationvllm
Host repository
ericrisco/rsc-harness
Host stars
84
Host language
JavaScript