BigHugger
sk Skill · lebsral

dspy-vllm

Use vLLM for high-throughput production serving of self-hosted models with DSPy via dspy.LM with openai/ prefix and api_base. Use when you want production LLM serving, tensor parallelism, multi-GPU inference, batch processing, or high-concurrency self-hosted models. Also used for vllm, vLLM, production serving, high throughput LLM, tensor parallelism, self-hosted production, PagedAttention, local production server,…

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
0
pythondockerbashPython
Host repository
lebsral/DSPy-Programming-not-prompting-LMs-skills
Host stars
11
Host language
Python