GH Repository · bentoml
OpenLLM
Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.
- stars
- 12,535
- 30-day movement
- +41/day
- Related entries
- 60
- Connections
- 1
otherPythonmlopsmistralpythonllmllama3-2-visionllmopsllamapython/uvllm-servingvicunafine-tuningopen-source-llmbentomlllama3-2model-inferencellm-inferencellama2llm-opsllama3-1openllm
OpenLLM is a Python tool from the BentoML project for serving open-source LLMs such as DeepSeek, Llama, Mistral, and Vicuna as OpenAI-compatible API endpoints in the cloud. It is built on the BentoML stack and covers model inference, serving, and fine-tuning workflows.
Use it when you want to self-host open-source LLMs behind an OpenAI-compatible API without changing client code.
Use it to
- Serve open-source LLMs as OpenAI-compatible endpoints
- Deploy LLM inference in the cloud via BentoML
- Fine-tune supported open-source models
- Swap hosted models into existing OpenAI-style clients
- Standardize LLM serving in MLOps pipelines
For ML engineers and teams deploying open-source LLMs
- Role
- other
- Language
- Python
- Licence
- Apache-2.0
- Forks
- 841
- Open issues
- 6
- Last push
- 2026-09-14
- Latest release
- v0.0.21 · 2023-06-04
topicsllm-servingmodel-inferenceopenai-compatiblebentomlfine-tuningllmops