BigHugger
GH Repository · bentoml

OpenLLM

Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.

stars
12,535
30-day movement
+41/day
Related entries
60
Connections
1
otherPythonmlopsmistralpythonllmllama3-2-visionllmopsllamapython/uvllm-servingvicunafine-tuningopen-source-llmbentomlllama3-2model-inferencellm-inferencellama2llm-opsllama3-1openllm

OpenLLM is a Python tool from the BentoML project for serving open-source LLMs such as DeepSeek, Llama, Mistral, and Vicuna as OpenAI-compatible API endpoints in the cloud. It is built on the BentoML stack and covers model inference, serving, and fine-tuning workflows.

Use it when you want to self-host open-source LLMs behind an OpenAI-compatible API without changing client code.

Use it to

  • Serve open-source LLMs as OpenAI-compatible endpoints
  • Deploy LLM inference in the cloud via BentoML
  • Fine-tune supported open-source models
  • Swap hosted models into existing OpenAI-style clients
  • Standardize LLM serving in MLOps pipelines

For ML engineers and teams deploying open-source LLMs

Role
other
Language
Python
Licence
Apache-2.0
Forks
841
Open issues
6
Last push
2026-09-14
Latest release
v0.0.21 · 2023-06-04
topicsllm-servingmodel-inferenceopenai-compatiblebentomlfine-tuningllmops