GH Repository · lemonade-sdk
lemonade
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk
- stars
- 5,727
- 30-day movement
- +3813/day
- Related entries
- 60
- Connections
- 4
dockerC++rocmgenaiairyzenvulkanmcpllmamdllm-inferencellamagpulocal-serverradeononnxruntimeqwennpumcp-servermistralopenai-api
Lemonade is an SDK and local server that serves optimized LLMs from your own GPUs and NPUs, exposing them through an OpenAI-compatible API and an MCP server. It targets AMD hardware (Radeon, Ryzen, ROCm) with backends like ONNX Runtime and Vulkan.
You want to run LLM inference locally on AMD GPU/NPU hardware and plug it into agents via MCP.
Use it to
- Serve local LLMs behind an OpenAI-compatible API
- Expose local models to agents via MCP
- Run inference on AMD Radeon GPUs and Ryzen NPUs
- Discover and run local AI apps
For Developers running LLMs locally on AMD hardware
- Role
- mcp-server
- Language
- C++
- Licence
- Apache-2.0
- Forks
- 501
- Open issues
- 404
- Last push
- 2026-09-17
- Latest release
- v7.0.0 · 2025-05-16
topicsllm-inferencelocal-servermcp-serveramdgpunpu