BigHugger
GH Repository · lemonade-sdk

lemonade

Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk

stars
5,727
30-day movement
+3813/day
Related entries
60
Connections
4
dockerC++rocmgenaiairyzenvulkanmcpllmamdllm-inferencellamagpulocal-serverradeononnxruntimeqwennpumcp-servermistralopenai-api

Lemonade is an SDK and local server that serves optimized LLMs from your own GPUs and NPUs, exposing them through an OpenAI-compatible API and an MCP server. It targets AMD hardware (Radeon, Ryzen, ROCm) with backends like ONNX Runtime and Vulkan.

You want to run LLM inference locally on AMD GPU/NPU hardware and plug it into agents via MCP.

Use it to

  • Serve local LLMs behind an OpenAI-compatible API
  • Expose local models to agents via MCP
  • Run inference on AMD Radeon GPUs and Ryzen NPUs
  • Discover and run local AI apps

For Developers running LLMs locally on AMD hardware

Role
mcp-server
Language
C++
Licence
Apache-2.0
Forks
501
Open issues
404
Last push
2026-09-17
Latest release
v7.0.0 · 2025-05-16
topicsllm-inferencelocal-servermcp-serveramdgpunpu