BigHugger
sk Skill · markoblogo

local-inference-tuning

Select and tune a local LLM inference engine for the user's hardware. Use when setting up or auditing local/private model serving, choosing between MLX, llama.cpp, Ollama, and vLLM, estimating model fit, deciding cache/storage policy, tuning batching and KV cache flags, running smoke benchmarks, or exposing an OpenAI-compatible local endpoint.

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
0
bashPython
Host repository
markoblogo/abvx-agent-skills
Licence
MIT
Host stars
16
Host language
Python