BigHugger
sk Skill · rlaope

omh-inference-serving

[omh] OMH Inference Serving workflow: choose the serving engine and quantization from decision tables, prepare deployment as an idempotent runbook with observed-only verification, and measure the endpoint with the standard TTFT/TPOT/goodput protocol. Use when the user says: inference-serving, inference serving, serve this model, serve the model, model serving, serving endpoint, vllm, llama.cpp.

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
0
bashPython
Host repository
rlaope/oh-my-hermes
Host stars
2,778
Host language
Python