BigHugger
sk Skill · AIops-tools

inference-aiops

Use this skill whenever the user needs to operate a GPU inference cluster — vLLM (OpenAI API + Prometheus /metrics) and Ray Serve / Ray Jobs (Ray dashboard), plus the single-process serving engines SGLang and TGI (Text Generation Inference): a one-shot cluster overview (deployments + total replicas + queue backpressure), request metrics (TTFT / TPOT / e2e latency + token totals), queue depth, KV-cache stats…

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
1
Pythonbashinferencegovernancemcpaiops
Host repository
AIops-tools/Inference-AIops
Allowed tools
Bash
Compatible with
Standalone, self-governed GPU-inference operations. The governance harness (audit, policy, token/runaway budget, undo, not authorizes). The fragile prod ops — scale_replicas_down, scale_to_zero, drain_replica, lora_unload, model_sleep, replica_restart, model_undeploy
Licence
MIT
Host language
Python