sk Skill · AIops-tools
inference-aiops
Use this skill whenever the user needs to operate a GPU inference cluster — vLLM (OpenAI API + Prometheus /metrics) and Ray Serve / Ray Jobs (Ray dashboard), plus the single-process serving engines SGLang and TGI (Text Generation Inference): a one-shot cluster overview (deployments + total replicas + queue backpressure), request metrics (TTFT / TPOT / e2e latency + token totals), queue depth, KV-cache stats…
Open on skills.sh ↗read 2026-09-17
- installs 8w
- 0
- 30-day movement
- starts with the next reading
- Related entries
- 1
- Connections
- 1
Pythonbashinferencegovernancemcpaiops
- Host repository
- AIops-tools/Inference-AIops
- Allowed tools
- Bash
- Compatible with
- Standalone, self-governed GPU-inference operations. The governance harness (audit, policy, token/runaway budget, undo, not authorizes). The fragile prod ops — scale_replicas_down, scale_to_zero, drain_replica, lora_unload, model_sleep, replica_restart, model_undeploy
- Licence
- MIT
- Host language
- Python