sk Skill · ajyadav013
llm-inference-optimization
Self-hosted LLM serving performance — quantization (AWQ/GPTQ/FP8/INT4), speculative decoding, prefill vs decode, KV cache and prefix caching. Use when sizing GPUs, cutting inference cost, or diagnosing TTFT versus inter-token latency.
Open on skills.sh ↗read 2026-09-15
- installs 8w
- 0
- 30-day movement
- starts with the next reading
- Related entries
- 1
- Connections
- 0
Python
- Host repository
- ajyadav013/claude-kit
- Host stars
- 12
- Host language
- Python