BigHugger
sk Skill · ajyadav013

llm-inference-optimization

Self-hosted LLM serving performance — quantization (AWQ/GPTQ/FP8/INT4), speculative decoding, prefill vs decode, KV cache and prefix caching. Use when sizing GPUs, cutting inference cost, or diagnosing TTFT versus inter-token latency.

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
0
Python
Host repository
ajyadav013/claude-kit
Host stars
12
Host language
Python