BigHugger
sk Skill · aws-samples

vllm-setup

Stand up a vLLM inference server for an open-weight coding model on a multi-GPU EC2 node (reference: g6e.12xlarge, 4xL40S). Drives the full flow end to end — verify the GPU node, install vLLM and its OS/Python dependencies (including the two Deep Learning AMI-specific fixes), serve a model with tensor parallelism and tool calling, confirm inference works, and optionally install opencode to drive it as a coding…

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
0
bashShell
Host repository
aws-samples/sample-claude-code-multi-model
Version
1.0
Licence
Apache-2.0
Host stars
13
Host language
Shell