Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Deploy and serve LLMs with vLLM behind an OpenAI-compatible endpoint, with tool calling enabled for agent workloads.
.claude/skills/prism-shadow-vllm/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-10 | ✗→✓ | ▲ Improved | -10% | 0% |
| case-12 | ✗→✓ | ▲ Improved | -9% | 0% |
| case-20 | ✗→✓ | ▲ Improved | -5% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 25% | 0% |
| case-11 | ✓→✗ | ▼ Worse | -11% | 0% |
vLLM serves open-weight LLMs on local GPUs with high-throughput inference behind an OpenAI-compatible API, ready for chat and agent workloads.
If the user's message only invokes this skill (e.g. "use vllm skill") without a concrete request, ask the user what they want. Do not run any command until the goal is clear.
Ask the user which model to serve; if they have no preference, recommend the small default Qwen/Qwen3.5-0.8B. Also ask what context length the workload needs.
vLLM needs an NVIDIA or AMD GPU. Engine choice follows the user's preference: Ollama also runs on GPUs and is the simpler default — pick vLLM for high-throughput serving, and Ollama on macOS or CPU-only machines, which vLLM does not serve. Confirm the hardware first:
bashnvidia-smi # NVIDIA: GPU model and free VRAM (AMD ROCm: rocm-smi) python3 --version # a recent Python is required
The model must fit the available VRAM — model size and context length drive the serve flags below.
curl http://localhost:8000/v1/models.penguin config model add ... --client-type openai-chat --base-url http://localhost:8000/v1 — a served model is not visible to Penguin until added.penguin config model list.Use a fresh virtual environment (or uv):
bashpython3 -m venv .venv && source .venv/bin/activate pip install vllm
bashvllm serve Qwen/Qwen3.5-0.8B --port 8000
This exposes an OpenAI-compatible API at http://localhost:8000/v1. Key flags:
--served-model-name <name> — the model id clients request (defaults to the model path).--api-key <key> — require this bearer token on every request.--max-model-len <n> — context window; agent sessions need a large one.--gpu-memory-utilization <0..1> — fraction of VRAM to claim (default 0.9).--tensor-parallel-size <n> — shard across n GPUs.--dtype <auto|bfloat16|float16> and --quantization <awq|gptq|fp8> — precision and quantized weights.If the port is taken, pick a free one — never kill a process already listening on it.
Agent harnesses (PenguinHarness included) send tools with their requests. vLLM must opt in at startup:
bashvllm serve Qwen/Qwen3.5-0.8B --enable-auto-tool-choice --tool-call-parser hermes
Choose the parser for the model family — e.g. hermes for Qwen models, llama3_json for Llama models. Without these flags, requests that set tool_choice fail with 400 "auto" tool choice requires --enable-auto-tool-choice and --tool-call-parser to be set.
bashcurl http://localhost:8000/v1/models
Model configuration is the penguin CLI's job — penguin config model add registers an endpoint and penguin config model list shows what has been registered. A served model is not visible to Penguin until you add it:
bashpenguin config model add --provider custom --client-type openai-chat \ --base-url http://localhost:8000/v1 --model-id <served-model-name> --api-key <key> penguin config model list # the new entry should now be listed
--gpu-memory-utilization or --max-model-len, or serve a quantized model.--max-model-len (bounded by VRAM).400 on tool calls: restart the server with the tool-calling flags above.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-10 | fail→pass | 16,192 | 9,251 | -43% | 1 | 1 | 0% | 2,015 | 1,812 | -10% | 0 | 0 | — |
case-11 | pass→fail | 14,752 | 8,802 | -40% | 1 | 1 | 0% | 1,843 | 1,648 | -11% | 0 | 0 | — |
case-12 | fail→pass | 13,207 | 4,053 | -69% | 1 | 1 | 0% | 2,040 | 1,857 | -9% | 0 | 0 | — |
case-13 | pass→pass | 13,534 | 8,116 | -40% | 1 | 1 | 0% | 1,549 | 1,505 | -3% | 0 | 0 | — |
case-14 | pass→pass | 12,742 | 7,394 | -42% | 1 | 1 | 0% | 1,327 | 1,497 | +13% | 0 | 0 | — |
case-15 | pass→pass | 10,488 | 1,854 | -82% | 1 | 1 | 0% | 1,100 | 1,460 | +33% | 0 | 0 | — |
case-16 | pass→pass | 9,772 | 2,554 | -74% | 1 | 1 | 0% | 863 | 1,627 | +89% | 0 | 0 | — |
case-17 | pass→pass | 6,433 | 3,061 | -52% | 1 | 1 | 0% | 1,184 | 1,633 | +38% | 0 | 0 | — |
case-18 | pass→pass | 3,686 | 6,472 | +76% | 1 | 1 | 0% | 669 | 1,373 | +105% | 0 | 0 | — |
case-19 | pass→pass | 12,140 | 8,214 | -32% | 1 | 1 | 0% | 2,170 | 1,716 | -21% | 0 | 0 | — |
case-20 | fail→pass | 18,921 | 12,950 | -32% | 1 | 1 | 0% | 3,736 | 3,558 | -5% | 0 | 0 | — |
case-21 | pass→fail | 19,099 | 6,782 | -64% | 1 | 1 | 0% | 2,474 | 2,295 | -7% | 0 | 0 | — |
case-01 | fail→fail | 10,895 | 8,236 | -24% | 1 | 1 | 0% | 367 | 1,493 | +307% | 0 | 0 | — |
case-02 | fail→fail | 17,172 | 16,354 | -5% | 1 | 1 | 0% | 380 | 1,501 | +295% | 0 | 0 | — |
case-03 | fail→fail | 10,548 | 12,076 | +14% | 1 | 1 | 0% | 246 | 1,549 | +530% | 0 | 0 | — |
case-04 | pass→pass | 9,069 | 7,826 | -14% | 1 | 1 | 0% | 1,681 | 1,526 | -9% | 0 | 0 | — |
case-05 | pass→pass | 24,520 | 11,716 | -52% | 1 | 1 | 0% | 2,354 | 2,451 | +4% | 0 | 0 | — |
case-06 | pass→pass | 16,877 | 4,591 | -73% | 1 | 1 | 0% | 2,166 | 2,070 | -4% | 0 | 0 | — |
case-07 | fail→pass | 8,544 | 8,758 | +3% | 1 | 1 | 0% | 1,475 | 1,839 | +25% | 0 | 0 | — |
case-08 | pass→pass | 21,267 | 8,307 | -61% | 1 | 1 | 0% | 2,789 | 2,675 | -4% | 0 | 0 | — |
case-09 | pass→pass | 17,294 | 6,147 | -64% | 1 | 1 | 0% | 2,206 | 2,333 | +6% | 0 | 0 | — |
case-22 | pass→pass | 18,803 | 11,405 | -39% | 1 | 1 | 0% | 2,017 | 2,038 | +1% | 0 | 0 | — |
case-23 | pass→pass | 3,806 | 7,542 | +98% | 1 | 1 | 0% | 744 | 1,498 | +101% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 20 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +9 percentage points is the difference between those two pass rates over the 20 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.