serving-llms-vllm
Pass
Audited by Gen Agent Trust Hub on Oct 1, 2026
Risk Level: SAFEEXTERNAL_DOWNLOADSDYNAMIC_EXECUTIONINDIRECT_PROMPT_INJECTIONPRIVILEGE_ESCALATION
Full Analysis
- [EXTERNAL_DOWNLOADS]: The skill instructs users to install several machine learning and benchmarking packages including
vllm,torch,transformers,locust,autoawq,auto-gptq, andflash-attnfrom established package registries. It also references fetching pre-trained models from trusted organizations on Hugging Face, such asmeta-llamaandTinyLlama. - [DYNAMIC_EXECUTION]: The documentation describes the use of the
--trust-remote-codeflag when serving models. This feature allows vLLM to execute custom Python code provided within model repositories (e.g., for specialized model architectures). While a standard practice in the ML ecosystem, it represents a dynamic execution surface that requires the model source to be verified. - [INDIRECT_PROMPT_INJECTION]: The skill sets up an inference server that processes arbitrary text prompts, which is a known attack surface for indirect prompt injection.
- Ingestion points: Untrusted data is ingested via an OpenAI-compatible API server and from batch input files like
prompts.txt(SKILL.md). - Boundary markers: Present; the server configuration uses
SamplingParamswith definedstoptokens to manage the termination of generated output. - Capability inventory: The skill enables model inference, serving, and system monitoring functionalities (
SKILL.md). - Sanitization: Absent; no specific prompt sanitization or safety filtering logic is included in the deployment instructions beyond the model's native guardrails.
- [PRIVILEGE_ESCALATION]: The troubleshooting guide suggests using
sudo ufw allow 8000to resolve connectivity issues. This involves using administrative privileges to modify system firewall rules, which is a standard administrative task for server configuration (references/troubleshooting.md).
Audit Metadata