vllm-ascend

Installation
SKILL.md

vLLM-Ascend - LLM Inference Serving

vLLM-Ascend is a plugin for vLLM that enables efficient LLM inference on Huawei Ascend AI processors. It provides Ascend-optimized kernels, quantization support, and distributed inference capabilities.


Quick Start

Offline Batch Inference

import os

# Required for vLLM-Ascend: set multiprocessing method before importing vLLM
os.environ["VLLM_WORKER_MULTIPROC_METHOD"] = "spawn"

from vllm import LLM, SamplingParams
Installs
62
GitHub Stars
143
First Seen
Mar 1, 2026
vllm-ascend — ascend-ai-coding/awesome-ascend-skills