llm-deployment

Installation
SKILL.md

Overview

Running LLMs in production — local inference with Ollama/llama.cpp to high-throughput serving with vLLM/TGI. Quantization, GPU optimization, OpenAI-compatible APIs.

Capabilities

  • Local deployment (Ollama, llama.cpp, LM Studio)
  • High-throughput serving (vLLM, TGI)
  • Quantization (GGUF, GPTQ, AWQ)
  • GPU memory optimization (FlashAttention, PagedAttention)
  • OpenAI-compatible API endpoints

When to Use

  • Self-hosted LLM for privacy/cost
  • High-throughput API serving
  • Running on consumer GPUs (24GB or less)

When NOT to Use

Installs
2
GitHub Stars
12
First Seen
Jun 20, 2026
llm-deployment — oyi77/1ai-skills