nvidia-cuda-performance

Installation
SKILL.md

NVIDIA CUDA performance

Shorten the declared result path by changing measured work, movement, placement, scheduling, or representation—and prove the effect outside the profiler.

Scope

Use this skill for NVIDIA CUDA kernels and CUDA-backed applications, including RTX 3090/SM86 specialization, LLM inference, and intra-host multi-GPU execution.

Route a CPU-only or ordinary single-binary bottleneck to $perform-like-jeff-and-sanjay. Route multi-node networking, storage fabrics, and general ML training to domain-specific methods. Keep host inventories, measurements, and winning configurations in the target project rather than this portable skill.

Treat every optimization as a gated move. A technique becomes a candidate only after evidence supports its precondition. Counters such as utilization, occupancy, stalls, cache hit rate, queue time, roofline position, and NCCL bandwidth are clues—not diagnoses.

Use these evidence labels:

  • SOURCE GAP: a stable technical claim lacks adequate primary evidence.
  • INSTANCE VALUE: the answer depends on the target host, build, artifact, or runtime dispatch.
  • EMPIRICAL OPTIMUM: the winner depends on the measured workload.

Process

Installs
1
Repository
whamp/skills
GitHub Stars
1
First Seen
Aug 12, 2026
nvidia-cuda-performance — whamp/skills