nvidia-cuda-performance
NVIDIA CUDA performance
Shorten the declared result path by changing measured work, movement, placement, scheduling, or representation—and prove the effect outside the profiler.
Scope
Use this skill for NVIDIA CUDA kernels and CUDA-backed applications, including RTX 3090/SM86 specialization, LLM inference, and intra-host multi-GPU execution.
Route a CPU-only or ordinary single-binary bottleneck to $perform-like-jeff-and-sanjay. Route multi-node networking, storage fabrics, and general ML training to domain-specific methods. Keep host inventories, measurements, and winning configurations in the target project rather than this portable skill.
Treat every optimization as a gated move. A technique becomes a candidate only after evidence supports its precondition. Counters such as utilization, occupancy, stalls, cache hit rate, queue time, roofline position, and NCCL bandwidth are clues—not diagnoses.
Use these evidence labels:
- SOURCE GAP: a stable technical claim lacks adequate primary evidence.
- INSTANCE VALUE: the answer depends on the target host, build, artifact, or runtime dispatch.
- EMPIRICAL OPTIMUM: the winner depends on the measured workload.