systems-profiling
Installation
SKILL.md
Systems Performance Profiling Skill
Use this skill to investigate CPU time, call paths, cache behavior, memory errors, heap growth, and frame or scheduling latency in native programs. It covers the systems tools listed in this repository's root README: gprof, Linux perf, Tracy, Valgrind Memcheck/Cachegrind/Callgrind/Massif, and Apple Instruments.
Start with a reproducible baseline
- Define one question and one workload: for example, “which call path consumes CPU during level loading?” or “which allocation remains live after ten requests?” Do not mix startup, warm-up, and steady-state measurements.
- Build a profiling binary with symbols and a known build identity. For GCC/Clang, a useful baseline is
-g -O2; keep frame pointers (-fno-omit-frame-pointer) when stack quality matters and disable or account for LTO/inlining when symbol attribution is more important than production fidelity. Use the same optimization mode as the target when validating a fix. - Warm up once, then capture several identical runs. Record OS/kernel, CPU, compiler and linker versions, flags, input data, tool version, and whether the run was under a debugger, VM, container, or emulator.
- Prefer a release-like build with debug information over an unoptimized debug build. Optimization changes inlining, code layout, locking, and cache behavior; an unoptimized binary can answer “where is work?” but not reliably “what will production do?”
- Keep the exact executable, shared libraries, dSYM/DWARF files, and source revision that produced a capture. A profile without matching symbols can still show addresses, but function names, source lines, and inline frames may be wrong or missing.
A profile is evidence, not a diagnosis. First compare total runtime, CPU utilization, allocations, cache misses, or frame time with an unprofiled baseline. Then inspect the hottest or most rapidly growing paths, change one thing, and repeat the same workload.
Sampling versus instrumentation
- Sampling periodically interrupts or observes a running process (Linux perf's normal
record, Time Profiler, and gprof's timer component). It usually has lower overhead and better timing fidelity, but it can miss very short functions, under-sample blocked threads, and depends on usable unwind information. A percentage is an estimate over samples, not an exact count of calls. - Instrumentation/tracing inserts hooks or observes individual events (gprof's
-pgcall hooks, Tracy zones, Valgrind's dynamic translation, and Instruments Allocations/Leaks templates). It exposes call counts, lifetimes, and event order, but can add substantial CPU, memory, synchronization, and I/O overhead. Use it for causality and memory correctness, not absolute latency comparisons. - Hybrid tools should be read as such: gprof uses instrumented call edges and timer samples; Instruments has sampled and event-based templates; perf can also collect tracepoints or hardware events in addition to sampling. State which mode produced a result.