production-audit
Production Audit
Using this skill: announce "Using production-audit", make a todo per phase in
## Phasesplus one per applicable dimension Phase 0 resolves, and do not skip the gates. This skill's worth is its process, not a hand-reproduced outcome. If you were told to "run production-audit", run it, do not improvise its result. (Suite standard: https://github.com/horizon-foundry/foundry/blob/main/reference/skill-authoring.md)
Overview
A whole-application audit that ends in a verdict: safe to ship; ready to ship, risks noted; or do not ship. Core principle: evidence or it doesn't ship. Every finding cites code, and every Critical, High, and verdict-driving finding survives an adversarial refutation attempt or gets downgraded. The audit never modifies code.
Every finding carries a kind: risk or improvement. A risk has a path to harm in production (exposure, data or payment corruption, a crash or unavailability, abuse, or a hidden active failure). An improvement is safe today but makes the system more robust, observable, consistent, accessible, or faster. The distinction is load-bearing: only risks drive the verdict; an improvement never caps it. A report with no risks is clear within its assessed scope even if it lists a dozen improvements, because a punch list of betterments is not a reason to hold a release. This is what keeps the report honest and readable: risks to weigh first, improvements to schedule second, never one undifferentiated wall of problems. Think of it as bugs versus tech debt.
Classify kind with two questions, in order (in the suite repo this is the field make validate gates on, so decide it deliberately):
| Question | Answer | kind |
|---|---|---|
| 1. Is severity low or informational? | Yes | improvement, always (minor enough is tech debt or polish by definition: a low rate-limit gap mitigated upstream, a rare recoverable race, an accessibility refinement) |
| 2. Severity is medium or higher: is there a harm path reachable today (live exposure, corruption path, crash or availability loss, abuse, hidden active failure)? | Yes | risk (the shipper must consciously accept it) |
| No (missing tests, missing CI, observability thinness, robustness the app lacks but does not bleed from) | improvement |
The verdict states what evidence it rests on. It carries an assessedScope, and a static-only run names that base beneath the recommendation ("static review, runtime not exercised") and lists the checks it skipped under notAssessed, so it reads as a human recommendation without overclaiming a runtime-verified result. A verdict may not claim more than the run could see.