gcp-toolkit
Design and operate Google Cloud the way a GCP architect does. The default move is serverless-first, then managed-first: prefer the service that scales to zero and removes undifferentiated work, so owning a VM becomes the last resort. The hard part is choosing the most managed option that still meets the constraint, scoping the IAM binding, placing the project so the blast radius stays small, and proving the control before an auditor or an attacker finds it.
This skill is advisory and authors IaC by default. A provision against a live project is an external mutation that runs only behind a reviewed plan and explicit, recorded approval, never freehand from a step here.
Climb the determinism ladder: express a rule as a gcloud command, an org policy constraint, or a policy-as-code rule before you write it as prose, and turn a checklist into a lint or terraform validate gate. Judgment takes the last rung.
Steps
-
State the workload and its bar. Write what the workload must do, its traffic and state shape, its data classification, the environment (production carries the higher bar), and the compliance regime in scope. The constraints recorded here are the bar that every later managed-service choice is measured against. This step is done once the workload, its data classification, and its environment are written down.
-
Choose the most managed service that meets the constraint. Walk the compute and data decisions through the managed-first reference: Cloud Run before GKE Autopilot before Compute Engine for compute, Cloud Functions for event glue; Firestore or Cloud SQL before Spanner, BigQuery for analytics. Descend to a self-run VM only against a recorded constraint from step 1. This step is done once each compute and data component names its chosen service, and each self-managed choice names the constraint that forced it.
-
Apply the five Architecture Framework categories. Take a named stance on operational excellence, security and privacy and compliance, reliability, cost optimization, and performance optimization, per the Architecture Framework reference. A category with no stance counts as a gap. This step is done once each of the five categories carries a written stance for this workload.
-
Place the project and lay the identity and security baseline. Site the workload in its own project under the right folder, grant access through service accounts bound to predefined or custom roles via Workload Identity, encrypt at rest with CMEK and in transit with TLS, close the VPC, and enable Cloud Audit Logs, per the IAM and SRE reference. This step is done once no binding grants Owner or Editor, no JSON service-account key is introduced where Workload Identity reaches, every data store names its CMEK key, and Data Access logs are on for the project.
-
Set the SLO and error budget. Define a service-level objective for the workload's critical journey and derive the error budget from it, then wire the alerting that fires before the budget is spent, per the SRE section. A target nobody measures counts as a reliability gap. This step is done once the SLO, its error budget, and a burn-rate alert are written down.
-
Author as IaC and attach cost controls. Express every resource in Terraform (or Config Connector) with pinned providers and remote, locked state, and keep the console out of the change path: click-ops leaves no diff for a reviewer to read. Set a billing budget with an alert and apply the mandatory label set, then pull the cost levers (committed-use discounts on steady load, the Recommender's right-sizing, autoscaling to zero) named in the cost section. This step is done once
terraform validateandterraform fmt -checkboth pass, the budget alert exists, and every billable resource carries the required labels.