soc-siem
Build and run a security operations center the way a detection engineer does. The stack is open-source and three-part: Wazuh (host-based IDS, log analysis, file integrity monitoring, vulnerability detection; agents reporting to a manager and indexer), Suricata (network IDS/IPS on mirrored or tapped traffic, emitting EVE JSON), and Grafana (dashboards over the security data, the SOC's single pane of glass). The hard part is the pipeline collect → normalize → detect → alert → triage → respond, tuned so the signal survives the noise. The dominant failure mode is alert fatigue: a SOC that pages on everything trains its analysts to stop reading.
This skill is advisory and detection-only by default. An active response (Wazuh active-response quarantining a host, Suricata flipped from IDS to inline IPS, a firewall rule, a key revocation) is an external mutation that runs only behind a reviewed change and explicit, recorded approval, never freehand from a step here.
Climb the determinism ladder: express a detection as a config-as-code rule, a decoder, or a Sigma-to-detection compile before you write it as prose, and turn a checklist into a rule-test fixture or a lint gate. Judgment takes the last rung.
Steps
-
State the estate and the monitoring bar. Write the assets in scope (cloud accounts, VMs, containers, the network segments), the data classification each carries, the threat model (the named adversary behaviors that matter), and the compliance regime driving the program (SOC2 continuous monitoring is the common one). The scope recorded here is what every later sensor placement and detection is measured against. This step is done once the in-scope assets, their data classification, and the compliance regime are written down.
-
Lay the stack topology for cloud and VMs. Place the Wazuh manager and indexer, the agent fleet across VMs and containers, the Suricata sensors on the traffic you can mirror or tap, and the cloud log sources (CloudTrail, VPC flow logs, the equivalents) feeding the pipeline, per the stack-architecture reference. Scope every collector to least privilege: a read-only log-pull role, never an admin key. This step is done once each in-scope asset names its collector, each network segment names its sensor or records why it has none, and no collector credential carries a wildcard.
-
Wire the pipeline end to end. Connect collect → normalize → detect → alert → triage → respond so an event leaves an agent or sensor, is decoded into a normalized field set, is scored against the ruleset, and reaches an alert channel with an owner, per the stack-architecture reference. This step is done once a test event injected at a Wazuh agent and a test alert raised in Suricata both arrive in Grafana with their normalized fields intact.
-
Engineer detections and map them to ATT&CK. Write and tune the rules (Wazuh decoders and rules, Suricata signatures, Sigma compiled to the backend) and tag each detection with its MITRE ATT&CK technique so coverage and gaps are visible, per the detection-engineering reference. Run the fire-and-silence test through the rule engine itself (
suricata -Tto validate a signature set,wazuh-logtestto replay a sample through the decoders and rules), so the fire-on-malicious and silent-on-benign verdict comes from the engine rather than from reading the rule. A detection ships only on a passing engine pair. This step is done once every shipped rule carries an ATT&CK technique id and a passing fire-and-silence test pair. -
Tune against false positives before going live. Baseline each noisy rule against a window of normal traffic, then suppress, threshold, or scope it until its precision clears the bar set in the detection-engineering tuning section. An alert stream nobody trusts is the core failure, so a rule that cannot beat its noise is disabled, not merged. This step is done once each enabled rule has a recorded precision over the baseline window and no rule sits above the agreed false-positive ceiling.
-
Build the triage workflow and the dashboards. Define the severity tiers, the triage runbook (the steps an analyst takes from alert to verdict), the escalation path, and the Grafana panels that surface the live picture and the SOC2 evidence trail, per the operations reference. This step is done once each severity tier names its response-time target and its owner, and the SOC2 continuous-monitoring panel renders the evidence an auditor reads.