Learning outcomes
- Establish the alert's scope, urgency, evidence quality, and protected consequence.
- Reconstruct identity, authorization, route, application, and control-plane activity on one timeline.
- Select a proportionate containment or escalation action and preserve the evidence that supports it.
- Feed investigation results back into detection logic, telemetry, and response procedures.
Operating objective
Triage must decide whether an alert represents a credible harmful condition, how quickly the team must act, which assets and authority may be affected, and what evidence or containment comes next. Investigation must support that decision with reproducible facts. It must not become an open-ended search for everything that could have happened.
Start with the alert contract: threat hypothesis, subject, resource, action, time window, detection version, contributing signals, confidence, severity, owner, and expected response. Protect severe alerts from automatic downgrade when a required source is missing. A telemetry failure can increase uncertainty and urgency.
Signals and evidence
Preserve the original alert and query results. Record source identifiers, event times and ingestion times, time zones, clock quality, query text, filters, result count, analyst actions, and missing sources. Build one timeline across identity-provider events, session creation and refresh, Pomerium decisions, policy and route versions, control-plane changes, endpoint and network evidence, and the application's final object action.
Confirm identity without assuming one field names one person. Compare subject, account, session, device, workload, issuer, audience, credential, source network, route, method, resource, and result. Test alternate explanations. A new location can mean account compromise, a corporate proxy, travel, or an address-allocation change. An accepted gateway request does not prove that the application allowed the intended object action.
Response and recovery
- Acknowledge the alert and preserve its exact inputs.
- Check whether required telemetry is present, timely, and internally consistent.
- Identify the protected asset, harmful consequence, current authority, and plausible blast radius.
- Search for the same subject, credential, session, route, resource, indicator, and behavior before and after the alert.
- Seek independent corroboration and record evidence that contradicts the initial hypothesis.
- Classify the result with a defined disposition and confidence. Do not use
false positivefor a real but permitted event that the rule intentionally detects. - Contain when expected harm exceeds the cost of a reversible action. Preserve evidence before destructive remediation when delay is safe.
- Escalate with a concise evidence record, open questions, affected owners, and the exact requested decision.
- Verify containment and recovery with negative tests for old authority and positive tests for approved access.
- Update the detection test set, telemetry contract, runbook, and risk record.
Design tradeoffs and residual risk
Fast triage favors a small evidence set and can miss a distributed campaign. Broad investigation consumes time and can delay containment. Automated enrichment reduces manual work and can import stale, incorrect, or excessive personal data. Immediate credential revocation limits misuse and can destroy access needed for recovery.
Residual risk includes valid-account abuse that resembles normal work, direct paths outside gateway evidence, compromised source logs, identifier mismatch, clock error, short retention, and harmful application actions that access logs cannot see.
Pomerium boundary
Pomerium can record requests, identity context, policy decisions, routes, and component signals for traffic that it processes. Operators must preserve and correlate identity-provider, endpoint, network, cloud, application, and control-plane evidence. Pomerium does not determine incident severity, prove the user's real-world identity, or show a final object mutation unless the application records it.
Exercise
Stage an alert for a disabled subject whose existing session still reaches a sensitive route. Add one benign event with a similar source address and remove one expected telemetry source. Investigate from the preserved alert through the application's final action.
State confidence, blast radius, missing evidence, containment choice, and verification plan. Then change one clock by five minutes and repeat the timeline. Record which conclusions remain supported.
Evaluation checklist
- Does the alert state a protected consequence, affected resource, time window, and evidence contract?
- Can another analyst reproduce the timeline, queries, and disposition from preserved records?
- Did the investigation test alternate explanations and record contradictory evidence?
- Is containment proportionate, reversible where possible, and verified against old authority?
- Did the result improve the detection, telemetry, procedure, or risk decision?
Next learning unit
Digital Forensics for Incident Response
Identify, collect, preserve, examine, analyze, and report digital evidence with stated scope, methods, time, integrity, and uncertainty.
