Skip to main content

Triage and Investigate Security Alerts

Turn an access alert into a supported finding, a bounded response decision, and useful feedback for the detection.

Learning outcomes

  • Establish the alert's scope, urgency, evidence quality, and protected consequence.
  • Reconstruct identity, authorization, route, application, and control-plane activity on one timeline.
  • Select a proportionate containment or escalation action and preserve the evidence that supports it.
  • Feed investigation results back into detection logic, telemetry, and response procedures.

Operating objective

Triage must decide whether an alert represents a credible harmful condition, how quickly the team must act, which assets and authority may be affected, and what evidence or containment comes next. Investigation must support that decision with reproducible facts. It must not become an open-ended search for everything that could have happened.

Start with the alert contract: threat hypothesis, subject, resource, action, time window, detection version, contributing signals, confidence, severity, owner, and expected response. Protect severe alerts from automatic downgrade when a required source is missing. A telemetry failure can increase uncertainty and urgency.

Signals and evidence

Preserve the original alert and query results. Record source identifiers, event times and ingestion times, time zones, clock quality, query text, filters, result count, analyst actions, and missing sources. Build one timeline across identity-provider events, session creation and refresh, Pomerium decisions, policy and route versions, control-plane changes, endpoint and network evidence, and the application's final object action.

Confirm identity without assuming one field names one person. Compare subject, account, session, device, workload, issuer, audience, credential, source network, route, method, resource, and result. Test alternate explanations. A new location can mean account compromise, a corporate proxy, travel, or an address-allocation change. An accepted gateway request does not prove that the application allowed the intended object action.

Response and recovery

  1. Acknowledge the alert and preserve its exact inputs.
  2. Check whether required telemetry is present, timely, and internally consistent.
  3. Identify the protected asset, harmful consequence, current authority, and plausible blast radius.
  4. Search for the same subject, credential, session, route, resource, indicator, and behavior before and after the alert.
  5. Seek independent corroboration and record evidence that contradicts the initial hypothesis.
  6. Classify the result with a defined disposition and confidence. Do not use false positive for a real but permitted event that the rule intentionally detects.
  7. Contain when expected harm exceeds the cost of a reversible action. Preserve evidence before destructive remediation when delay is safe.
  8. Escalate with a concise evidence record, open questions, affected owners, and the exact requested decision.
  9. Verify containment and recovery with negative tests for old authority and positive tests for approved access.
  10. Update the detection test set, telemetry contract, runbook, and risk record.

Design tradeoffs and residual risk

Fast triage favors a small evidence set and can miss a distributed campaign. Broad investigation consumes time and can delay containment. Automated enrichment reduces manual work and can import stale, incorrect, or excessive personal data. Immediate credential revocation limits misuse and can destroy access needed for recovery.

Residual risk includes valid-account abuse that resembles normal work, direct paths outside gateway evidence, compromised source logs, identifier mismatch, clock error, short retention, and harmful application actions that access logs cannot see.

Pomerium boundary

Pomerium can record requests, identity context, policy decisions, routes, and component signals for traffic that it processes. Operators must preserve and correlate identity-provider, endpoint, network, cloud, application, and control-plane evidence. Pomerium does not determine incident severity, prove the user's real-world identity, or show a final object mutation unless the application records it.

Exercise

Stage an alert for a disabled subject whose existing session still reaches a sensitive route. Add one benign event with a similar source address and remove one expected telemetry source. Investigate from the preserved alert through the application's final action.

State confidence, blast radius, missing evidence, containment choice, and verification plan. Then change one clock by five minutes and repeat the timeline. Record which conclusions remain supported.

Evaluation checklist

  • Does the alert state a protected consequence, affected resource, time window, and evidence contract?
  • Can another analyst reproduce the timeline, queries, and disposition from preserved records?
  • Did the investigation test alternate explanations and record contradictory evidence?
  • Is containment proportionate, reversible where possible, and verified against old authority?
  • Did the result improve the detection, telemetry, procedure, or risk decision?

Next learning unit

Sources and further reading

Keep learning

Authorization and PolicySecurity Operations and Risk

Authorization Decision Log

Record enough structured evidence to explain and test an access decision without storing credentials or excess personal data.

Learn this term
Security Operations and Risk

Security Telemetry

Security telemetry uses logs, metrics, traces, and events to answer defined detection, investigation, and control questions.

Learn this term

Get a Personalized Demo

Schedule a Call with a Pomerium Engineer

Get a Demo