SOC 2 Control Implementation Guide

System Operations

Availability and Performance Monitoring for SOC 2

AOT availability, ingestion health, telemetry latency, processing capacity, storage pressure, alerting thresholds, and customer-facing monitoring commitments are measured and escalated when thresholds are exceeded.

Use this guide to put the control into operation, decide what records to retain, and check that an auditor can trace the evidence back to the work your team performed.

Maintained by GreenHat Security · Reviewed August 21, 2026

What this control should accomplish

The team can see whether the service and its critical processing paths are healthy, detect customer-impacting degradation early, and route actionable alerts to someone who can respond.

First SOC 2 program

A credible starting point

Monitor a few customer-critical journeys, service errors, latency, queue depth, storage, and dependency health. Put actionable thresholds on an owned on-call route and record the resulting incidents or tuning decisions.

As the company scales

Make it repeatable

Use service objectives, synthetic checks, dependency telemetry, capacity signals, alert ownership, and trend reviews to distinguish real customer risk from noise across regions and product components.

How to implement Availability and Performance Monitoring

  1. 1

    Identify critical service journeys

    Map the customer actions and background processing paths whose failure would materially affect availability, including the dependencies needed for each path.

    You should end up with: Prioritized service and dependency map

  2. 2

    Select useful health signals

    Choose availability, error, latency, saturation, queue, ingestion, storage, and dependency measures that reveal whether each critical path works.

    You should end up with: Health-signal catalog with owners and data sources

  3. 3

    Set actionable thresholds

    Define warning and urgent conditions using observed behavior and business impact, then document who receives each alert and what first response is expected.

    You should end up with: Approved thresholds and alert-routing matrix

  4. 4

    Test the response path

    Trigger safe test conditions or synthetic failures to confirm alerts fire, reach the correct person, contain enough context, and create a durable response record.

    You should end up with: Alert-path test with timing and findings

  5. 5

    Review trends and tune

    Review customer impact, missed conditions, noisy alerts, capacity pressure, and recurring failures; assign improvements and retain the decision record.

    You should end up with: Dated monitoring review and action tracker

Evidence to keep, and what it should prove

Build the evidence set in three layers: what defines the control, who approved or reviewed it, and what proves it operated. Collect operating records when the work happens so they remain dated, attributable, correctly scoped, and traceable to the underlying activity.

Before sharing, remove unrelated personal or customer data, never expose passwords, tokens, or secret values, preserve enough source context to authenticate the record, and use the secure exchange approved for the engagement.

Operating / technical evidence

Dated proof that the control actually ran, such as tickets, logs, settings, exports, reports, and test results.

Availability dashboards

  • Confirm what the record proves

    Critical services and customer journeys are measured over a defined period using signals that expose availability, error, latency, and saturation conditions.

  • Include this context

    Service and environment

  • Include this context

    Covered period

  • Include this context

    Availability or health measure

  • Include this context

    Data source

  • Include this context

    Threshold or objective

  • Include this context

    Last refresh

Weak evidence to avoid

A current uptime percentage with no service boundary, underlying data source, date range, or threshold.

uptime/health checks

  • Confirm what the record proves

    Independent or service-level checks ran at known intervals and recorded whether critical endpoints or workflows responded as expected.

  • Include this context

    Check name

  • Include this context

    Target and environment

  • Include this context

    Check interval

  • Include this context

    Expected result

  • Include this context

    Execution timestamp

  • Include this context

    Observed result

Weak evidence to avoid

A green status badge that does not identify the tested target, interval, response condition, or history.

ingestion metrics

  • Confirm what the record proves

    The team monitors event flow, queue depth, delay, failures, and backlog conditions that could degrade processing or hide service problems.

  • Include this context

    Pipeline or queue

  • Include this context

    Metric name

  • Include this context

    Covered period

  • Include this context

    Expected range

  • Include this context

    Observed value

  • Include this context

    Source

Weak evidence to avoid

One current message count without trend history, expected range, queue identity, or processing delay.

alert thresholds

  • Confirm what the record proves

    Monitoring conditions have documented trigger values, severity, ownership, and escalation behavior tied to practical service impact.

  • Include this context

    Alert name

  • Include this context

    Signal and scope

  • Include this context

    Trigger condition

  • Include this context

    Severity

  • Include this context

    Route or owner

  • Include this context

    Effective version

Weak evidence to avoid

A vendor-default alert enabled without an owner, service scope, business rationale, or effective date.

incident/escalation tickets

  • Confirm what the record proves

    Triggered service conditions were acknowledged, assessed, escalated when needed, and resolved with an attributable operating record.

  • Include this context

    Ticket ID

  • Include this context

    Trigger and detected time

  • Include this context

    Affected service

  • Include this context

    Owner

  • Include this context

    Impact and decision

  • Include this context

    Resolution time

Weak evidence to avoid

A chat message saying the service recovered with no linked alert, owner, impact assessment, timestamps, or ticket.

Which records should you prepare for the audit?

Type 1

Evidence at the as-of date

The current critical-service map, active health checks, monitoring thresholds and routes, and a recent alert-path test or monitoring review showing the design in use at the review date.

Type 2

Evidence across the review period

Every availability, performance, ingestion, capacity, or dependency alert generated by the covered monitoring rules during the review period, plus every scheduled monitoring review and alert-path test due during that period, including acknowledged, escalated, tuned, failed, and missed occurrences.

Completeness check

Export alert history directly from every covered monitoring platform and reconcile each alert ID to its on-call or incident disposition; separately reconcile the monitoring-review and test calendar to completed records, preserving failed, missed, and rescheduled items rather than omitting them.

Build the record set from

  • Application performance monitoring
  • Cloud and dependency metrics
  • Synthetic monitoring
  • On-call platform
  • Incident ticketing

Keep these fields for each record

  • Alert or review ID
  • Service and environment
  • Rule or check
  • Triggered or due time
  • Acknowledgment
  • Owner
  • Impact decision
  • Resolution or tuning result

How an auditor may test this control

Use this checklist to prepare for procedures an auditor may perform. The exact steps and sample selection depend on your engagement scope and the service auditor's professional judgment.

  • Confirm the intended control outcome

    Determine whether the control is designed to achieve this result: The team can see whether the service and its critical processing paths are healthy, detect customer-impacting degradation early, and route actionable alerts to someone who can respond.

  • Confirm ownership and operating cadence

    Compare the documented owner with the intended role (Security Operations / Engineering / Infrastructure), then compare dated records with the stated cadence: Continuous/ongoing; periodic review by risk.

  • Establish the complete audit record set

    Export alert history directly from every covered monitoring platform and reconcile each alert ID to its on-call or incident disposition; separately reconcile the monitoring-review and test calendar to completed records, preserving failed, missed, and rescheduled items rather than omitting them.

  • Prepare the as-of-date evidence for a Type 1 engagement

    The current critical-service map, active health checks, monitoring thresholds and routes, and a recent alert-path test or monitoring review showing the design in use at the review date.

  • Prepare period evidence for a Type 2 engagement

    Every availability, performance, ingestion, capacity, or dependency alert generated by the covered monitoring rules during the review period, plus every scheduled monitoring review and alert-path test due during that period, including acknowledged, escalated, tuned, failed, and missed occurrences.

  • Inspect the operating / technical evidence

    Dated proof that the control actually ran, such as tickets, logs, settings, exports, reports, and test results.

    • Inspect Availability dashboards

      For each selected record, confirm it demonstrates Critical services and customer journeys are measured over a defined period using signals that expose availability, error, latency, and saturation conditions.

      • Service and environment
      • Covered period
      • Availability or health measure
      • Data source
      • Threshold or objective
      • Last refresh
    • Inspect uptime/health checks

      For each selected record, confirm it demonstrates Independent or service-level checks ran at known intervals and recorded whether critical endpoints or workflows responded as expected.

      • Check name
      • Target and environment
      • Check interval
      • Expected result
      • Execution timestamp
      • Observed result
    • Inspect ingestion metrics

      For each selected record, confirm it demonstrates The team monitors event flow, queue depth, delay, failures, and backlog conditions that could degrade processing or hide service problems.

      • Pipeline or queue
      • Metric name
      • Covered period
      • Expected range
      • Observed value
      • Source
    • Inspect alert thresholds

      For each selected record, confirm it demonstrates Monitoring conditions have documented trigger values, severity, ownership, and escalation behavior tied to practical service impact.

      • Alert name
      • Signal and scope
      • Trigger condition
      • Severity
      • Route or owner
      • Effective version
    • Inspect incident/escalation tickets

      For each selected record, confirm it demonstrates Triggered service conditions were acknowledged, assessed, escalated when needed, and resolved with an attributable operating record.

      • Ticket ID
      • Trigger and detected time
      • Affected service
      • Owner
      • Impact and decision
      • Resolution time
  • Trace the control from design to operation

    Use the categories that apply to this control: connect any policy or design artifact to its approval or review record, then trace a selected operating record through execution, result, and any exception or remediation.

Common implementation and evidence gaps

  • Infrastructure is green while a customer-critical transaction is failing.
  • Thresholds use defaults that are either constantly noisy or too late to prevent impact.
  • Alerts go to a chat channel without a named responder or escalation path.
  • Dashboards omit queues, ingestion delay, storage pressure, or external dependencies.
  • The team tunes alerts but does not retain why thresholds changed.

Before you call this control ready

  • Can each critical customer journey be tied to a health signal and accountable owner?
  • Does a safe test reach the on-call person and produce a traceable record?
  • Are thresholds based on business impact and observed operating ranges?
  • Can the team identify monitoring blind spots and show how they are being addressed?
  • Do recent incidents reveal any conditions that monitoring should have caught earlier?

Trust Services Criteria references

These identifiers help you navigate related Trust Services Criteria. They do not reproduce the criteria or prove that this control fully addresses them in your environment.

  • CC7.2
  • A1.1

Confirm final scope, mappings, and testing expectations with your service auditor. SOC 2® is an AICPA trademark; GreenHat Security is not affiliated with or endorsed by AICPA.