SOC 2 Control Implementation Guide

Availability / Resilience

Capacity and Scalability Planning for SOC 2

Monitoring includes capacity, scalability, ingestion health, storage pressure, queue/backlog indicators, and service health signals needed to identify resilience risk and trigger remediation.

Use this guide to put the control into operation, decide what records to retain, and check that an auditor can trace the evidence back to the work your team performed.

Maintained by GreenHat Security ยท Reviewed August 21, 2026

What this control should accomplish

The team understands current headroom and foreseeable bottlenecks for critical services, then acts before compute, storage, connection, queue, throughput, or provider limits cause customer impact.

First SOC 2 program

A credible starting point

Monitor a short list of bottleneck signals for each critical service, document acceptable headroom, review growth trends regularly, and create an owned task when a threshold or provider limit approaches.

As the company scales

Make it repeatable

Combine demand forecasts, load tests, autoscaling boundaries, quota management, dependency capacity, and scenario reviews to plan changes before launches and sustained growth.

How to implement Capacity and Scalability Planning

  1. 1

    Map capacity constraints

    For each critical service, identify compute, memory, database, connection, storage, queue, throughput, rate-limit, and third-party constraints with an owner.

    You should end up with: Service capacity and constraint map

  2. 2

    Set headroom and warning levels

    Use observed peaks, recovery time, scaling delay, customer commitments, and provider quotas to define when investigation and action should begin.

    You should end up with: Capacity thresholds with rationale

  3. 3

    Monitor trends and limits

    Track peaks and sustained use, backlog growth, storage trajectory, throttling, autoscaling behavior, and hard provider limits.

    You should end up with: Capacity dashboard covering leading and hard-limit signals

  4. 4

    Review upcoming demand

    Consider customer growth, launches, migrations, retention changes, seasonal load, and dependency plans before material events.

    You should end up with: Dated capacity forecast and assumptions

  5. 5

    Validate scaling behavior

    Run safe load or failover exercises to confirm the system scales as expected and identify the component that becomes limiting first.

    You should end up with: Scaling test with measured limits and findings

  6. 6

    Track capacity actions

    Assign owners and dates for quota increases, architecture changes, cleanup, tuning, or accepted constraints and verify completion.

    You should end up with: Capacity remediation tracker

Evidence to keep, and what it should prove

Build the evidence set in three layers: what defines the control, who approved or reviewed it, and what proves it operated. Collect operating records when the work happens so they remain dated, attributable, correctly scoped, and traceable to the underlying activity.

Before sharing, remove unrelated personal or customer data, never expose passwords, tokens, or secret values, preserve enough source context to authenticate the record, and use the secure exchange approved for the engagement.

Approval / review evidence

Records showing that an accountable person reviewed, approved, challenged, or accepted the work.

scaling reviews

  • Confirm what the record proves

    Known launches, customer growth, migrations, retention changes, and failure scenarios are translated into capacity decisions and owned work.

  • Include this context

    Review date

  • Include this context

    Services and scenarios

  • Include this context

    Demand assumptions

  • Include this context

    Constraints identified

  • Include this context

    Decision

  • Include this context

    Actions and owners

Weak evidence to avoid

Meeting notes say capacity looks fine without assumptions, measured constraints, decisions, owners, or due dates.

Operating / technical evidence

Dated proof that the control actually ran, such as tickets, logs, settings, exports, reports, and test results.

Capacity dashboards

  • Confirm what the record proves

    Critical services expose current and historical headroom across the resources and hard limits most likely to constrain customer operations.

  • Include this context

    Service and environment

  • Include this context

    Resource or limit

  • Include this context

    Covered period

  • Include this context

    Current and peak values

  • Include this context

    Warning threshold

  • Include this context

    Data source

Weak evidence to avoid

A current CPU chart with no service scope, peak history, threshold, storage, queue, database, or provider-limit context.

utilization reports

  • Confirm what the record proves

    The team reviews sustained and peak demand, growth rate, headroom, and likely exhaustion dates instead of relying only on a live dashboard.

  • Include this context

    Covered services

  • Include this context

    Reporting period

  • Include this context

    Peak and sustained use

  • Include this context

    Growth assumption

  • Include this context

    Remaining headroom

  • Include this context

    Reviewer

Weak evidence to avoid

An average-utilization percentage with no peaks, growth trend, quota comparison, service owner, or review date.

backlog/queue monitoring

  • Confirm what the record proves

    The team can detect processing pressure and delay that resource averages may hide, with thresholds tied to timely response.

  • Include this context

    Queue or pipeline

  • Include this context

    Depth and age measures

  • Include this context

    Covered period

  • Include this context

    Threshold

  • Include this context

    Owner or route

  • Include this context

    Observed exceptions

Weak evidence to avoid

A current queue count with no oldest-item age, trend, trigger, owner, or evidence of response to breaches.

remediation tickets

  • Confirm what the record proves

    Identified capacity risks become accountable changes with due dates, decision history, completion proof, and validation of added headroom.

  • Include this context

    Risk and service

  • Include this context

    Triggering measure

  • Include this context

    Owner

  • Include this context

    Opened and due dates

  • Include this context

    Treatment

  • Include this context

    Completion validation

Weak evidence to avoid

A backlog item called improve scalability with no measured risk, target, owner, due date, or validation result.

Which records should you prepare for the audit?

Type 1

Evidence at the as-of date

The current service constraint map, active capacity thresholds and dashboards, latest forecast or scaling review, and open capacity actions showing the process is designed and in use at the review date.

Type 2

Evidence across the review period

Every scheduled capacity or scaling review due during the review period, every capacity threshold or hard-limit alert generated for covered services, and every material launch, migration, retention, or demand event designated to trigger a capacity assessment, including missed, deferred, and closed occurrences.

Completeness check

Reconcile full-period threshold and provider-limit alert exports with on-call or engineering records, then reconcile the approved review calendar and designated launch or migration triggers to capacity assessments; account for every missed review, suppressed alert, deferred decision, and action without validation.

Build the record set from

  • โ€ข Cloud and application metrics
  • โ€ข Database, queue, and storage monitoring
  • โ€ข Provider quota consoles
  • โ€ข Product launch or change calendar
  • โ€ข Engineering issue tracker

Keep these fields for each record

  • โ€ข Occurrence ID
  • โ€ข Service and constraint
  • โ€ข Alert, review, or trigger type
  • โ€ข Detected or due time
  • โ€ข Measured headroom
  • โ€ข Owner
  • โ€ข Decision
  • โ€ข Action and validation

How an auditor may test this control

Use this checklist to prepare for procedures an auditor may perform. The exact steps and sample selection depend on your engagement scope and the service auditor's professional judgment.

  • Confirm the intended control outcome

    Determine whether the control is designed to achieve this result: The team understands current headroom and foreseeable bottlenecks for critical services, then acts before compute, storage, connection, queue, throughput, or provider limits cause customer impact.

  • Confirm ownership and operating cadence

    Compare the documented owner with the intended role (Operations Lead / Infrastructure Owner), then compare dated records with the stated cadence: Continuous for backups/monitoring; quarterly/annual testing as defined.

  • Establish the complete audit record set

    Reconcile full-period threshold and provider-limit alert exports with on-call or engineering records, then reconcile the approved review calendar and designated launch or migration triggers to capacity assessments; account for every missed review, suppressed alert, deferred decision, and action without validation.

  • Prepare the as-of-date evidence for a Type 1 engagement

    The current service constraint map, active capacity thresholds and dashboards, latest forecast or scaling review, and open capacity actions showing the process is designed and in use at the review date.

  • Prepare period evidence for a Type 2 engagement

    Every scheduled capacity or scaling review due during the review period, every capacity threshold or hard-limit alert generated for covered services, and every material launch, migration, retention, or demand event designated to trigger a capacity assessment, including missed, deferred, and closed occurrences.

  • Inspect the approval / review evidence

    Records showing that an accountable person reviewed, approved, challenged, or accepted the work.

    • Inspect scaling reviews

      For each selected record, confirm it demonstrates Known launches, customer growth, migrations, retention changes, and failure scenarios are translated into capacity decisions and owned work.

      • Review date
      • Services and scenarios
      • Demand assumptions
      • Constraints identified
      • Decision
      • Actions and owners
  • Inspect the operating / technical evidence

    Dated proof that the control actually ran, such as tickets, logs, settings, exports, reports, and test results.

    • Inspect Capacity dashboards

      For each selected record, confirm it demonstrates Critical services expose current and historical headroom across the resources and hard limits most likely to constrain customer operations.

      • Service and environment
      • Resource or limit
      • Covered period
      • Current and peak values
      • Warning threshold
      • Data source
    • Inspect utilization reports

      For each selected record, confirm it demonstrates The team reviews sustained and peak demand, growth rate, headroom, and likely exhaustion dates instead of relying only on a live dashboard.

      • Covered services
      • Reporting period
      • Peak and sustained use
      • Growth assumption
      • Remaining headroom
      • Reviewer
    • Inspect backlog/queue monitoring

      For each selected record, confirm it demonstrates The team can detect processing pressure and delay that resource averages may hide, with thresholds tied to timely response.

      • Queue or pipeline
      • Depth and age measures
      • Covered period
      • Threshold
      • Owner or route
      • Observed exceptions
    • Inspect remediation tickets

      For each selected record, confirm it demonstrates Identified capacity risks become accountable changes with due dates, decision history, completion proof, and validation of added headroom.

      • Risk and service
      • Triggering measure
      • Owner
      • Opened and due dates
      • Treatment
      • Completion validation
  • Trace the control from design to operation

    Use the categories that apply to this control: connect any policy or design artifact to its approval or review record, then trace a selected operating record through execution, result, and any exception or remediation.

Common implementation and evidence gaps

  • Average utilization hides short peaks, queue buildup, or tenant-specific pressure.
  • Autoscaling is assumed to remove capacity risk without checking quotas, warm-up time, or downstream limits.
  • Storage growth and retention changes are absent from planning until a hard limit is near.
  • Forecasts are discussed but do not create owned engineering actions.
  • Load tests measure one component while a dependency becomes the real bottleneck.

Before you call this control ready

  • Can each critical service name its first likely bottleneck and remaining headroom?
  • Do alerts fire early enough to act before a hard limit or customer impact?
  • Are cloud and third-party quotas visible alongside ordinary utilization metrics?
  • Does the latest forecast include known launches, growth, migrations, and retention changes?
  • Can recent capacity risks be traced to an owner, decision, and completed action?

Trust Services Criteria references

These identifiers help you navigate related Trust Services Criteria. They do not reproduce the criteria or prove that this control fully addresses them in your environment.

  • A1.1

Confirm final scope, mappings, and testing expectations with your service auditor. SOC 2ยฎ is an AICPA trademark; GreenHat Security is not affiliated with or endorsed by AICPA.