First SOC 2 program
A credible starting point
Monitor a short list of bottleneck signals for each critical service, document acceptable headroom, review growth trends regularly, and create an owned task when a threshold or provider limit approaches.
Availability / Resilience
Monitoring includes capacity, scalability, ingestion health, storage pressure, queue/backlog indicators, and service health signals needed to identify resilience risk and trigger remediation.
Use this guide to put the control into operation, decide what records to retain, and check that an auditor can trace the evidence back to the work your team performed.
Maintained by GreenHat Security ยท Reviewed August 21, 2026
The team understands current headroom and foreseeable bottlenecks for critical services, then acts before compute, storage, connection, queue, throughput, or provider limits cause customer impact.
First SOC 2 program
Monitor a short list of bottleneck signals for each critical service, document acceptable headroom, review growth trends regularly, and create an owned task when a threshold or provider limit approaches.
As the company scales
Combine demand forecasts, load tests, autoscaling boundaries, quota management, dependency capacity, and scenario reviews to plan changes before launches and sustained growth.
For each critical service, identify compute, memory, database, connection, storage, queue, throughput, rate-limit, and third-party constraints with an owner.
You should end up with: Service capacity and constraint map
Use observed peaks, recovery time, scaling delay, customer commitments, and provider quotas to define when investigation and action should begin.
You should end up with: Capacity thresholds with rationale
Track peaks and sustained use, backlog growth, storage trajectory, throttling, autoscaling behavior, and hard provider limits.
You should end up with: Capacity dashboard covering leading and hard-limit signals
Consider customer growth, launches, migrations, retention changes, seasonal load, and dependency plans before material events.
You should end up with: Dated capacity forecast and assumptions
Run safe load or failover exercises to confirm the system scales as expected and identify the component that becomes limiting first.
You should end up with: Scaling test with measured limits and findings
Assign owners and dates for quota increases, architecture changes, cleanup, tuning, or accepted constraints and verify completion.
You should end up with: Capacity remediation tracker
Build the evidence set in three layers: what defines the control, who approved or reviewed it, and what proves it operated. Collect operating records when the work happens so they remain dated, attributable, correctly scoped, and traceable to the underlying activity.
Before sharing, remove unrelated personal or customer data, never expose passwords, tokens, or secret values, preserve enough source context to authenticate the record, and use the secure exchange approved for the engagement.
Records showing that an accountable person reviewed, approved, challenged, or accepted the work.
Confirm what the record proves
Known launches, customer growth, migrations, retention changes, and failure scenarios are translated into capacity decisions and owned work.
Include this context
Review date
Include this context
Services and scenarios
Include this context
Demand assumptions
Include this context
Constraints identified
Include this context
Decision
Include this context
Actions and owners
Weak evidence to avoid
Meeting notes say capacity looks fine without assumptions, measured constraints, decisions, owners, or due dates.
Dated proof that the control actually ran, such as tickets, logs, settings, exports, reports, and test results.
Confirm what the record proves
Critical services expose current and historical headroom across the resources and hard limits most likely to constrain customer operations.
Include this context
Service and environment
Include this context
Resource or limit
Include this context
Covered period
Include this context
Current and peak values
Include this context
Warning threshold
Include this context
Data source
Weak evidence to avoid
A current CPU chart with no service scope, peak history, threshold, storage, queue, database, or provider-limit context.
Confirm what the record proves
The team reviews sustained and peak demand, growth rate, headroom, and likely exhaustion dates instead of relying only on a live dashboard.
Include this context
Covered services
Include this context
Reporting period
Include this context
Peak and sustained use
Include this context
Growth assumption
Include this context
Remaining headroom
Include this context
Reviewer
Weak evidence to avoid
An average-utilization percentage with no peaks, growth trend, quota comparison, service owner, or review date.
Confirm what the record proves
The team can detect processing pressure and delay that resource averages may hide, with thresholds tied to timely response.
Include this context
Queue or pipeline
Include this context
Depth and age measures
Include this context
Covered period
Include this context
Threshold
Include this context
Owner or route
Include this context
Observed exceptions
Weak evidence to avoid
A current queue count with no oldest-item age, trend, trigger, owner, or evidence of response to breaches.
Confirm what the record proves
Identified capacity risks become accountable changes with due dates, decision history, completion proof, and validation of added headroom.
Include this context
Risk and service
Include this context
Triggering measure
Include this context
Owner
Include this context
Opened and due dates
Include this context
Treatment
Include this context
Completion validation
Weak evidence to avoid
A backlog item called improve scalability with no measured risk, target, owner, due date, or validation result.
Type 1
The current service constraint map, active capacity thresholds and dashboards, latest forecast or scaling review, and open capacity actions showing the process is designed and in use at the review date.
Type 2
Every scheduled capacity or scaling review due during the review period, every capacity threshold or hard-limit alert generated for covered services, and every material launch, migration, retention, or demand event designated to trigger a capacity assessment, including missed, deferred, and closed occurrences.
Reconcile full-period threshold and provider-limit alert exports with on-call or engineering records, then reconcile the approved review calendar and designated launch or migration triggers to capacity assessments; account for every missed review, suppressed alert, deferred decision, and action without validation.
Use this checklist to prepare for procedures an auditor may perform. The exact steps and sample selection depend on your engagement scope and the service auditor's professional judgment.
Determine whether the control is designed to achieve this result: The team understands current headroom and foreseeable bottlenecks for critical services, then acts before compute, storage, connection, queue, throughput, or provider limits cause customer impact.
Compare the documented owner with the intended role (Operations Lead / Infrastructure Owner), then compare dated records with the stated cadence: Continuous for backups/monitoring; quarterly/annual testing as defined.
Reconcile full-period threshold and provider-limit alert exports with on-call or engineering records, then reconcile the approved review calendar and designated launch or migration triggers to capacity assessments; account for every missed review, suppressed alert, deferred decision, and action without validation.
The current service constraint map, active capacity thresholds and dashboards, latest forecast or scaling review, and open capacity actions showing the process is designed and in use at the review date.
Every scheduled capacity or scaling review due during the review period, every capacity threshold or hard-limit alert generated for covered services, and every material launch, migration, retention, or demand event designated to trigger a capacity assessment, including missed, deferred, and closed occurrences.
Records showing that an accountable person reviewed, approved, challenged, or accepted the work.
For each selected record, confirm it demonstrates Known launches, customer growth, migrations, retention changes, and failure scenarios are translated into capacity decisions and owned work.
Dated proof that the control actually ran, such as tickets, logs, settings, exports, reports, and test results.
For each selected record, confirm it demonstrates Critical services expose current and historical headroom across the resources and hard limits most likely to constrain customer operations.
For each selected record, confirm it demonstrates The team reviews sustained and peak demand, growth rate, headroom, and likely exhaustion dates instead of relying only on a live dashboard.
For each selected record, confirm it demonstrates The team can detect processing pressure and delay that resource averages may hide, with thresholds tied to timely response.
For each selected record, confirm it demonstrates Identified capacity risks become accountable changes with due dates, decision history, completion proof, and validation of added headroom.
Use the categories that apply to this control: connect any policy or design artifact to its approval or review record, then trace a selected operating record through execution, result, and any exception or remediation.
These identifiers help you navigate related Trust Services Criteria. They do not reproduce the criteria or prove that this control fully addresses them in your environment.
Confirm final scope, mappings, and testing expectations with your service auditor. SOC 2ยฎ is an AICPA trademark; GreenHat Security is not affiliated with or endorsed by AICPA.