Question-led guide · evaluation
When is a capacity forecast safe to use for decisions?
Evaluate capacity forecasts at the decision horizon, including asymmetric underprediction cost, data revisions, and fallback controls.
Direct answer
A capacity forecast is useful when it supports a named provisioning or admission decision at the right horizon and its errors are tolerable in the expensive direction. Backtest on historical periods without future leakage, then test deployment changes, exceptional demand, and missing metrics. Record the decision the forecast would have made, not only its average prediction error. Keep a capacity floor and an operator override for uncertain periods.
Start with the action and its lead time
Buying reserved capacity, scheduling a scale-up, and rejecting new work have different deadlines. Identify when the decision must be made and when capacity can actually arrive. A forecast for tomorrow is irrelevant to an autoscaler that reacts in minutes unless it addresses a different constraint. State the resource bottleneck, usable headroom, and whether a missed peak harms users or merely delays optional work.
Average error conceals expensive tails
A model that is usually close can still miss the rare peaks that cause a queue to run past its deadline. Report error by horizon, day type, workload class, and direction. Underprediction near saturation has a different consequence from modest overprediction in a cheap pool. Keep the baseline comparison simple: yesterday’s same period or a guarded moving window may perform adequately.
A holiday forecast meets a slow scale-up
In a constructed service, ordinary demand is 600 tasks per minute and provisioned capacity is 800. A promotion lifts demand to 1,200; new workers take twenty minutes to become useful. The model’s monthly error is small because most periods are quiet, but its promotion forecast is 750. A decision based on the aggregate score leaves the queue exposed precisely when lead time matters.
Use a decision ledger rather than one accuracy score
Record the action that would have followed every forecast at its issue time.
| Field | Example review value |
|---|---|
| Forecast issue and horizon | 09:00 for a 10:00 capacity decision |
| Available evidence | Data visible at 09:00 only |
| Proposed capacity | 800 tasks per minute |
| Actual peak | 1,200 tasks per minute |
| Decision cost | Queue delay and rejected admissions |
Backtest without borrowing the future
Freeze data at the forecast issue time. Reproduce missing observations, late telemetry, deployment history, and holidays as they were known then. Compare forecast-led decisions with the guarded baseline and record how often the capacity floor prevented harm. A model revision should be assessed on held-out periods and then observed under limited rollout before it can reduce spare capacity.
Operate uncertainty as a capacity rule
If the input distribution changes or the prediction interval widens, preserve a stated minimum and escalate the decision to an operator. Autoscaling itself has measurement windows and readiness delay; it cannot conjure an unavailable downstream quota. Stop using the forecast when its data contract fails and reopen it only after a fresh backtest across the affected workload shape.
Evidence boundary for capacity forecasts
- Kubernetes horizontal autoscaling: Kubernetes documents metric-based scaling behavior and its control loop. It does not validate this invented demand forecast or lead time.
- Google SRE managing load: Google SRE discusses load management as a combination of operating mechanisms. The account does not prescribe a forecast error threshold for this service.
The service rates are fictional. The ledger requires measured startup times, queue limits, and real decision costs.
Evidence
Autoscaling uses observed metrics and control calculations rather than instant capacity.
Kubernetes documents metric-based scaling behavior and its control loop.
Primary source · official-doc · checked Oct 7, 2026
Limit: It does not validate this invented demand forecast or lead time.
Reliable load management combines strategies and accounts for traffic and capacity under change.
Google SRE discusses load management as a combination of operating mechanisms.
Primary source · official-doc · checked Oct 7, 2026
Limit: The account does not prescribe a forecast error threshold for this service.
Limitations
Forecast usefulness depends on the local workload, lead time, and error cost. The illustrative rates are not benchmark or capacity recommendations.
FAQ
- Is low mean absolute error enough to automate provisioning?
- No. Test the actual capacity decision, especially underprediction near saturation and periods with long scale-up delay.
- What should happen when the model input is missing?
- Use a documented capacity floor or safe fallback and mark the forecast unavailable rather than inventing a confident value.
Related guides
Continue within AIOps alerting and operations, or use one of these adjacent diagnostics:
Editorial QA: automated native-English, structure, source-presence, and link checks completed . This record is not an independent expert endorsement. Review boundary.
