Kubernetes 1.37 promotes HPAScaleToZero to beta, extending the HorizontalPodAutoscaler (HPA) so supported workloads can move from zero replicas to an appropriate count and later return to zero. The change is aimed at intermittently used or resource-intensive workloads, including queue consumers and workloads that may otherwise leave dedicated CPU or GPU capacity idle.
Scale-to-zero is not a general replacement for capacity planning. It depends on the right metrics provider, introduces startup delay, and leaves no serving Pods while the workload is at zero. Kubernetes Services also do not queue requests for a workload with zero replicas. Those constraints are central to deciding where the feature fits.
What changed in Kubernetes 1.37
HPAScaleToZero was alpha in Kubernetes 1.36 and is listed as a beta feature in Kubernetes 1.37. The Kubernetes feature-gate reference lists it as enabled by default in 1.37. Operators should still verify the configuration of their own clusters, particularly where control-plane settings are managed by a distribution or service provider.
An HPA can use minReplicas: 0 to permit a workload to scale down completely. Scale-up from zero is supported when the HPA uses object or external metrics. Resource metrics such as CPU and memory cannot provide this behavior by themselves because those metrics depend on running Pods. The scale-from-zero Kubernetes Enhancement Proposal documents that limitation.
The feature applies to scalable workload resources, such as a Deployment or StatefulSet, that provide the Kubernetes scale subresource. When the HPA holds a workload at zero, it records a ScaledToZero=True condition. After the workload scales back up, that condition becomes False with the reason NotScaledToZero.
Requirements for scale-to-zero
The HPAScaleToZero feature gate must be enabled on both the kube-apiserver and kube-controller-manager. Although the gate is enabled by default in the upstream Kubernetes 1.37 feature-gate reference, cluster administrators should confirm that their control-plane configuration matches the expected version and settings.
The HPA must set its minimum replica count to zero. The workload should also be managed without a competing replica value in its Deployment or StatefulSet manifest. Kubernetes documentation recommends removing spec.replicas from manifests when an HPA controls the workload; applying a manifest that contains that field can override the HPA-managed count and produce undesirable behavior.
External metrics require an External Metrics API provider. For Prometheus-based setups, the Prometheus Adapter external metrics documentation describes configuring external rules and namespace behavior. Those rules and the availability of the metrics API are operational dependencies: without usable external metrics, the HPA cannot make the decisions expected by a scale-to-zero configuration.
What scale-up from zero looks like
Scale-up from zero uses the regular HPA reconciliation process. The scale-from-zero proposal identifies the default kube-controller-manager synchronization period as 15 seconds, with additional delay possible when the metrics pipeline has not yet produced fresh data. Workload initialization adds another part of the startup experience.
This means the feature should not be treated as an instantaneous activation mechanism. A queue consumer may be a good candidate when delayed processing is acceptable, while a latency-sensitive request path needs a separately designed approach for startup and request handling. The supplied Kubernetes sources do not establish a single recommended stabilization profile for every HPA configuration. In particular, the downscale stabilization window does not apply to scale-up from zero, but operators should test behavior for their own metric and workload patterns.
Operational implications for developers and platform teams
The most direct benefit is eliminating idle Pods and the resources assigned to them. The scale-from-zero proposal identifies potential cost and energy benefits for workloads that are frequently idle or use expensive dedicated CPU or GPU resources. It does not provide a universal savings percentage, and actual results depend on workload demand, startup behavior, metric freshness, and the surrounding infrastructure.
At zero replicas, however, the workload has no running Pods to serve traffic. Kubernetes Services do not buffer requests as part of HPAScaleToZero. If requests must wait while a workload starts, buffering or queueing must be provided by an external system or by the application itself. Relying on a Service alone will not supply that behavior.
Monitoring should include HPA conditions, especially ScaledToZero and ScalingActive. These conditions help distinguish an HPA-managed zero state from a manually set replica count and can expose problems with metric collection or evaluation. Teams should also monitor the external metrics provider because it becomes part of the scaling path.
Version-skew and managed-service behavior require local confirmation. The upstream documentation specifies the control-plane feature-gate requirements, but the supplied material does not document additional restrictions for particular Kubernetes distributions or managed Kubernetes services.
What you should do
- Identify suitable workloads. Start with intermittently used workers, queue consumers, or resource-intensive workloads where idle capacity is significant and startup delay is acceptable.
- Choose an available metric. Use an object or external metric that remains meaningful when the workload has no Pods. Do not expect CPU or memory metrics alone to trigger scale-up from zero.
- Confirm the metrics API. Install or verify the External Metrics API provider and configure the required rules. Prometheus Adapter users should review its external-metrics and namespace configuration.
- Verify the feature gate. Confirm that HPAScaleToZero is enabled on both the API server and controller manager. Kubernetes 1.37 lists the gate as enabled by default, but cluster-level verification remains appropriate.
- Configure the HPA deliberately. Set
minReplicas: 0, use a scalable workload resource, and removespec.replicasfrom workload manifests managed by the HPA. - Test failure and recovery paths. Exercise scale-down, scale-up, metric-provider failure, and manually setting replicas to zero. Observe HPA conditions and account for reconciliation, metric freshness, and application initialization delays.
- Design request handling separately. If users or upstream systems send requests while the workload is at zero, provide application-level or external buffering rather than assuming the Kubernetes Service will queue them.
Beta availability and limits
HPAScaleToZero is available as a beta feature in Kubernetes 1.37, following its alpha phase in 1.36. Beta status means teams can evaluate and use the capability, but it should not be presented as a generally stable feature without qualification.
The core behavior is clear: an HPA can use object or external metrics and minReplicas: 0 to scale a supported workload down to zero and back up. The practical limits are equally important. The feature does not add Service-level request buffering, does not make CPU or memory metrics sufficient for scale-up from zero, and does not come with documented universal savings or cold-start benchmarks.



