Kubernetes In-Place Pod Resizing: CPU and Memory Changes
Kubernetes in-place Pod resizing lets an operator change a running container’s CPU or memory requests and limits without replacing the Pod. It is useful when a workload’s resource needs change after startup, but it is not the same as adding replicas or nodes. This guide explains how the resize request moves through the API server, scheduler, kubelet, and container runtime, and what operators should verify before using it.
What Is Kubernetes In-Place Pod Resizing?
In-place Pod resizing is a Kubernetes feature for updating the CPU and memory resources assigned to containers in an existing Pod. The Pod keeps its identity, network namespace, and volumes; Kubernetes updates the container’s resource allocation, and a configured resize policy can require a container restart when applying the change. The Kubernetes project introduced the feature as an alpha capability before graduating it to stable in Kubernetes 1.35.
The official Pod resize documentation describes two separate views of resource values. The Pod’s spec.containers[].resources holds the desired requests and limits. The status.containerStatuses[].resources field reports the resources currently configured for the running container. A resize is not complete merely because the desired specification changed; compare those fields and inspect Pod resize conditions to see whether the node applied it.
Resizing is limited to CPU and memory resources. It does not change a Pod’s node, add storage, or dynamically resize every other resource type. It is also different from changing a Deployment’s Pod template: updating that template normally triggers a rollout that creates replacement Pods.
Why In-Place Resizing Exists
Before this feature, changing a Pod’s resource requests or limits generally meant deleting and recreating the Pod. A controller such as a Deployment would replace it, but the application could experience a restart, lose local process state, and repeat startup work. A rolling update can reduce disruption, yet it still replaces Pods and may require spare capacity to keep the service available.
Some workloads have resource needs that vary after launch. A service may need more CPU during a temporary processing phase, or an operator may discover that a long-running worker needs a higher memory limit. Recreating a healthy Pod just to adjust those values adds operational steps and can be especially inconvenient for applications with expensive initialization.
Resizing does not remove the need for capacity planning. A node must have enough resources to accommodate the new allocation. If it does not, Kubernetes can leave the request pending or report it as infeasible rather than silently overcommitting the node. A changed request also affects scheduling and accounting; it does not make additional CPU or memory appear.
How Pod Resizing Works
An operator updates the resource fields on the Pod through its resize subresource. The API server validates the request and records the desired values. The kubelet on the node evaluates whether it can apply them, then asks the container runtime to adjust the running container. The runtime implements applicable limits through mechanisms such as Linux cgroups; the Open Container Initiative runtime specification describes the Linux resource controls available at that execution layer.
The scheduler has to account for a resize that has not finished. While a Pod has a pending or incomplete resize, the scheduler considers the maximum of the desired, allocated, and actual requests. This avoids treating a Pod as if its old, smaller request were still the only resource claim while a larger allocation is being attempted.
The Kubernetes enhancement proposal documents the feature’s API design, safety constraints, and how the kubelet handles requests that cannot be applied immediately.
The kubelet reports progress through Pod conditions:
PodResizePendingwithDeferred: The resize cannot be granted right now but may become possible later, for example after another Pod releases capacity. The kubelet retries it.PodResizePendingwithInfeasible: The request cannot be satisfied on the current node, such as when it asks for more resources than the node can provide.PodResizeInProgress: The kubelet has accepted and allocated the new resources, but the container runtime is still applying them.
These options solve different scaling problems:
| Concern | In-place Pod resize | HPA replica scaling | Deployment rollout |
|---|---|---|---|
| What changes | CPU or memory assigned to an existing Pod | Number of Pod replicas | Pod template and replacement Pods |
| Pod identity | Preserved | New replicas are created or removed | Pods are replaced |
| Application restart | Depends on the container’s resize policy | New replicas start; existing replicas remain | Replacement containers start |
| Main constraint | Node capacity and runtime support | Metrics, replica limits, and cluster capacity | Rollout strategy and available replacement capacity |
| Typical use | Adjust one running workload’s resource allocation | Add or remove application instances | Release a new application configuration or image |
In-place resizing can complement autoscaling, but it does not replace the Horizontal Pod Autoscaler or a node autoscaler. Replica scaling changes how many Pods run; node autoscaling provides more machines when Pods cannot fit. Resizing changes resource values for an existing Pod.
Components and Key Concepts
- Desired and actual resources: A changed
specis the request, not proof of completion. Checkstatus.containerStatuses[].resourcesfor the values currently applied to a running container. The status may lag while a resize is pending or in progress. - Resize subresource: Use the Pod’s
/resizesubresource to update mutable CPU and memory resource fields. Thekubectlclient needs to support the--subresource resizeoption; the Kubernetes documentation specifies client version 1.32 or later for that flag. - Kubelet and container runtime: The kubelet manages Pod lifecycle on a node and coordinates the resize. The runtime and operating system must support applying the requested changes. A running application may still need its own configuration or restart to use newly available memory effectively.
- Requests and limits: Requests influence scheduler placement and resource accounting; limits constrain container use. Raising either value can be blocked by node headroom or policy, and a resize does not move the Pod to a different node.
- Resize policy:
resizePolicyselects whether changing CPU or memory requires a container restart. The default isNotRequired, which attempts to apply the change without restarting.RestartContainerrestarts that container to apply the resource update. If CPU and memory are changed together and either policy requires a restart, the container restarts. - QoS class: A Pod’s Quality of Service class is fixed when it is created. Resource changes must continue to satisfy that class’s rules; for example, a Guaranteed Pod must keep requests equal to limits for CPU and memory.
Only regular application containers and restartable sidecars can be resized; init containers that have completed and ephemeral containers cannot. Requests and limits that were set cannot be removed entirely through a resize. Windows Pods and Pods managed by static CPU or memory manager policies are also among the documented limitations.
Real-World Use Cases
Long-running services may need an operator to increase CPU for a temporary workload phase without draining every instance. A resize can preserve Pod identity, although the operator still needs a rollout or restart plan if the application itself must reload configuration.
Batch and data-processing workers can receive a resource adjustment after the work profile becomes clearer. Operators can avoid replacing a worker solely to adjust its container allocation, while monitoring its progress and service-level behavior.
Right-sizing after observation is another use. If measurements show a container needs more memory or has excessive CPU allocation, an operator can change the running Pod’s resources and observe the effect before updating the workload template.
This is not automatically a feedback controller: Kubernetes does not infer the correct resource values for each Pod. A human, automation system, or higher-level controller must choose the target values and decide how to make that change durable.
Getting Started: Resize a Test Pod
Use a disposable cluster and a standalone Pod to test the resize API. In Kubernetes 1.35 and later, the feature is stable. On earlier versions, check the release-specific feature-gate requirements and ensure the control plane and node support the feature. Use kubectl 1.32 or later for the resize subresource command shown below.
Save this manifest as resize-demo.yaml:
apiVersion: v1
kind: Pod
metadata:
name: resize-demo
spec:
restartPolicy: Always
containers:
- name: app
image: registry.k8s.io/pause:3.10
resources:
requests:
cpu: 500m
memory: 256Mi
limits:
cpu: 500m
memory: 256Mi
resizePolicy:
- resourceName: cpu
restartPolicy: NotRequired
- resourceName: memory
restartPolicy: RestartContainer
Create the Pod and wait for it to run:
kubectl apply -f resize-demo.yaml
kubectl wait --for=condition=Ready pod/resize-demo --timeout=120s
Increase its CPU request and limit without restarting the container:
kubectl patch pod resize-demo --subresource resize --type merge --patch \
'{"spec":{"containers":[{"name":"app","resources":{"requests":{"cpu":"750m"},"limits":{"cpu":"750m"}}}]}}'
kubectl get pod resize-demo -o yaml
Compare the CPU values under spec.containers[].resources and status.containerStatuses[].resources. Also check the container’s restartCount; with the CPU policy set to NotRequired, it should not increase for this resize. If status does not converge, inspect the Pod conditions and events:
kubectl describe pod resize-demo
kubectl get events --sort-by=.lastTimestamp
To test the memory policy, increase both the request and limit so the Pod retains its original Guaranteed QoS class. The configured RestartContainer policy means the container restarts to apply this change:
kubectl patch pod resize-demo --subresource resize --type merge --patch \
'{"spec":{"containers":[{"name":"app","resources":{"requests":{"memory":"384Mi"},"limits":{"memory":"384Mi"}}}]}}'
kubectl get pod resize-demo -o yaml
Treat this as an API and runtime test, not a performance benchmark. Real applications may respond differently to memory changes, and reducing a memory limit while usage is high can be unsafe. Once finished, remove the test Pod:
kubectl delete pod resize-demo
For production, update the workload’s template or its source of truth separately. A direct change to a Deployment-managed Pod does not update the Deployment template, so a later replacement can return to the template’s original values. Confirm node headroom, namespace quotas, Pod conditions, and application behavior before automating changes. For capacity limits at the node layer, see Kubernetes cluster autoscaling and node provisioning.
Common Misconceptions
“In-place means the container can never restart.” The Pod is not replaced, but a container may restart if its resizePolicy is RestartContainer. That choice can be useful when the application cannot safely adapt to a changed memory limit.
“A successful API update means the runtime applied it.” The desired specification may change before the node can grant or apply the resize. Check actual container resources and Pod conditions, not only the output of kubectl patch.
“Resizing is a substitute for replicas or nodes.” It changes the resource allocation of a Pod already assigned to a node. It neither increases replica count nor provisions more node capacity. Use the right control loop for the dimension that needs to change.
“A memory limit reduction is always safe.” A running process can already be using more memory than the target limit. Kubernetes documents memory reduction as best-effort under a no-restart policy, not a guarantee that the application will remain unaffected.
Related Articles
- Linux cgroups: CPU, memory, and container limits explains the kernel resource controls that container runtimes configure.
- Kubernetes cluster autoscaling and node provisioning covers how Pods’ resource requests interact with node capacity.
- KEDA and event-driven Kubernetes autoscaling explains scaling the number of worker Pods from external demand.
- Kubernetes architecture explained introduces the scheduler, kubelet, and other cluster components.
Changelog
- Initial publication; feature behavior and limitations checked against the Kubernetes documentation and enhancement proposal.

