Learn / Kubernetes survival kit / Rollouts and rollbacks
Rollouts and rollbacks
How a Deployment replaces Pods safely, readiness probes as the real safety net, and rolling back a bad release fast.
Rolling updates: gradual, not all-at-once
When you change a Deployment’s Pod spec (most commonly, a new container image tag), the default RollingUpdate strategy replaces old Pods with new ones incrementally, not by tearing everything down and starting over. Two settings control the pace:
spec:
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 1 # how many EXTRA Pods above the desired replica count are allowed mid-rollout
maxUnavailable: 0 # how many Pods are allowed to be unavailable mid-rollout
maxUnavailable: 0 with maxSurge: 1 is a common “zero downtime” pattern: Kubernetes creates
one new Pod before removing an old one, so capacity never drops below the desired count,
just briefly exceeds it while the swap happens.
The real safety net: readiness probes
A rolling update is only as safe as Kubernetes’ ability to know a new Pod is actually working. By default, a container is considered “ready” the moment its process starts - which is often well before your application has finished loading config, warming a cache, or connecting to a database. Without a readiness probe, the rollout happily routes live traffic to a Pod that’s still starting up, and requests fail during every deploy.
spec:
containers:
- name: hello
image: myregistry/hello:1.5
readinessProbe:
httpGet:
path: /healthz
port: 8080
initialDelaySeconds: 3
periodSeconds: 5
With this in place, the Service (previous lesson) only routes traffic to Pods that are passing
/healthz - and the rollout controller only considers a new Pod successfully rolled out once
it’s ready, which is also what makes an automatic rollout failure detectable at all.
Watching and controlling a rollout
kubectl rollout status deployment/hello # blocks until the rollout finishes or fails
kubectl rollout history deployment/hello # lists revisions
kubectl rollout pause deployment/hello # stop mid-rollout, e.g. to inspect a partial state
kubectl rollout resume deployment/hello # continue a paused rollout
rollout status is the command to actually use in a deploy script or CI pipeline instead of
polling kubectl get pods and eyeballing it - it exits non-zero if the rollout fails (e.g.
new Pods never become ready), which is exactly the signal a pipeline needs to stop and alert.
Rolling back
Every update to a Deployment’s Pod template creates a new ReplicaSet revision; the old ReplicaSet is kept (scaled to zero) rather than deleted, which is what makes a fast rollback possible:
kubectl rollout undo deployment/hello # back to the previous revision
kubectl rollout undo deployment/hello --to-revision=3 # back to a specific revision
This performs the same rolling-update mechanism in reverse - gradually replacing current Pods
with Pods matching the target revision’s spec, respecting the same maxSurge/maxUnavailable
settings. It’s usually the fastest first response to “the new release is broken” - revert
first with a single command, investigate the root cause after traffic is stable again, not
before.
Key takeaways
- A Deployment's default rolling update replaces old Pods with new ones gradually, controlled by maxSurge and maxUnavailable, instead of an all-at-once replacement.
- A readiness probe is what makes a rollout actually safe - without one, Kubernetes considers a new Pod 'ready' the instant its container starts, even if the application inside hasn't finished booting.
- kubectl rollout status watches a rollout to completion (or failure) instead of guessing from kubectl get pods; kubectl rollout undo reverts to the previous ReplicaSet immediately.
- Every Deployment update is recorded as a new ReplicaSet revision, which is what a rollback actually reverts to - the old ReplicaSet's Pod spec, not a re-deploy of old code from scratch.
Quick check
3 questions - see how much stuck.