Canary deploy
Also called: canary release, canary, progressive rollout.
Releasing a new version to a small share of real traffic first, for example 10%, while the old version serves the rest. If its error rate and latency look as good as the old version's, the share grows step by step to 100%; if not, all traffic goes back to the old version. A bad release then hurts only the few users who were sent to it.
Each step sends a share of traffic to v2 and watches its error rate for 10 minutes (shortened here). Over 1% errors means an automatic rollback.
Meters: errors measured over each 10-minute step (tick: the 1% limit). The dots are a small sample of the requests.
Traffic: v1 100% · v2 0% · Drawn requests: v1 0, v2 0, failed 0
v1 serves everyone. Release v2 to try it on 10% of the traffic first.
Say it in a prompt
Release the checkout service with Argo Rollouts as a canary: setWeight 10, pause 10 minutes, setWeight 50, pause 10 minutes, then 100. Run a background analysis on Prometheus the whole time: abort and roll back if the canary's 5xx rate is over 1% or its p99 latency is over 800 ms. Vague vs precise prompt
Vague prompt
deploy the new checkout version safely Typical resultSwaps every pod to the new version at once with a rolling update. Health checks pass, so a bug in the payment step reaches 100% of users until someone notices and redeploys by hand.
Precise prompt
Canary with Argo Rollouts: 10% for 10 min, 50% for 10 min, then 100%; abort and roll back if the canary's 5xx rate is over 1% or p99 over 800 ms. Typical resultA bad build reaches only 10% of checkouts for at most 10 minutes before it rolls back by itself, and a good build reaches everyone in about 20 minutes.
Seen on
- Google SRE Workbook: Defines canarying as a partial and time-limited deployment of a change, evaluated against the rest of the service (the control), which keeps running the old version.
- Argo Rollouts docs: A canary rollout is a list of steps such as setWeight (the share of traffic for the new version) and pause; a background analysis that fails aborts the rollout.
You might describe it as
- try the new version on a few users first
- roll out slowly and roll back if errors go up
- send 10% of traffic to the new release
Not to be confused with
- Health check
A canary deploy asks "is the new version as good as the old one for real users?"; a health check asks "is this pod alive and ready?".