Blue-green and canary deployments: make releases safer
A release can start successfully and still break the customer’s checkout. Blue-green and canary deployments reduce particular release risks by controlling where the new version runs and how traffic reaches it. They need reliable checks, compatible data changes and a rehearsed recovery plan before anyone can promise a smooth rollout.

Blue-green keeps an existing and a replacement environment available for a traffic switch. Canary exposes a limited part of traffic or a selected cohort to the new version before wider adoption. Both patterns provide opportunities to detect problems, but neither makes every release safe automatically.
Choose the pattern around the service’s state, traffic and operating requirements. A background worker, a customer-facing API and a mobile client have different release constraints. The aim is to protect meaningful customer operations, not merely keep a deployment indicator green.
01Choose the release pattern for the failure you fear
| Approach | How exposure changes | Risk to plan around |
|---|---|---|
| Blue-green | Switch traffic from the old environment to the prepared one | Shared data changes and safe switchback |
| Canary | Expose a limited cohort, then expand | Representative traffic and sensitive evaluation |
| Rolling update | Replace instances progressively | Old and new versions coexist |
| Feature flag | Control a supported behavior independently of deployment | Flag state does not undo data writes |
For an illustrative ordering service, a blue-green switch can let the team verify a replacement application environment before serving broad traffic. A canary can reveal how a changed pricing operation behaves for a controlled cohort. The decision depends on what can be observed and reversed safely.
List the possible failure modes: startup errors, incorrect business output, slow responses, overload and incompatible messages or data. A container becoming ready addresses only some of those problems. Match each significant risk with a check and a recovery action.
Decide whether routing can hold a user or tenant consistently on one version where needed. Arbitrary request-level splitting may make a stateful journey alternate between implementations. Cohort routing should support the real workflow without turning a small test into an unrepresentative special environment.
Keep the deployment scope explicit. Switching the API does not update old mobile clients, browser assets already cached or independent workers. The cloud migration guide discusses related transition planning; a release should name every component that can coexist during the change.
02Prepare a blue-green environment that is comparable
Build the replacement environment from controlled configuration and an identified artifact. Compare important settings, secrets access and external dependencies with the current environment. A staging success tells little about production when traffic, permissions and service connections differ materially.
Check the actual routing path, including load balancers, DNS or gateway behavior used by the system. Establish how connections drain and how long requests may remain active. A traffic switch is not necessarily an instantaneous move of every session or in-flight operation.
Budget the temporary capacity and operational effort for coexistence. The spare environment must be capable of receiving the intended traffic. Keeping a smaller replacement running successfully under a health probe does not show that it can support the normal workload after the switch.
Prevent unintended duplicated side effects. If both environments run scheduled jobs or message consumers, they may both process the same work unless ownership and deduplication are designed. A replacement environment should not send customer emails merely because it has production connectivity for validation.
Keep the old environment available for the agreed observation period, while understanding its limitations. It can support switchback only if it remains compatible with the state created by the new version. A retained artifact is necessary recovery material, but is not a complete rollback plan.
03Make a canary small enough to limit risk and useful enough to judge
Google’s SRE workbook explains canarying as a release-evaluation approach based on limited exposure and comparison. Use representative traffic and signals relevant to the changed behavior. A quiet canary that receives no meaningful requests provides little evidence.
Choose the cohort with a reason. Internal accounts can validate functionality, but may not reflect real permissions, data sizes or user behavior. Consider tenant boundaries and the commercial impact of exposure. Define who is included and how the team can identify affected operations.

Compare the canary with a suitable baseline using consistent metrics. Inspect errors and latency for the new version, but also business outcomes such as accepted orders and correct calculations. Aggregate service success can conceal a failure concentrated in the exposed group.
Choose the observation conditions before starting. Traffic volume, expected operation frequency and delayed failures affect how long evaluation needs to run. Do not adopt a universal small percentage or fixed duration as proof of safety for every application.
Define promotion, pause and abort decisions with named ownership. Automated evaluation can help, provided signals and thresholds are appropriate. If a key metric is missing or unreliable, the rollout should not promote merely because no alarm happened to fire.
04Keep database and message changes compatible
AWS’s blue-green whitepaper discusses compatible schema changes during coexistence. Its deployment principles are useful here; current service capabilities should be checked separately. Plan readers, writers and serialized data together. A version rollback can fail when the older code cannot understand state already written by the new release.
Use a staged migration where appropriate: add a compatible structure, deploy code that can work during the transition, move data with validation and remove the old structure only after the compatibility window closes. Each phase should have an owner and an acceptance condition.
For the ordering service, renaming a required field and deleting the old one in a single step can break old workers. A transition that supports both forms gives the team room to change readers and writers safely. Document when the old representation is no longer needed.
Review queue messages, cached records and public API contracts as well as database tables. A delayed message created before the rollout can reach a new worker, while an older worker may receive a message from the new API. Mixed-version behavior should be a deliberate supported state.
Do not use database restoration as a casual substitute for application rollback. Restoring an older snapshot can remove legitimate transactions created after it. Recovery of data and reversal of code have different consequences, which need a business-approved plan and reconciliation.
05Test readiness, rollout and rollback separately
Readiness asks whether an instance should receive traffic. Rollout evaluation asks whether the release works acceptably. Recovery asks whether the team can restore a supported operating state. Keep those questions separate so a simple health endpoint does not become an unsupported safety claim.
Kubernetes Deployment documentation describes rolling updates and rollout status. It reports a stalled deployment condition but does not automatically roll back solely because the progress deadline is exceeded. Any automatic recovery must come from a configured higher-level process.
- Build and identify the artifact and compatible migration steps.
- Validate the new environment’s configuration and meaningful workflows.
- Exercise mixed-version reads, writes and message handling.
- Rehearse routing reversal and the handling of in-flight work.
- Confirm that monitoring detects the failure scenarios being guarded against.
Use controlled failure exercises suited to the system. A slow dependency, failed authorization check or incompatible response can test whether the team notices the right problem. Record the evidence rather than declaring the process safe because the normal rollout completed once.
Keep recovery permissions and artifacts ready before broad exposure. A rollback instruction is of limited value when no on-call person can access the router or retrieve the prior configuration. Reduce the number of improvised actions required while customers are already affected.
06Operate a release with a clear stop decision
Publish a concise release record: artifact, changed behavior, migration state, cohort, evaluation signals, owner and recovery limits. This gives support and operations enough context to investigate a problem without reconstructing the release from scattered chat messages.
Expand exposure only when the agreed evidence supports it. Continue checking the stages after the first promotion, because a larger workload can reveal a new failure. Preserve version labels and cohort context in observability so the behavior remains attributable during coexistence.

When a guardrail fails, stop the rollout and follow the defined recovery path. Assess customer impact and queued operations before switching traffic blindly. A reversal that repeats payments or discards new records can create a second incident even if the old API starts responding normally.
Review the release after completion and retire temporary resources deliberately. Remove outdated flags and compatibility code only when their supported window ends. Keeping every old environment indefinitely is costly, while deleting it immediately can remove a needed recovery option.
Connect releases to the wider support process, including cloud security responsibilities. Routing, secrets, permissions and logging remain operational decisions. A safer deployment pattern helps only when the team can maintain the controls that make it work.
07Questions about blue-green and canary releases
Do these patterns guarantee zero downtime?
No. They reduce particular risks, but dependencies, state changes, routing and capacity can still cause disruption. Verify the customer operation and recovery limits.
Is canary always better than blue-green?
Choose according to the change and evaluation capability. Canary supports limited exposure; blue-green supports prepared environments and traffic switching. They can be combined.
What percentage should a canary use?
There is no universal safe percentage. Limit impact while ensuring enough representative traffic for the behavior being evaluated.
Can a feature flag replace rollback?
It can control supported behavior, but it does not necessarily reverse data writes, messages or side effects. Plan those consequences separately.
Does Kubernetes automatically undo a stalled rollout?
A Deployment reports the stalled condition. Automatic rollback requires a suitable additional process; do not assume the progress deadline alone performs recovery.
Can I restore the database to undo a release?
Restoration can remove valid later transactions. Treat data recovery as a separate plan with business consequences and reconciliation, not an ordinary code rollback.
What should be in the release record?
Identify the artifact, migration state, exposure cohort, checks, decision owner and recovery constraints. Keep enough context to operate and investigate the transition.