Home› Blog› AWS Daily Intelligence #46 - The most cautious dep…
AWS Daily Intelligence AWS

AWS Daily Intelligence #46 - The most cautious deployment the API accepts cannot finish

Verified against current vendor documentation on 3 October 2026. Pricing, limits and API behaviour were checked against the official docs on that date. Cloud services change fast — if you are reading this much later, treat the specifics as a starting point and re-check the linked sources.

Executive summary

ECS now offers blue/green, linear and canary deployments for services using VPC Lattice, “enabled for both new and existing ECS services in all AWS Regions where VPC Lattice is available.”

The capability itself is unsurprising and welcome: Lattice joins ALB, NLB and Service Connect as a traffic-shifting resource, so the same three primitives apply — a target group, a listener and a rule are each “an Elastic Load Balancing or VPC Lattice resource”. If your service-to-service traffic already runs through Lattice, you no longer have to put a load balancer in front of it just to deploy safely.

What is worth the post is the configuration surface underneath, because two of its documented ranges multiply into a third documented limit and overshoot it.

What changed

Four strategies now, “ROLLING | BLUE_GREEN | LINEAR | CANARY”, with the newer three differing only in how traffic moves:

StrategyTraffic movementConfig object
BLUE_GREEN All at once, after a bake bakeTimeInMinutes — “You must provide this parameter”
LINEAR Equal increments with a wait between each linearConfiguration — only valid for LINEAR
CANARY “a small percentage… then shifts the remaining traffic all at once” canaryConfiguration — only valid for CANARY

And Lattice slots in at the testing phase as well as the shifting phase: “Your traffic-shifting resource (an Application Load Balancer, Network Load Balancer, Service Connect, or VPC Lattice) directs test requests to the green environment while production traffic remains on blue.” The test-traffic path is the part people under-use, and it is available before a single production request moves.

Architecture

Diagram: how an ECS linear deployment moves traffic, and the three timeouts that bound it. Amazon ECS now supports blue/green, linear and canary deployment strategies for services using Amazon VPC Lattice, enabled for both new and existing services in all Regions where VPC Lattice is available, with the valid strategy values being ROLLING, BLUE_GREEN, LINEAR and CANARY. The traffic-shifting resource can be an Application Load Balancer, a Network Load Balancer, Service Connect, or VPC Lattice, and a target group, a listener and a rule are each either an Elastic Load Balancing or a VPC Lattice resource. A linear deployment runs through eleven lifecycle stages: RECONCILE_SERVICE, PRE_SCALE_UP, SCALE_UP, POST_SCALE_UP, TEST_TRAFFIC_SHIFT, POST_TEST_TRAFFIC_SHIFT, PRE_PRODUCTION_TRAFFIC_SHIFT, PRODUCTION_TRAFFIC_SHIFT, POST_PRODUCTION_TRAFFIC_SHIFT, BAKE_TIME and CLEAN_UP. Hook support is uneven: SCALE_UP, BAKE_TIME and CLEAN_UP support no lifecycle hooks at all, while TEST_TRAFFIC_SHIFT and PRODUCTION_TRAFFIC_SHIFT support Lambda hooks only and therefore cannot take a pause hook, which means the moment traffic is actually moving cannot be gated on a human. Hooks configured for PRODUCTION_TRAFFIC_SHIFT or PRE_PRODUCTION_TRAFFIC_SHIFT are invoked at every production traffic shift step. The configuration has two independent dials: step percent, which takes a Double with valid values from three point zero to one hundred, and step bake time, with valid values from zero to one thousand four hundred and forty minutes; the last step bake time is skipped once traffic has shifted one hundred percent. Three timeouts bound the result: each lifecycle stage and each traffic shift step may last up to twenty-four hours before the system times out, fails the deployment and initiates a rollback; CloudFormation enforces a thirty-six hour limit on the entire deployment; and the overall deployment timeout is thirty days, with pause hooks configurable up to twenty thousand one hundred and sixty minutes or fourteen days. The central finding is that the two dials at their documented extremes exceed the third limit: a three percent step needs thirty-four steps and therefore thirty-three bake periods, and thirty-three periods of one thousand four hundred and forty minutes is forty-seven thousand five hundred and twenty minutes, or thirty-three days, against a thirty day overall deployment timeout, so the most cautious linear deployment the API accepts times out before it completes. Working backwards, thirty days across thirty-three bakes allows about one thousand three hundred and nine minutes per step rather than the documented one thousand four hundred and forty, and CloudFormation's thirty-six hours across thirty-three bakes allows only about sixty-five minutes per step. A closing note records that the deployment circuit breaker can only be used for services using the rolling update ECS deployment type, so rollback for these strategies comes from CloudWatch alarms and from the stage timeouts rather than from the circuit breaker, and that linear deployments temporarily run both revisions simultaneously, which may double resource usage during a deployment.
Two independent dials, three timeouts, and one combination the API accepts but cannot complete.

Two dials, documented independently

“Step percent — The percentage of traffic to shift in each increment during a linear deployment. This field takes Double for value, and valid values are from 3.0 to 100.0.”

“Step bake time — The duration to wait between each traffic shift increment during a linear deployment. Valid values are from 0 - 1440 minutes.”

Both are legitimate. A 3% step is the most gradual shift the API will take, and 1,440 minutes is a full day of observation between increments — which is exactly what someone deploying something frightening would reach for.

And three timeouts, on the same page

“Each lifecycle stage can last up to 24 hours and in addition each traffic shift step in PRODUCTION_TRAFFIC_SHIFT can last upto 24 hours… The system times out, fails the deployment, and then initiates a rollback after a stage reaches 24 hours.”

“While the 24-hour stage limit remains in effect, CloudFormation enforces a 36-hour limit on the entire deployment.”

“For pause hooks, you can configure the timeout up to 20,160 minutes (14 days). The overall deployment timeout is 30 days.”

Multiply the dials and you land outside the third limit

A 3% step needs 34 increments to reach 100%, and the waits sit between them — “Note, that last step bake time is skipped once traffic is shifted 100.0%” — so 33 bake periods.

At the maximum step bake time that is 33 × 1,440 = 47,520 minutes, which is 33 days. The overall deployment timeout is 30 days.

So the most cautious linear deployment the API will accept cannot complete. Every individual value is inside its documented range, each stage is inside its 24-hour cap, and the deployment still dies three days short of the end — and what it does on timeout is roll back, which at that point means discarding a revision that has been serving a majority of production traffic for weeks.

ConstraintBake periodsMax step bake time that fits
Documented range331,440 min — does not fit
30-day overall timeout33~1,309 min (21.8 h)
CloudFormation, 36 h total33~65 min

The CloudFormation row is the one that will actually bite, because most people deploy ECS services from a template rather than from the API. Thirty-six hours for the entire deployment, spread across 33 increments, is about 65 minutes of observation per step. Set 1,440 and not even two increments fit inside the limit.

Business value

Two things, and the smaller one is the cleaner win. If your service-to-service traffic already runs through Lattice, you have until now needed a load balancer in front of it purely to get a managed deployment strategy — an extra resource, an extra hop and an extra bill that existed for the benefit of the deploy rather than the request path. That resource can go.

The larger one is that traffic shifting becomes available on the substrate where cross-VPC and cross-account calls already live. The strategies are the same ones ECS offers elsewhere, so this is not a new capability to learn — it is an existing capability reaching the place a service mesh put your traffic.

Security considerations

Nothing here changes who can call what; the authorisation model for Lattice is unchanged. What changes is the exposure window during a release, and it changes in a good direction: the test-traffic path runs first and separately, so a new revision can be exercised end-to-end before a single production request reaches it. “Your traffic-shifting resource… directs test requests to the green environment while production traffic remains on blue.”

The thing worth gating deliberately is the pause hook. A pause hook is a human holding a deployment open, and it waits for someone to call ContinueServiceDeployment — so whoever holds that permission decides when a release proceeds. Scope ecs:ContinueServiceDeployment like the approval step it is, not like a read-only deployment action.

And one failure mode is a security matter rather than a reliability one: a linear deployment with no alarms configured will shift to 100% on a revision that is serving errors, so long as its tasks are healthy. If the new revision has a broken authorisation path, task health will not notice.

Lifecycle hooks are uneven, and unevenly in the places you want them

Eleven stages, and the hook support column is the thing to read before designing a gate:

StageHook supportWhat it means
PRE_SCALE_UP, POST_SCALE_UPYesLambda or pause, before any traffic moves
SCALE_UPNoGreen launching tasks; nothing can gate it
TEST_TRAFFIC_SHIFTLambda onlyNo pause hook while test traffic migrates
PRE_PRODUCTION_TRAFFIC_SHIFTYes“invoked at every traffic shift step”
PRODUCTION_TRAFFIC_SHIFTLambda onlyNo pause hook while production traffic moves
BAKE_TIME, CLEAN_UPNoNo gate before the blue revision is torn down

The pattern is consistent: you can gate around a traffic shift but not during one. A pause hook — which “pause[s] the deployment and wait[s] for you to call ContinueServiceDeployment to proceed” — belongs on PRE_PRODUCTION_TRAFFIC_SHIFT, and that one fires “at every production traffic shift step”.

Which is worth doing the arithmetic on too: a pause hook on that stage at 3% steps means 34 manual approvals for one deployment. It is a per-step hook, not a per-deployment one.

And BAKE_TIME and CLEAN_UP supporting no hooks means there is no programmatic last look before the old revision scales to zero. The bake time itself is the only thing standing between a bad release and an irreversible one.

What rolls this back is not the circuit breaker

The announcement mentions circuit-breaker integration, and the API reference is narrower about where the circuit breaker applies: “The deployment circuit breaker can only be used for services using the rolling update (ECS) deployment type.”

For the three traffic-shifting strategies the rollback path is the alarms configuration plus the stage timeouts — the monitoring phase is described as “Monitor application health, performance metrics, and alarm states during the bake time period. A rollback operation is initiated when issues are detected”, and the 24-hour stage cap “times out, fails the deployment, and then initiates a rollback” on its own.

Worth being precise about, because the two mechanisms behave differently. The circuit breaker reacts to tasks failing to reach steady state; CloudWatch alarms react to whatever you told them to watch. A linear deployment with no alarms configured has no automatic failure detection during traffic shifting at all — it will shift all the way to 100% on a revision that is serving errors, as long as its tasks are healthy.

That is the same shape as #37, which found that ECS rollback targets the most recent COMPLETED deployment and can strand itself when there isn't one. Deployment safety here is a set of things you configure, not a property you get.

Cost considerations

No charge is stated for the capability, and the real cost is capacity: “Linear deployments temporarily run both the blue and green service revisions simultaneously, which may double your resource usage during deployments.”

Combine that with the timing arithmetic and the bill becomes a function of your caution. A 3% step with a one-hour bake is about 33 hours at double capacity. The same service deployed blue/green with a 30-minute bake is about half an hour. Choosing the gentlest strategy is choosing to run two copies of production for a day and a half.

That is the trade to state plainly rather than discover: linear deployments buy confidence with capacity and wall-clock time, and both scale with the number of steps.

Operational considerations

The deployment now has a duration you have to plan around. A rolling update finishes in minutes. A linear deployment with 33 increments and an hour between each runs for a day and a half, with both revisions live throughout. That changes what a change-freeze window means and what "deploy on Friday" costs.

Stage timeouts fail closed, and they fail late. Each stage caps at 24 hours and “times out, fails the deployment, and then initiates a rollback”. A pause hook nobody answers therefore does not sit there indefinitely — it burns a day and then reverts, which is the right default and a surprising one if you expected it to wait.

Know which path deploys the service. CloudFormation's 36-hour ceiling on the whole deployment is far tighter than the 30-day API ceiling, so the same configuration behaves differently depending on whether a template or a direct API call drove it. That is the single most important operational fact in this release.

Pause-hook timeouts are generous enough to be dangerous. Configurable “up to 20,160 minutes (14 days)” against a 24-hour stage cap — so a pause hook can be configured to wait far longer than the stage it sits in will survive.

Tradeoffs

StrategyBuys youCosts you
ROLLING Speed, and the deployment circuit breaker No traffic-percentage control; no test-traffic phase
BLUE_GREEN One bake, one cutover, fast rollback Full exposure the moment traffic moves
CANARY A small exposed population, then one cutover Only two exposure levels, so a problem that scales with load may not appear in the canary
LINEAR Graduated exposure with a gate between each step Double capacity for the whole run, wall-clock measured in hours or days, and a product that has to fit inside three timeouts

The axis nobody states is that caution costs capacity multiplied by time, and linear deployments scale both with the number of steps. A 3% step is not 33 times safer than a 100% step; it is 33 times longer and twice as expensive throughout.

Implementation guidance

Pick the step percent from your traffic volume, not from your nerves. The point of a small step is to expose the new revision to enough requests to see a problem. If 3% of your traffic is forty requests a minute, a 3% step tells you almost nothing and costs you 33 waits.

Then check the product against your deployment path. Steps × step bake time has to fit inside 36 hours if CloudFormation drives the deployment, and inside 30 days otherwise. There is nothing that warns you at configuration time — both values are individually valid.

Configure alarms. Without them, traffic shifting has no automatic abort, and the circuit breaker is not available for these strategies.

Put pause hooks on PRE_PRODUCTION_TRAFFIC_SHIFT only if you mean it per step. For a single human gate, use POST_TEST_TRAFFIC_SHIFT — after the green revision has taken 100% of test traffic and before any production request moves.

If you already run Lattice, drop the load balancer you added for deployments. That is the actual win in this announcement: the traffic-shifting primitives are now the same three objects on either substrate.

Best practices

Derive the step percent from request volume, not from anxiety. A step is only a test if enough traffic crosses it to surface a fault. Below a few hundred requests per increment you are paying in hours for a sample that cannot tell you anything.

Compute steps × step bake time and compare it against your deployment path's ceiling — 36 hours under CloudFormation, 30 days otherwise. Make it a check in CI if you template ECS services, because nothing in the API will do it for you.

Always configure alarms. For these strategies they are the only automatic abort; the circuit breaker is not available. An alarm on target-group 5xx and on the service's own error metric is the minimum that makes graduated exposure mean anything.

Put the human gate at POST_TEST_TRAFFIC_SHIFT, not at PRE_PRODUCTION_TRAFFIC_SHIFT. The first is one approval after the green revision has taken all test traffic; the second fires at every step.

Use Lambda hooks for the per-step checks and keep them fast. TEST_TRAFFIC_SHIFT and PRODUCTION_TRAFFIC_SHIFT accept Lambda only, so automated validation is the only thing that can run while traffic is actually moving.

Drop the load balancer you added only to enable deployments once the Lattice path is proven. That is the concrete saving in this release.

Who should adopt this

Anyone running ECS services behind VPC Lattice can now use the managed strategies directly, on existing services, with no migration.

Anyone already using LINEAR with a small step and a long bake should multiply the two numbers today and compare against 36 hours if the service is deployed by CloudFormation. This is the item here that can already be configured wrong.

Anyone who added an ALB purely to get blue/green on a Lattice-fronted service can remove it.

Key takeaways

Lattice becoming a traffic-shifting resource is a clean, uncomplicated addition, and it removes a load balancer that existed only to enable deployments.

The thing to carry away is about limits rather than about Lattice: stepPercent and stepBakeTimeInMinutes are documented independently, and their product is bounded by a third number documented further down the same page. At the extremes the API accepts a configuration that runs 33 days against a 30-day ceiling, and under CloudFormation the real ceiling is 36 hours for everything.

Which is the same failure the last two posts found from different directions — today's #71 is about defaults that are individually reasonable and collectively expensive. Here the values are individually valid and collectively impossible. Nothing checks the combination but you.

Official AWS references

Comments

How was your experience?
Your feedback helps improve this site.
PoorExcellent