Home› Blog› Week 21 - Blue/Green on ECS: The Rollback That Nev…
AWS Weekly Lab AWS Terraform

Week 21 - Blue/Green on ECS: The Rollback That Never Had To Happen

Verified against current vendor documentation on 29 September 2026. Pricing, limits and API behaviour were checked against the official docs on that date. Cloud services change fast — if you are reading this much later, treat the specifics as a starting point and re-check the linked sources.

Why — The Problem This Solves

You have a website running. You want to put a new version of it live. The obvious way is to stop the old one and start the new one, which means a gap where the site is down, and no way back if the new version turns out to be broken except doing the whole thing again in reverse.

Blue/green is the other way. Start the new version alongside the old one, wait until it looks healthy, then point the front door at it. The old version stays running for a while. If something is wrong, you point the door back — no redeploy, no gap.

The half everyone builds is the traffic switch. The half almost nobody tests is the pointing back.

So this week ran a small containerised web service on AWS behind a load balancer, released a new version of it, and then deliberately shipped a broken one — to find out which part of the machinery actually protected users. The answer was not the part built for the job.

What You Need to Know — Skills & Tools

  • Amazon ECS on Fargate — runs containers with no servers to manage
  • Application Load Balancer — the front door; it points at one of two target groups
  • ECS native blue/green — ECS shipped this in July 2025, canary and linear in October 2025. As of March 2026 AWS recommends it over CodeDeploy for new work, so there is no CodeDeploy here
  • Terraform + HCP Terraform — the AWS provider needs 6.4.0 or later for the blue/green block

Architecture — How It Fits Together

One load balancer, two target groups, one service that moves between them.

Client curl / browser Application Load Balancer one listener rule blue target group current version green target group new version ECS service Fargate task sets CloudWatch alarm on 4xx → automatic rollback it never fired, and that is this week's finding During a deployment ECS rewrites the listener rule. The dashed arrow is the switch.

Blue and green are interchangeable slots, not fixed roles. ECS put the first deployment in green, which surprised me — Terraform names blue as the primary.

How We Built It — Step by Step

What we deployed

Sixteen AWS resources. Only the second row exists because of blue/green — everything else you need to run any container behind a load balancer:

WhatCountIts job
Application Load Balancer, listener, listener rule3The front door. ECS rewrites the rule to switch versions
Target groups — blue and green2The two slots a version can occupy. This is the blue/green part
ECS cluster, task definition, service3Runs the container on Fargate; the service carries the deployment strategy
Security groups2Public HTTP to the ALB; the ALB only to the tasks
CloudWatch log group, alarm2Task logs, and the 4xx alarm wired to roll a deployment back
IAM roles + policy attachments4One role pulls the image and writes logs; the other lets ECS rewrite the listener rule

HCP reports 20 rather than 16, because its count includes the four data sources that look up the default VPC and its subnets.

Default VPC on purpose — a NAT gateway costs about $0.045/hr and would have been the biggest line on the bill for containers that only pull a public image.

The application is deliberately trivial: busybox from ECR Public writes an environment variable into a file and serves it. A "new version" is therefore just a variable change, which keeps the deployment itself as the subject rather than a container build.

How we wrote it

The Terraform file tree for this week, showing environments/dev with main.tf, variables.tf and outputs.tf, modules/service with a 347-line main.tf, and two shell scripts, each annotated with its job and line count
Every file, and what it is for. environments/dev says what is deployed; modules/service says how it is built. Changing the version is a variable change in one; changing the architecture is a change in the other.

Of those 347 lines, the part that makes this blue/green rather than a rolling update is nine:

deployment_configuration {
  strategy             = "BLUE_GREEN"
  bake_time_in_minutes = 5      # how long the old version stays alive
}

alarms {
  enable      = true
  rollback    = true            # roll back if the alarm goes off
  alarm_names = [aws_cloudwatch_metric_alarm.errors.alarm_name]
}

Two details that are easy to get wrong. The provider must be 6.4.0 or later; on earlier versions the block is silently ignored and you get a rolling update that looks correct in the file. And ECS switches the listener rule, not the listener, so the rule needs ignore_changes = [action] or the next Terraform plan proposes shifting traffic back.

How we deployed it

HCP Terraform, VCS-driven, OIDC credentials, no static keys.

HCP Terraform run list for week-21-dev showing the applied run and an earlier errored run
Twenty resources. The errored run above it is real — see Challenges.

What appeared in the console

The green target group in the AWS console with one registered task, showing 1 healthy and 0 unhealthy on port 8080
One of the two slots, holding the live version. Its twin sits empty until a deployment fills it. That pair is the whole of blue/green — everything else is ordinary container plumbing.
The ECS service in the console showing deployment strategy blue/green with a five minute bake time
The service carries the strategy, not a separate deployment controller. This is what replaces CodeDeploy.

How we tested it

Two deployments. A normal one, then a broken one. Throughout both, a script polled the load balancer every two seconds and recorded what a caller actually got — because the console shows task sets and percentages, and only a client sees the service.

The ECS console during a deployment showing the new task set in progress alongside the active one
Mid-deployment: both versions running, traffic still on the old one.
Terminal output from the polling script for both deployments: 62 requests served v1 and 86 served v2 with no failures, then 214 requests all served by the old version during the broken deployment
The client's view of both releases. No failed request in either.

What success looked like: the good release cut over at t+182s — 62 requests on v1, 86 on v2, nothing failed. The broken release reached nobody: 214 requests, every one served by the previous version.

Challenges — What Actually Went Wrong

The rollback never happened, and that is the finding

The ECS console showing the new deployment failing, with the target reported unhealthy due to a response code mismatch
Target.ResponseCodeMismatch — "Health checks failed with these codes: [404]".

The broken version failed its health check, so ECS never sent it any traffic. No client request reached it, so there were no errors, so the alarm stayed OK and rolled back nothing. Alarm history confirms no state change during the deployment.

Users were protected — by the health check, not by the mechanism built for the job. If you only ever test rollback with a fault your health check catches, you have tested the health check.

A failing deployment does not give up on its own

deploymentCircuitBreaker is off by default. The failed deployment did not roll back; it retried, cycling tasks for over ten minutes until I reverted the change. Turn the circuit breaker on.

Two small ones

terraform apply returned applied with zero new tasks running and the old version still serving — it hands off to ECS and returns. Gate a pipeline on that and you have gated on nothing.

And the first apply failed on AmazonECSInfrastructureRolePolicyForLoadBalancers: it has no /service-role/ path, unlike AmazonECSTaskExecutionRolePolicy. ECS then reports it as "Unable to assume role and validate the specified targetGroupArn", which sends you looking at target groups instead of at the role.

Security — Controls at Every Layer

  • No static AWS credentials. HCP Terraform assumes a role via OIDC
  • Tasks reachable only through the ALB. The task security group accepts traffic from the load balancer's group and nothing else
  • Scoped IAM. Two AWS managed policies, one per role, no inline wildcards
  • Known gap: tasks run in public subnets with public IPs to avoid NAT costs. Production puts them in private subnets

Cost

ItemRate
Application Load Balancer$0.0225/hr + $0.008/LCU-hr
Fargate$0.04048/vCPU-hr + $0.004445/GB-hr
CloudWatch alarm$0.10/month (first 10 free)

This build has a real hourly meter, and the load balancer is most of it — it bills whether or not anyone visits. About $0.05/hour with both versions running; roughly $16/month for an ALB left up. Ran about three hours.

Billed total: $0.1552. Read from Cost Explorer once AWS posted the charges — Elastic Load Balancing $0.1126 and ECS/Fargate $0.0425 across 29–30 September, with $0.00 on the 1st, which is what a verified teardown looks like on a bill.

The load balancer was 73% of it, for a service that served a few hundred requests. That is the shape worth remembering: the meter is the front door, not the containers behind it.

Cleanup

Destroyed via HCP, then verified independently of the teardown script: no load balancer, no cluster, no target groups, no leftover network interfaces, workspace at zero resources.

bash week-21-bluegreen-ecs/scripts/cleanup.sh
# CLEAN - nothing from Week 21 remains.

Key Takeaways

  • ECS does blue/green itself now. No CodeDeploy for new work — one less service in the failure path
  • Test rollback with a fault your health check cannot see. Otherwise you are testing the health check and learning nothing about the alarm
  • Turn on the deployment circuit breaker. Without it a failing deployment retries instead of reverting
  • The load balancer is the cost. It bills hourly with no traffic at all, so tear it down
  • Watch a deployment from the client's side. The console shows task sets; only a caller shows the service

Comments

How was your experience?
Your feedback helps improve this site.
PoorExcellent