AWS ECS Fargate Deployment¶
Deploy Django Keel with serverless containers using AWS ECS Fargate and Terraform.
Overview¶
AWS ECS Fargate provides serverless container deployment with:
- No server management - AWS manages infrastructure
- Auto-scaling - Scale based on CPU/memory/requests
- High availability - Multi-AZ deployment
- AWS integration - RDS, S3, CloudWatch, Secrets Manager
- Terraform IaC - Infrastructure as code
- Load balancing - Application Load Balancer included
Prerequisites¶
- AWS account with appropriate permissions
- AWS CLI configured (
aws configure) - Terraform installed (v1.5+)
aws-ecs-fargateselected indeployment_targetsduring generation
Quick Start¶
You can run everything with the generated deploy/ecs/deploy.sh script, which checks dependencies, builds and pushes the image, and runs Terraform for you. The steps below show the manual flow.
1. Configure Terraform Variables¶
Edit terraform.tfvars:
# Project Configuration
project_name = "your-project"
environment = "production"
aws_region = "us-east-1"
# Container Configuration
container_cpu = 512 # 0.5 vCPU
container_memory = 1024 # 1 GB
# Django Configuration
django_secret_key = "<generated-secret>"
django_debug = false
allowed_hosts = "yourdomain.com"
# Auto-scaling Configuration
desired_count = 2
min_capacity = 2
max_capacity = 10
# Database Configuration
use_aurora_serverless = false
db_instance_class = "db.t3.micro"
db_allocated_storage = 20
db_backup_retention_period = 7
db_deletion_protection = true
# Redis / Celery
enable_redis = true
redis_node_type = "cache.t3.micro"
enable_celery = true
celery_worker_count = 2
# Domain Configuration (optional)
domain_name = "" # Leave empty for ALB URL only
create_dns_zone = false
See variables.tf for the full list (S3 buckets, monitoring, Sentry, Stripe, spot capacity, tags).
2. Initialize Terraform¶
3. Plan Deployment¶
Review the resources that will be created: - VPC with public/private subnets - RDS PostgreSQL database - ElastiCache Redis cluster - ECS cluster and service - Application Load Balancer - CloudWatch log groups - IAM roles and policies
4. Deploy¶
Type yes to confirm. Deployment takes ~10 minutes.
Your app will be available at the ALB DNS name (output after apply).
Architecture¶
Internet
↓
Application Load Balancer (ALB)
↓
ECS Service (Fargate)
├── Task 1 (Private Subnet AZ-A)
├── Task 2 (Private Subnet AZ-B)
└── Task 3 (Auto-scaled)
↓
RDS PostgreSQL (Multi-AZ)
ElastiCache Redis
S3 (Media/Static files)
Project Structure¶
Generated when aws-ecs-fargate is in deployment_targets:
deploy/ecs/
├── deploy.sh # One-command deployment script
├── README.md
└── terraform/
├── main.tf # Main configuration & providers
├── variables.tf # Input variables
├── outputs.tf # Output values
├── network.tf # VPC, subnets, NAT gateways
├── ecs.tf # ECS cluster, services, auto-scaling
├── alb.tf # Application Load Balancer
├── database.tf # RDS PostgreSQL / Aurora Serverless
├── storage.tf # S3 buckets and ElastiCache Redis
├── security.tf # Security groups, secrets, Route53
├── ecr.tf # Container registry
├── monitoring.tf # CloudWatch alarms and dashboard
└── terraform.tfvars.example # Example variables file
Terraform Configuration¶
The configuration is a set of flat .tf files (no modules):
network.tf¶
- VPC with CIDR block
- Public subnets (ALB)
- Private subnets (ECS tasks, RDS)
- Internet Gateway
- NAT Gateways
- Route tables
database.tf¶
- RDS PostgreSQL or Aurora Serverless v2 (
use_aurora_serverless) - Automated backups (
db_backup_retention_period, default 7 days) - Encryption at rest
- Security group (private subnets only)
ecs.tf¶
- ECS cluster
- Task definitions (web, and Celery when enabled)
- Services with desired count
- Auto-scaling policies (CPU, memory, request count)
- CloudWatch logs
alb.tf¶
- Application Load Balancer
- Target group (ECS tasks)
- Health checks (
/health/) - HTTPS listener when a custom domain is configured
- HTTP → HTTPS redirect
Environment Variables¶
Sensitive values are set in terraform.tfvars and stored in AWS Secrets Manager (security.tf). The task definitions inject django_secret_key (as DJANGO_SECRET_KEY) and the database URL (as DATABASE_URL). The sentry_dsn and stripe_secret_key secrets are created in Secrets Manager but are not wired into the task definitions by default; add them to the task's secrets in ecs.tf if your app needs them. Non-sensitive settings are passed as plain environment variables from tfvars: allowed_hosts as DJANGO_ALLOWED_HOSTS and django_debug as DEBUG. These names match what config/settings reads.
Deploying Updates¶
1. Build and Push Image¶
# Login to ECR
aws ecr get-login-password --region us-east-1 | \
docker login --username AWS --password-stdin \
123456789.dkr.ecr.us-east-1.amazonaws.com
# Build image
docker build -t your-project:latest .
# Tag for ECR
docker tag your-project:latest \
123456789.dkr.ecr.us-east-1.amazonaws.com/your-project:latest
# Push
docker push 123456789.dkr.ecr.us-east-1.amazonaws.com/your-project:latest
2. Update ECS Service¶
# ECS automatically detects new image
aws ecs update-service \
--cluster your-project-production-cluster \
--service your-project-production-app \
--force-new-deployment
Or use Terraform directly (flat config, no modules):
Auto-Scaling¶
Configure in terraform.tfvars:
ecs.tf sets up target-tracking policies on CPU utilization (70%), memory utilization (80%), and ALB request count per target. Adjust the target values in ecs.tf if needed.
Database Migrations¶
Run migrations before deploying new tasks:
# Option 1: ECS Exec into running task
aws ecs execute-command \
--cluster your-project-production-cluster \
--task <task-id> \
--container your-project \
--command "python manage.py migrate" \
--interactive
# Option 2: Run the dedicated migrate task (it already runs
# `python manage.py migrate --noinput` — no command override needed).
# Terraform exposes this task-definition name as the `ecs_migrate_task_definition` output.
aws ecs run-task \
--cluster your-project-production-cluster \
--task-definition your-project-production-migrate \
--launch-type FARGATE \
--network-configuration '{
"awsvpcConfiguration": {
"subnets": ["subnet-abc123"],
"securityGroups": ["sg-abc123"]
}
}'
Background Workers (Celery)¶
A Celery worker service is already defined in ecs.tf. Enable it via tfvars:
This creates a separate Fargate service running celery -A config worker using the same container image and secrets as the web service.
Monitoring¶
CloudWatch Logs¶
View logs:
# All tasks share one log group; streams are prefixed per task
# (app, celery-worker, celery-beat, migrate)
aws logs tail /ecs/your-project-production --follow
# Filter to just the celery worker streams
aws logs tail /ecs/your-project-production --log-stream-name-prefix celery-worker --follow
CloudWatch Metrics¶
Monitor in AWS Console: - ECS → Clusters → your-project → Metrics - Key metrics: CPU, Memory, Request count
Alarms¶
Terraform creates alarms for: - High CPU utilization (> 80%) - High memory utilization (> 80%) - ECS service unhealthy tasks - RDS high connections
Custom Domain¶
HTTPS is only available with a custom domain. Set it in tfvars:
# terraform.tfvars
domain_name = "yourdomain.com"
create_dns_zone = true # Let Terraform manage a Route53 hosted zone
Terraform requests an ACM certificate and, when create_dns_zone = true, validates it automatically via Route53 and points DNS at the ALB. If your DNS lives elsewhere, set create_dns_zone = false, create the ACM validation records shown in the outputs at your DNS provider, and point your domain at the ALB:
Disaster Recovery¶
Database Backups¶
RDS automatic backups: - Retention: 7 days (configurable) - Automated snapshots daily - Point-in-time recovery available
Manual snapshot:
aws rds create-db-snapshot \
--db-instance-identifier your-project-db \
--db-snapshot-identifier manual-backup-$(date +%Y%m%d)
Restore from Backup¶
aws rds restore-db-instance-from-db-snapshot \
--db-instance-identifier your-project-db-restored \
--db-snapshot-identifier manual-backup-20250109
Update Terraform to point to restored database.
Cost Optimization¶
Estimated Monthly Costs¶
Minimal setup (1 task, db.t4g.micro, 1GB traffic): - ECS Fargate: ~$15 - RDS db.t4g.micro: ~$15 - ALB: ~$20 - NAT Gateway: ~$35 - ElastiCache (cache.t4g.micro): ~$12 - Total: ~$97/month
Optimization Tips¶
- Use Spot instances for non-production
- Scheduled scaling - Scale down at night
- Use S3 for static files - Cheaper than EFS
- Single NAT Gateway for non-prod (not HA)
- RDS reserved instances - Up to 60% savings
- CloudFront CDN - Reduce ALB data transfer costs
Troubleshooting¶
Task Keeps Restarting¶
# Check task stopped reason
aws ecs describe-tasks \
--cluster your-project-production-cluster \
--tasks <task-id>
# Check logs
aws logs tail /ecs/your-project-production --since 1h
Common causes:
- Health check failing (/health/ endpoint)
- Database connection issues
- Missing environment variables
- Out of memory (increase container_memory)
Database Connection Timeout¶
# Check security group rules
aws ec2 describe-security-groups \
--group-ids sg-abc123
# Verify RDS is in same VPC
aws rds describe-db-instances \
--db-instance-identifier your-project-db
High Costs¶
# Analyze costs
aws ce get-cost-and-usage \
--time-period Start=2025-01-01,End=2025-01-31 \
--granularity MONTHLY \
--metrics BlendedCost \
--group-by Type=SERVICE
Check: - NAT Gateway data transfer - ALB idle time - Underutilized RDS instances
Production Checklist¶
- [ ] Enable RDS Multi-AZ
- [ ] Configure automated backups
- [ ] Set up CloudWatch alarms
- [ ] Enable ECS Container Insights
- [ ] Configure WAF rules on ALB
- [ ] Set up VPC Flow Logs
- [ ] Enable AWS Config for compliance
- [ ] Use Secrets Manager for sensitive data
- [ ] Configure auto-scaling policies
- [ ] Set up CI/CD pipeline (GitHub Actions/GitLab CI)
- [ ] Document runbooks for common issues
- [ ] Test disaster recovery procedures
Best Practices¶
- Use Terraform workspaces for multiple environments
- Store state in S3 with DynamoDB locking
- Tag all resources for cost tracking
- Use IAM roles (not access keys) for ECS tasks
- Enable encryption for RDS, S3, secrets
- Implement blue/green deployments for zero downtime
- Use VPC endpoints to reduce NAT costs
- Regular security audits with AWS Trusted Advisor
Further Reading¶
Support¶
- AWS Support: Available in AWS Console
- Terraform Community: discuss.hashicorp.com
- Stack Overflow: aws-ecs