1. The Migration Isn't the Finish Line
We frequently get called in by CTOs three months after a successful AWS migration. The application is running, but the team is stressed: deployments are entirely manual, the monthly AWS bill is triple what was projected, and nobody actually knows what happens if an Availability Zone goes offline.
When you move to AWS, getting the code to run is the easy part. The hard part is transforming that infrastructure from a fragile liability into an observable, scalable platform.
If you've recently migrated, or are planning to, this is the exact 4-week operational playbook we use to establish control.
Week 1 — Establish Visibility
The Goal: Stop flying blind. Understand exactly who is doing what, and what it costs.
Before changing any architecture, you must establish baseline visibility. Most teams skip this step and spend months chasing ghosts.
- AWS Organizations & Accounts: Ensure production, staging, and shared services are isolated in separate AWS accounts under a central Organization.
- CloudTrail & IAM Audit: Enable global CloudTrail logging. Revoke all permanent IAM user credentials and enforce SSO (AWS IAM Identity Center) with least-privilege roles.
- Billing Visibility & Budgets: Enable AWS Cost Explorer. Set hard AWS Budgets with Slack/email alerts at 50%, 80%, and 100% of your projected spend.
- Tagging Strategy: Enforce mandatory tags (
Environment,Service,Owner) across all resources to enable cost allocation. - Identify Unused Resources: Hunt down orphaned EBS volumes, unattached Elastic IPs, and idle EC2 instances left over from the migration cutover.
Week 2 — Fix the Architecture
The Goal: Harden the network and eliminate single points of failure.
The Risk: "Lift and shift" often results in databases running in public subnets or EC2 instances exposed directly to the internet.
Now that you can see the environment, it's time to fix the structural flaws that occurred during the rush to migrate.
- Subnet Isolation: Move all EC2, EKS, and RDS instances strictly into Private Subnets.
- Security Groups: Remove all
0.0.0.0/0ingress rules. Instances should only accept traffic from the ALB/NLB or specific internal security groups. - Load Balancing & WAF: Ensure all public traffic routes through an Application Load Balancer protected by AWS WAF (Web Application Firewall).
- RDS Configuration: Enable automated backups, encryption at rest (KMS), and ensure the database is running in Multi-AZ mode for automated failover.
- NAT Gateway Dependencies: Ensure high-throughput services aren't routing internal AWS traffic (like S3 or DynamoDB requests) out through expensive NAT Gateways. Implement VPC Endpoints.
Week 3 — Make Deployments Boring
The Goal: Eliminate "ClickOps" and manual SSH deployments.
If your team is logging into the AWS Console to change security groups, or SSH-ing into instances to pull git repositories, your infrastructure is a ticking time bomb.
- Infrastructure as Code: Codify the entire baseline (VPCs, RDS, IAM, ALBs) using Terraform or OpenTofu. State files must be locked and version-controlled.
- CI/CD Pipelines: Implement GitHub Actions or GitLab CI to automate testing and artifact generation (e.g., Docker images pushed to ECR).
- Deployment Strategies: Transition from in-place updates to Immutable Infrastructure (Blue/Green or Rolling updates via ECS/EKS).
- Secrets Management: Remove all
.envfiles from instances. Applications must fetch credentials dynamically from AWS Secrets Manager or Parameter Store via IAM Roles. - Reliable Rollbacks: Ensure the team can revert to the previous working version in under 3 minutes without manual code changes.
Week 4 — Optimize Before the Bill Surprises You
The Goal: aggressively right-size infrastructure before the finance team asks questions.
AWS gives you infinite scale, which means it will happily let you spend infinite money. In Week 4, we focus entirely on Cloud FinOps.
- EC2 Right-sizing: Review CloudWatch metrics. If instances are running at 10% CPU, halve their size immediately.
- Graviton Migration: Move RDS and compatible containerized workloads to AWS Graviton (ARM) processors for an instant 20% price/performance gain.
- Spot Instances: Move fault-tolerant, stateless background workers to EC2 Spot Instances (saving up to 80%).
- CloudWatch Logs: Implement retention policies (e.g., 14 days) on all log groups. Infinite retention of debug logs is a common cause of billing spikes.
- Autoscaling: Configure dynamic Target Tracking scaling policies. Infrastructure should scale down to a minimal footprint at 3 AM.
The Production-Readiness Checklist
Before declaring an AWS environment "production ready", we force engineering teams to answer these 8 questions:
| Area | Question to Validate |
|---|---|
| Security | Is IAM strictly least-privilege, and are there zero hardcoded access keys? |
| Networking | Are workloads unnecessarily traversing expensive NAT Gateways for internal services? |
| Reliability | What exactly happens, step-by-step, if us-east-1a (an Availability Zone) goes completely offline? |
| Deployment | Can a junior engineer safely deploy—and more importantly, roll back—in under 5 minutes? |
| Observability | If the ALB throws a 502 Bad Gateway, can you trace the request down to the failing pod/instance in under 60 seconds? |
| Cost | Do we know precisely which microservice or environment is driving the AWS bill? |
| IaC | If the AWS account was deleted, could Terraform recreate the entire environment automatically? |
| Disaster Recovery | When was the last time the RDS snapshot restore process was actually tested? |
What "Production-Ready" Actually Means
Production-ready doesn't mean the application is running. It means the engineering team knows exactly what happens when something goes wrong.
Did your team just migrate to AWS?
If you recently migrated to AWS, we can review your architecture and give you a prioritized production-readiness roadmap. Stop scaling on technical debt.
Book an Infrastructure Assessment
