ECS to EKS Migration and 38% Cloud Cost Cut for a High-Traffic Indian E-Commerce Brand
Monolithic ECS moved to Kubernetes with spot capacity, IaC, and HPA. Cloud spend down 38%. Deployment lead time from four hours to twelve minutes.
cloud spend reduction
deployment lead time (from 4h)
availability during flash sale
A fast-growing Indian fashion e-commerce brand serving roughly 42,000 concurrent shoppers at peak was overspending on AWS by an audited 37%. Deploys took four hours on a monolithic ECS stack.
Their two big-sale events per year required manual capacity planning six weeks in advance. Even then, one of the previous four sales had suffered a 90-minute cart outage.
AUERON's platform engineering team migrated the workload to EKS. Horizontal autoscaling on real request metrics, spot-instance capacity for stateless services, Terraform-managed infrastructure, and a GitOps deploy pipeline.
Cloud spend dropped from roughly ₹73 lakhs to ₹45 lakhs per month. The first flash sale after cutover held 99.95% availability with zero manual capacity planning.
The Problem & Operational Risk
The client's platform ran on AWS ECS with a mix of Fargate and EC2 launch types, provisioned via a mostly-imperative combination of console clicks and a partial CloudFormation template. Scaling policies were CPU-based and slow to react.
During flash-sale spikes, the checkout service was routinely CPU-saturated for eight to twelve minutes before new tasks came online. Carts dropped. Gateway transactions refunded that the customer never intended.
Infrastructure cost visibility was worse than it looked. A "we've got Cost Explorer" answer masked the fact that roughly ₹28 lakhs per month was going to over-provisioned on-demand EC2 capacity kept warm for peak. Five of eight microservices were stuck on x86 instances long after their runtimes had ARM64 support. And NAT gateway egress was the single largest line item at ₹5.2 lakhs, driven by chatty inter-service calls that should never have crossed AZ boundaries.
Every deploy required a ClickOps approval flow through three engineers and averaged four hours end-to-end. Rollbacks were a manual re-deploy of the previous image tag. The engineering team had a running joke about "the Friday freeze," which was really just the accumulated fear of pushing on a day when senior engineers were less available.
CPU-based autoscaling is a symptom, not a signal. HPA on p95 latency reacted 4× faster to real load and eliminated the twelve-minute checkout brownouts.
Engineering Architecture & Solution
AUERON scoped the engagement in four phases.
Phase one wrote the current infrastructure into Terraform without changing behavior. A lift-and-encode step that gave the team ground truth to reason from.
Phase two stood up an EKS cluster with two node groups: an on-demand baseline for stateful services and a spot pool for stateless request-handling. Karpenter provisioned nodes on demand. HPA scaled workloads on p95 latency and request queue depth, not CPU.
Phase three migrated services one at a time behind a shadow-traffic split so behavior could be compared before cutover.
Phase four introduced Argo CD for GitOps, tore down the ClickOps deploy path, and moved image builds into a hardened GitHub Actions pipeline with Trivy scans and signed provenance.
Alongside the platform work, AUERON's SRE team ran a cost-reshaping pass. Moved three stateless services to Graviton (ARM64). Collapsed two AZ crossings that were driving NAT egress. Moved cold storage to S3 Intelligent-Tiering. Reserved a two-year Savings Plan for the baseline capacity that wasn't going anywhere. The Savings Plan alone recovered roughly ₹6 lakhs per month.
Key Architectural Takeaways
- CPU-based autoscaling is a symptom, not a signal. HPA on p95 latency reacted 4× faster to real load and eliminated the twelve-minute checkout brownouts.
- The single largest cost line was NAT gateway egress from unnecessary cross-AZ chatter, not compute. Cost audits should follow the network before the instance sheet.
- ARM64 (Graviton) delivered ~22% price-performance improvement for stateless services at zero code change beyond the Docker base image.
- GitOps with Argo CD removed the human-in-the-loop deploy step without adding risk. The same PR approvals still gate every change, they just also gate the deployment.
- The Friday freeze evaporated within three weeks of cutover once rollback became a single `argocd app rollback` command.
Explore More Case Studies
Fixing a Silent Fraud Model: MLOps Rescue for a Mid-Market Payments Platform
Real-Time Hospital Operations Analytics: 40-Minute Batch to 8-Second Streaming Across 12 Data Sources
Strangling a 380k-Line Django Monolith: Eight Services, 45-Minute Deploys to 6-Minute Deploys
Let's talk
Book your free consultation with an AUERON engineer
One senior engineer will respond within one business day.
Prefer email? hello@aueron.in