Skip to main content

Command Palette

Search for a command to run...

Desaster Recovery Strategy for Austrix Advisory

Updated
•37 min read•View as Markdown
A

Hi, I'm Aditya — a Cloud & DevOps Engineer passionate about automating everything from CI/CD pipelines to multi-cloud infrastructure. I specialize in AWS, Kubernetes, Terraform, and GitOps tools like Argo CD. I’ve helped teams scale applications, cut cloud costs by 90%, and build disaster-ready infra. I love sharing real-world DevOps lessons, cloud cost optimization tips, and infrastructure design patterns. Let’s connect and simplify the cloud — one YAML file at a time ☁️⚙️

Also give For production we use EKS and for dr we use EKS fargate compare both cases and which is best and followup questions for our fintceh wealth advisory app

ChatGPT said:

Got it. Let’s break this down carefully for your fintech wealth advisory platform, comparing EKS on EC2 (production) vs EKS Fargate (DR), then highlight follow-up questions that interviewers might ask.


1️⃣ EKS on EC2 (Production)

  • Node type: EC2 instances (e.g., t3.medium or m5.large)

  • Pros:

    • Full control over the nodes (OS, instance type, custom AMIs)

    • Can run stateful workloads with persistent storage

    • Cost-effective for consistent workloads due to reserved instances or autoscaling

    • Easier to integrate with monitoring, security agents, and sidecar containers

  • Cons:

    • Need to manage node scaling, patching, and lifecycle

    • Requires proper cluster autoscaler setup

    • Slightly more operational overhead

Why we use for production:

  • Production workloads are critical, high-performance, and require consistent control and monitoring. EC2 nodes give predictable performance and allow fine-tuned resource allocation for 15+ core microservices.

2️⃣ EKS Fargate (Disaster Recovery / DR)

  • Node type: Serverless, managed by AWS

  • Pros:

    • No need to manage EC2 nodes

    • Automatically scales pods as needed

    • Isolated compute per pod → better security and resource isolation

    • Ideal for DR where workload is not continuously running

  • Cons:

    • Slightly higher cost for always-on workloads

    • Limited control over OS-level configurations

    • Some performance-sensitive or stateful workloads might be harder to run

  • Why we use for DR:

    • DR cluster doesn’t run full-time; Fargate makes it operationally simpler to stand up on-demand if primary region fails.

    • Security isolation helps ensure DR cluster is protected from accidental misconfigurations.


3️⃣ Direct Comparison Table

FeatureEKS on EC2 (Production)EKS Fargate (DR)
ControlHigh (full OS/node control)Limited (AWS manages nodes)
ScalingManual / Cluster AutoscalerAutomatic per pod
CostLower for steady workloadsHigher for long-running workloads
MaintenancePatching & node management neededFully managed
SecurityCustomizableIsolated, AWS managed
Use CaseCritical production workloadsStandby DR workloads
Stateful SupportYesLimited (can use EBS, but more complex)

Verdict:

  • EKS on EC2 → better for production because you need performance, full control, and persistent workloads.

  • EKS Fargate → better for DR because you don’t want to manage nodes for a cluster that mostly sits idle.


4️⃣ Potential Follow-up Questions

  1. Why not use Fargate for production?

    • Cost and performance constraints, limited OS control for fine-tuning.
  2. How does autoscaling work in both clusters?

    • EC2: Cluster Autoscaler + HPA; Fargate: Pod scaling automatically managed by AWS.
  3. How do you handle stateful workloads in DR on Fargate?

    • Use EBS or EFS; preferably DR workloads are stateless or replicated from production DB.
  4. How do you replicate data from AUS (prod) to Singapore (DR)?

    • Cross-region RDS/Aurora replication or S3 replication, DynamoDB global tables for caching/state.
  5. How fast can you failover from AUS → Singapore?

    • Depends on DR RTO/RPO strategy; pre-provisioned Fargate cluster reduces failover time.
  6. How do you handle security & compliance differences between regions?

    • AUS → APRA compliance, Singapore → MAS regulations; IAM, KMS, VPC segregation, and network policies are replicated per region.
  7. Cost optimization strategy?

    • Production EC2 nodes can use spot/reserved/auto-scaling. DR Fargate remains idle until failover → pay-per-use.

Q. Is having EKS Fargate best strategy or we using EKS with nodes ?

It depends on the use case and workload characteristics:


1️⃣ EKS on EC2 nodes

  • Best for: Production workloads with high traffic, stateful services, or performance-sensitive microservices.

  • Pros:

    • Full control over OS, instance type, and networking

    • Can optimize cost with reserved/spot instances

    • Easier to integrate monitoring, logging, and sidecars

  • Cons:

    • You must manage nodes (patching, scaling, updates)

    • Slightly more operational overhead


2️⃣ EKS Fargate

  • Best for: DR or low-traffic workloads, stateless apps, or when you want fully managed compute.

  • Pros:

    • No node management

    • Auto-scales per pod

    • Strong isolation and security

    • Ideal for standby DR cluster (pay only when used)

  • Cons:

    • More expensive for always-on workloads

    • Less control over OS-level configs

    • Stateful workloads are trickier


💡 Recommendation for fintech wealth advisory app

  • Production: EKS on EC2 → stable, high-performance, persistent microservices.

  • DR: Fargate → low maintenance, secure, pay-per-use, quickly scalable in case AUS region fails.

So yes, using Fargate for DR is a strategic choice, but for primary production EKS with nodes is still better.


Question: Why do you use EKS on EC2 for production and Fargate for DR in your fintech wealth advisory platform?

Answer:

For our primary production workloads in Australia, we use EKS on EC2 nodes because our platform has multiple critical microservices that are performance-sensitive and often stateful. EC2 nodes give us full control over OS-level configurations, instance types, and persistent storage, which is essential for high availability and optimized cost using reserved or spot instances.

For our disaster recovery cluster in Singapore, we use EKS Fargate. The DR cluster is mostly idle until a failover event occurs, so Fargate allows us to avoid managing EC2 nodes and automatically scales pods as needed. It also provides stronger isolation and reduces operational overhead, which is ideal for a standby DR setup.

In summary, EKS on EC2 is better for consistent, high-performance production workloads, whereas Fargate is optimal for on-demand, low-maintenance DR workloads, balancing cost, security, and scalability.


Possible Follow-up Questions & Answers:

  1. Q: Why not use Fargate for production?
    A: Fargate is more expensive for always-on workloads and offers limited OS-level control, which is critical for fine-tuning performance and running stateful services in production.

  2. Q: How do you replicate data between AUS and Singapore?
    A: We use cross-region replication for RDS/Aurora and S3, and for caching/state, DynamoDB global tables.

  3. Q: How fast can you failover to DR?
    A: With a pre-provisioned Fargate cluster and replicated data, our RTO is minimized, and we can bring services online within minutes.

  4. Q: How do you ensure security and compliance across regions?
    A: We enforce IAM roles, KMS encryption, VPC isolation, and network policies according to APRA in Australia and MAS in Singapore.

  5. Q: How do you optimize costs in this setup?
    A: Production EC2 nodes leverage reserved or spot instances with auto-scaling, while DR Fargate only incurs cost when pods are running.

Q How traffic flows in your setup when everything is normal (production healthy) for your fintech wealth advisory app ?


Traffic Flow (Normal Scenario)

  1. Client request:

    • User accesses the app via browser or mobile.
  2. DNS / Route53:

    • Requests go to your primary region (Australia) via Route53, typically using weighted or failover routing.

    • Since production is healthy, all traffic is routed to AUS.

  3. Load Balancer:

    • ALB (Application Load Balancer) in front of EKS receives requests.

    • ALB handles SSL termination, routing to appropriate Kubernetes services (via ingress).

  4. Kubernetes Ingress:

    • Ingress rules in EKS direct traffic to the correct microservice pods running on EC2 nodes.
  5. Pods / Services:

    • Each microservice pod processes the request, accessing RDS/Aurora, Redis, or S3 as needed.
  6. Response back:

    • The response flows back through the ALB → Route53 → user.

Key Points

  • Traffic always goes to the primary region (EKS on EC2) as long as health checks pass.

  • DR (Fargate in Singapore) does not serve traffic unless AUS fails.

  • Route53 failover uses health checks on ALB endpoints to decide when to route traffic to DR.


Q. So using EKS Fargate is best strategy for DR ?

Answer:

Yes, using EKS Fargate for DR is a strategic choice. Since DR workloads are mostly idle until failover, Fargate eliminates the need to manage EC2 nodes, automatically scales pods on demand, and provides strong isolation and security. This reduces operational overhead and cost while ensuring rapid failover in case the primary production region fails.


Q. When app is running on EKS Fargate and once EKS node is running how you get the data generated on EKS Fargate to EKS node setup ? How is it ? Do both share same data sources or has separate one ?

Here’s how it works in a typical fintech setup with EKS on EC2 (prod) and EKS Fargate (DR):


1️⃣ Data Sources

  • Both clusters do NOT share the same local storage.

    • Fargate pods don’t have persistent local disks like EC2 nodes.

    • Local pod storage is ephemeral.

  • Shared data is always in external services:

    • Databases: RDS/Aurora (with cross-region replication)

    • Object storage: S3 buckets (cross-region replication)

    • Caches: Redis/Memcached (replicated or rehydrated in DR)


2️⃣ How data flows

  1. Production EKS (EC2) writes data to shared external stores (RDS, S3, etc.).

  2. DR EKS (Fargate) is mostly idle but has access to the same replicated data sources.

  3. On failover:

    • Route53 switches traffic to the DR region.

    • Fargate pods start serving requests using replicated data.

  4. No “syncing from Fargate → EC2” is needed in normal operations.


Key Point

  • Both clusters use the same replicated external data sources, not separate ones, to ensure data consistency and minimal RTO/RPO.

  • Fargate is essentially stateless; all critical data lives in managed, replicated services.


Q. That means if i have a app say ie Python, React and RDS for MySQL we have same RDS for both EKS node and EKS Fargate right ?

Exactly ✅.

For your Python + React app with MySQL (RDS):

  • Primary (EKS on EC2, AUS):

    • App reads/writes to RDS MySQL in AUS.
  • DR (EKS Fargate, Singapore):

    • Uses the same database replicated to Singapore (cross-region replication).

    • Fargate pods are stateless; they just connect to the replicated RDS when DR is active.

Key point for interviews:

Both clusters do not have separate MySQL instances for regular operations; DR relies on replicated, external RDS to ensure consistency and minimal failover time.

If you want, I can also give a crisp 1–2 sentence interview answer for this scenario. Do you want me to do that?

You said:

Q. So both workloads points to same RDS instance ? S3 buckets and other storages right ?

Yes, exactly. Here’s a clear interview-ready answer:


Answer:

Both workloads—production on EKS nodes and DR on EKS Fargate—point to the same external data sources, such as RDS for MySQL, S3 buckets, and caches. DR relies on replicated or cross-region versions of these resources, ensuring data consistency and allowing failover without data loss.


This makes it clear that both clusters share the same data layer, but DR is mostly idle until needed.

Q. Do you have both DR and Prod in diff VPC as you said you have different accounts environment wise right ?

Yes. Here’s a concise explanation suitable for interviews:


Answer:

Yes, our production and DR clusters are in separate VPCs and AWS accounts per environment. Production (EKS on EC2) resides in the AUS account/VPC, while DR (EKS Fargate) is in the Singapore account/VPC. This ensures isolation, security, and compliance, while still allowing DR to access replicated external resources like RDS, S3, and caches.


It emphasizes environment isolation and best practices for fintech compliance.

Q. So how do EKS Fargate access data from Singapore as its in diff VPC ,as you said about isolation ?

Here’s how it works in a secure, production-ready setup for your scenario:


Accessing Data Across VPCs / Accounts

Even though DR (EKS Fargate) is in a separate VPC/account, it can still access shared data in a secure way:

  1. Cross-Region / Cross-VPC Replication

    • RDS MySQL: Set up cross-region read replicas. Fargate connects to the Singapore replica directly.

    • S3 Buckets: Enable cross-region replication (CRR). DR pods read/write to the replicated bucket in Singapore.

  2. Networking Options

    • VPC Peering / Transit Gateway: Securely connect VPCs if needed.

    • PrivateLink / VPC endpoints: Access S3 or RDS without exposing traffic to the public internet.

  3. Security

    • IAM roles and KMS encryption ensure DR can access only the intended resources.

    • Network policies in Kubernetes restrict pod-level access.


Key Point (Interview Answer):

DR on EKS Fargate accesses replicated external resources (RDS, S3) either via cross-region replication or secure VPC connectivity (PrivateLink/VPC endpoints), without breaking isolation or security compliance.


Q. Which is best approach to have prod and DR in same VPC or in diff VPC and diff accounts and diff region ?

For a fintech wealth advisory platform, the best approach is to have production and DR in different VPCs, different accounts, and different regions. Here’s why:


1️⃣ Same VPC (or same account)

  • Pros:

    • Simpler networking, no cross-VPC setup needed.

    • Slightly easier access to shared resources.

  • Cons:

    • Risk of single point of failure (VPC misconfiguration, network outage).

    • Less isolation → security and compliance concerns.

    • Region-level disaster (AZ failure or region outage) affects both prod and DR.


  • Pros:

    • True isolation: DR is independent of prod failures.

    • Improves compliance and security (regulatory requirements like APRA, MAS).

    • Enables cross-region disaster recovery (RTO/RPO achievable).

    • Limits blast radius: misconfigurations in prod don’t affect DR.

  • Cons:

    • Requires cross-region replication of data (RDS, S3).

    • Slightly more complex networking setup (VPC endpoints, PrivateLink).


✅ Conclusion / Interview Answer:

The best practice for critical fintech workloads is to place production and DR in separate VPCs, accounts, and regions. This ensures maximum isolation, security, and compliance while allowing DR to access replicated resources for quick failover.


Q. So when failover occurs how does it switches you said you have separate pipeline right so do you have pipeline getting triggered from main/prod branch only and is you EKS Fargate cluster always running ?

Here’s the interview-ready explanation for your scenario:


1️⃣ Pipeline & Deployment

  • Primary Pipeline:

    • Triggered from main/prod branch.

    • Deploys microservices to EKS on EC2 in the primary region (Australia).

  • DR Pipeline:

    • Usually not continuously triggered.

    • Can be triggered manually or automatically during failover to deploy the latest stable images to the EKS Fargate DR cluster in Singapore.

  • Image Management:

    • Docker images are stored in a central registry (ECR) and used by both prod and DR pipelines.

    • Helm charts or manifests are updated via scripts to point to the correct image version.


2️⃣ Is EKS Fargate always running?

  • No, not fully active all the time:

    • Fargate DR pods can be kept idle (scaled to 0) to save cost.

    • On failover, the cluster scales pods up automatically using the same container images as prod.

  • Data:

    • DR accesses the replicated RDS/S3/data sources, so no data is lost.

3️⃣ Key Interview Answer (Crisp Version)

Our production pipeline deploys to EKS on EC2 from the main branch, while the DR pipeline is triggered only during failover or periodic sync. The EKS Fargate DR cluster is mostly idle to reduce cost, and scales up automatically to serve traffic using replicated external resources like RDS and S3.


Q. How your EKS Fargate DR cluster comes up and starts serving traffic ?

Here’s a clear step-by-step explanation


1️⃣ Triggering the DR Pipeline

  • The DR pipeline can be triggered manually during failover or automatically if health checks on production fail.

  • The pipeline ensures it’s using the latest stable Docker images from your container registry (ECR).


2️⃣ Preparing the Cluster

  • EKS Fargate cluster exists but pods may be scaled to 0 (idle) to save cost.

  • Pipeline applies the Helm manifests or Kubernetes YAMLs to the cluster:

    • Updates container images

    • Configures services, ingress, environment variables, secrets (from Secrets Manager)


3️⃣ Scaling Up Fargate Pods

  • Kubernetes automatically schedules pods on Fargate:

    • Fargate provisions the compute required to run each pod (serverless, no nodes to manage)

    • Pods mount secrets and connect to replicated data sources (RDS/S3)


4️⃣ Serving Traffic

  • Route53 detects production failure via health checks.

  • DNS switches traffic from AUS (EKS on EC2) → Singapore (EKS Fargate).

  • The pods are now live, serving the app using replicated external resources.


Key Points for Interviews

  • Pipeline → Registry → Fargate pods → Traffic switch.

  • Fargate is serverless, so the cluster doesn’t need pre-provisioned nodes.

  • Data comes from replicated external sources, not from the prod nodes themselves.

  • Cost-efficient: idle until failover.


Q. Give me step by step for when a failover occurs and your EKS Fargate DR cluster comes up. keep it simple ?


Failover Step-by-Step Flow

1️⃣ Production Failure Detected

  • AWS Route53 health checks detect that the primary production EKS (EC2) cluster in AUS is unhealthy.

  • Health check fails → triggers DNS failover to the DR region (Singapore).


2️⃣ DR Pipeline Triggered

  • The DR deployment pipeline runs.

    • Can be manual or automatic depending on setup.

    • Pulls the latest stable Docker images from your ECR registry.

    • Pulls the Helm charts or Kubernetes manifests configured for Fargate.


3️⃣ EKS Fargate Cluster Provisioning

  • The EKS Fargate cluster already exists in Singapore, but pods are scaled to 0 (idle) to save cost.

  • The pipeline applies manifests, which:

    • Schedule pods for each microservice.

    • Assign pods to Fargate compute (serverless nodes).

    • Mount secrets (from AWS Secrets Manager) and configure environment variables.


4️⃣ Pods Spin Up

  • Kubernetes schedules pods on Fargate.

  • Fargate automatically provisions the compute needed.

  • Pods connect to replicated external resources:

    • RDS MySQL read replicas or cross-region replicas

    • S3 buckets with cross-region replication

    • Caches (Redis/Memcached) if used


5️⃣ Traffic Starts Flowing

  • Route53 has already switched traffic to Singapore.

  • Requests now hit the Fargate DR pods, which serve the app normally.

  • From the user’s perspective, the app is live — minimal downtime.


6️⃣ Post-Failover

  • DR cluster can remain active until primary is restored.

  • Once prod is back, traffic switches back to AUS.

  • DR pods can scale down to 0 to save cost.


Key Points to Emphasize in an Interview

  • Fargate is serverless → no nodes to manage; pods are scheduled on-demand.

  • Data is not copied from prod nodes; it comes from replicated external sources (RDS, S3).

  • Pipeline ensures pods use the latest stable images.

  • Failover is fast because DR cluster exists, just mostly idle.


Q. But how it comes up with those deployments ,pods when ig you said its idle with 0 , and our ci and cd code for k8s manifest and helm is in cd repo right how it happend s? is it like when failover occurs does it route traffic to eks fragate and it triggers pipeline

Ah! Now I understand exactly what’s confusing — you’re asking how Fargate “comes up” from idle 0 pods and how CI/CD interacts with it during failover. Let me clarify step by step, exactly how it works in a typical setup:


1️⃣ Idle DR Cluster

  • The EKS Fargate DR cluster exists, but all microservice pods are scaled to 0 to save cost.

  • Cluster resources exist: VPC, subnets, security groups, EKS control plane.

  • No workload is running until the pipeline or autoscaling triggers pods.


2️⃣ CI/CD for DR

  • Your CI/CD repository contains:

    • Kubernetes manifests or Helm charts

    • Docker image references (from ECR)

    • Environment variables, secrets config

  • Pipeline is configured to deploy to the DR Fargate cluster whenever it runs.


3️⃣ Failover Occurs

  • Route53 health checks detect prod (EKS on EC2) failure.

  • Traffic is switched to Singapore DR cluster.


4️⃣ Triggering DR Deployment

There are two common strategies:

Option A: Manual/Automatic Pipeline Trigger

  • When failover occurs, the DR pipeline is triggered (manual or automated via monitoring).

  • Pipeline applies Helm charts / manifests to the Fargate cluster.

  • Kubernetes schedules pods on Fargate.

Option B: Pre-deployed “always-on” pods (optional)

  • Sometimes critical DR workloads are kept scaled to 1 or minimal replicas.

  • Pipeline already applied manifests; pods are running in DR but very low resource.

  • Failover simply routes traffic, no need to scale up.


5️⃣ How Pods Come Up on Fargate

  • Helm or manifests define deployments/services.

  • Kubernetes schedules the pods → Fargate provisions compute automatically.

  • Pods pull Docker images from ECR.

  • Pods mount secrets from Secrets Manager and connect to replicated RDS/S3.


6️⃣ Traffic Now Served

  • Route53 points traffic to the DR cluster.

  • Pods are live, app works normally.

  • Cost is minimized because Fargate scales pods only as needed.


Key Clarification

Failover does not automatically trigger pipeline just because traffic is routed.
The pipeline either:

  • Runs periodically to keep DR updated, or

  • Is triggered during failover (via monitoring/automation) to spin up the pods.


If you want, I can draw a clear diagram showing “pipeline → Helm/manifest → Fargate pods scale up → Route53 traffic” so you can visualize it for interviews.

Do you want me to do that?

You said:

but how it picks the pods and other details during failover ?

ChatGPT said:

Got it — you want to understand how Fargate “knows” which pods and deployments to run during failover, even when the DR cluster is idle. Here’s the clear explanation:


1️⃣ DR Cluster Has Manifests Predefined

  • Even if the pods are scaled to 0, the Kubernetes deployments, services, and Helm charts exist in the cluster.

  • These define:

    • Pod templates (images, env variables, secrets)

    • Number of replicas (can be 0 initially)

    • Services and ingress rules

    • ConfigMaps and volume mounts

Think of this as a blueprint — the cluster “knows” what to run, it just hasn’t allocated compute yet.


2️⃣ Pipeline / Helm Applies During Failover

  • The CD pipeline applies the manifests or Helm charts to the DR cluster.

  • Kubernetes reads the deployment specs:

    • Image versions (from ECR)

    • Number of replicas

    • Resource requirements, environment variables, secrets

  • Kubernetes then schedules pods according to those specs.


3️⃣ Fargate Automatically Allocates Compute

  • Fargate is serverless, so it automatically provisions compute for each pod.

  • Pods are instantiated using the images and config defined in the manifests.

  • Secrets and configmaps are mounted, and pods connect to replicated RDS/S3 resources.


4️⃣ Route53 Switches Traffic

  • Once pods are running, Route53 routes traffic to the DR cluster.

  • The pods are now live and serve requests exactly like production.


Key Point (Interview Version)

Even if the DR cluster is idle, the Kubernetes manifests/Helm charts already define all pods, services, and configurations. During failover, the CD pipeline applies these manifests, Fargate provisions compute, and pods spin up automatically with the correct images and environment.


If you want, I can make a step-by-step failover diagram showing DR idle → pipeline → pods spin up → traffic switch, which is very visual for interviews.

Do you want me to do that?

You said:

but how does it comes up automatically ?

ChatGPT said:

Ah! Now I understand — you’re asking how Fargate automatically spins up pods even when the DR cluster is idle. Let me break it down very clearly:


1️⃣ Kubernetes Keeps Deployment Specs in Cluster

  • Even if pods are scaled to 0, the Deployment/StatefulSet objects exist in Kubernetes.

  • These objects contain:

    • Container image to run

    • Number of replicas (can start from 0)

    • Environment variables, secrets, configmaps

    • Resource requests/limits

Think of it as a blueprint stored in the cluster — the cluster “knows” what to create when instructed.


2️⃣ How Pods Spin Up

There are two ways pods come up automatically:

A) Using a CD pipeline trigger

  • The failover monitoring system detects prod failure.

  • The CD pipeline applies the manifests or Helm charts to the DR cluster.

  • Kubernetes sees that replicas are 0 and updates them to the desired number (usually 1+).

  • Fargate provisions compute automatically and starts pods.

B) Using Kubernetes Horizontal/Cluster Autoscaler (optional)

  • You can configure HPA or KEDA to scale DR pods automatically based on metrics or a “failover” trigger.

  • Once triggered, the cluster spins up pods as defined in the Deployment spec.


3️⃣ Fargate Role

  • Fargate is serverless — it provisions compute on demand whenever a pod needs to run.

  • No nodes are required; the pod just appears with the defined image, secrets, and environment.


4️⃣ Key Takeaways

  • Pods don’t magically appear — they come up because:

    1. Kubernetes has the Deployment blueprint stored.

    2. Either the CD pipeline or autoscaler tells Kubernetes to create pods.

    3. Fargate provisions the compute automatically.

  • DR cluster can remain mostly idle, but failover triggers pod creation automatically.


If you want, I can make a simple diagram showing DR idle → failover trigger → CD pipeline → Fargate spins up pods → traffic starts, which makes it crystal clear for interviews.

Do you want me to do that?

You said:

but how does cd pipeline gets triggered automatically when failover occurs

ChatGPT said:

Ah, now we’re at the core of “automatic failover triggering the CD pipeline”. Let me explain step by step in a fintech DR context:


1️⃣ Health Checks Detect Failure

  • Route53 (or a monitoring system like CloudWatch/Prometheus) continuously monitors your primary production cluster (EKS on EC2 in AUS).

  • Health checks include:

    • ALB / service endpoints

    • Application response (HTTP 200 / custom checks)

  • If the health check fails, it triggers an alert or a failover event.


2️⃣ Failover Event Triggers Pipeline

There are a few ways this happens:

A) Event-driven CD pipeline (recommended)

  • Monitoring alert (CloudWatch Alarm / EventBridge / custom Lambda) detects prod failure.

  • Event triggers the CD pipeline for the DR cluster in Singapore.

    • Pipeline fetches manifests/Helm charts from CD repo.

    • Applies them to the DR cluster.

B) Pre-deployed DR with minimal pods (simpler option)

  • DR cluster already has pods running at low scale.

  • Route53 switches traffic automatically.

  • No pipeline trigger is needed; traffic just goes to the already running pods.


3️⃣ Pipeline Actions

When triggered, the pipeline:

  1. Pulls latest container images from ECR.

  2. Applies Helm charts or Kubernetes manifests to the DR cluster.

  3. Kubernetes schedules pods on Fargate (serverless compute).

  4. Pods connect to replicated RDS/S3 resources.


4️⃣ Key Points for Interviews

  • The CD pipeline does not automatically know failover happened; it is triggered by monitoring events (Route53 health checks, CloudWatch alarms, EventBridge).

  • This ensures DR pods are always up-to-date with latest images/configs.

  • Optionally, DR cluster can be pre-warmed with minimal pods to avoid pipeline trigger latency.


If you want, I can make a diagram showing “Route53 health check → CloudWatch/EventBridge → CD pipeline → DR Fargate pods → traffic switch”, which is very visual for interviews.

Do you want me to do that?

You said:

Q. or can we say we have eks fagate running with minimal load like we split traffic vai servcie mesh using split tarffic 5 % to secondary and rest to primary during primary it move 100% to eks fragte the dr setup

Yes — that’s actually a very common blue/green or canary-style DR strategy and works well for interviews. Here’s how you can frame it:


Interview-Ready Explanation

Strategy:

  • The EKS Fargate DR cluster is kept running at minimal load — just enough pods to be warm.

  • Using service mesh (Istio/Linkerd) or ALB weighted routing, traffic can be split between primary and DR:

    • Example: 95% to primary (EKS on EC2, AUS)

    • 5% to secondary (EKS Fargate, Singapore)

  • This keeps DR pods ready and tested in real traffic, reducing failover latency.

Failover Behavior:

  • If primary fails or becomes unhealthy:

    • Route53/service mesh shifts 100% traffic to DR Fargate automatically.

    • No cold-start delay because DR pods were already running.

Benefits:

  • DR is warm and validated continuously.

  • Reduces failover downtime.

  • Ensures DR pods are using latest images and configuration.

  • Can test DR under real traffic without affecting production significantly.


How to Phrase in an Interview

“Our DR cluster on EKS Fargate is kept running with minimal pods. Using a service mesh, we split traffic (e.g., 5% to DR) so DR is warm and validated. During failover, 100% of traffic is routed to DR instantly, ensuring near-zero downtime.”


Q . How much does it cost and for EKS node and for EKS Fargate as you said you use reserved instances with long term and EKS Fargate for same ?

Here’s a clear breakdown of costs for your setup — EKS on EC2 nodes vs EKS Fargate, with consideration for reserved instances and DR use case:


1️⃣ EKS on EC2 (Production)

Components of cost:

  1. EC2 instances:

    • You pay for compute per hour.

    • Using Reserved Instances (1–3 year term) reduces cost by ~30–60% vs on-demand.

    • Example: t3.medium ~ $0.0416/hour on-demand, Reserved Instance ~ $0.025/hour.

  2. EKS Control Plane:

    • AWS charges ~$0.10 per hour (~$72/month) per cluster.
  3. Other resources:

    • ALB, NAT Gateways, S3, RDS, CloudWatch, etc.

Cost optimization:

  • Use Reserved Instances / Savings Plans for predictable production workloads.

  • Cluster autoscaler can scale nodes based on load.


2️⃣ EKS Fargate (DR)

Components of cost:

  1. Fargate pods:

    • Billed per vCPU and GB of memory per second.

    • No EC2 nodes to manage.

  2. EKS Control Plane:

    • Same $0.10/hour per cluster.

Cost example:

  • Suppose DR has 15 microservices, 2 vCPU + 4 GB memory per pod, 1 replica per microservice:

    • 15 pods × (2 vCPU + 4 GB) × ~per-second pricing → ~$0.20–0.25/hour (approx).
  • Idle DR cluster scaled to 0 → you pay almost nothing except EKS control plane.

Cost optimization:

  • Scale pods to 0 until failover → only pay for EKS control plane.

  • Keep a few pods running (minimal load, e.g., 5% traffic) for warm DR.


3️⃣ Key Comparison

FeatureEKS on EC2 (Production)EKS Fargate (DR)
Compute costPaid per EC2 instance; reserved reduces costPaid per vCPU+memory; serverless
ControlFull control over nodesServerless, no node control
ScalingManual/AutoscalerAutomatic per pod
Idle costAlways pay for nodesCan scale to 0 → minimal cost
Use caseAlways-on productionDR / failover / warm standby

💡 Interview-Ready Answer

“For production EKS on EC2, we use reserved instances to reduce compute costs, while paying for the EKS control plane and associated AWS resources. For DR on EKS Fargate, we keep pods idle or at minimal load, paying mainly for the EKS control plane, and only scale pods during failover. This makes Fargate cost-efficient for standby DR while production remains optimized with reserved EC2 nodes.”


Q. So for EKS node we use combination of both m5.xlarge and r5 large routing traffic respectively to pods depending upon workloads son considering its a fintech wealth advisory platform give setup of master node and worker node EKS Fargate RAM/CPU during idel as well when running 100% workload during failover ?

Got it! Let’s create a detailed, interview-ready setup for your fintech wealth advisory platform with both EKS on EC2 (prod) and EKS Fargate (DR), including master/worker nodes, instance types, and pod specs for idle vs failover:


1️⃣ EKS on EC2 (Production — AUS)

Master / Control Plane

  • Managed by AWS → no EC2 nodes to manage.

  • Always-on, highly available across 3 AZs.

  • Handles scheduling, API server, etc.

  • Cost: ~$0.10/hour per cluster.

Worker Nodes (EC2)

  • Combination of M5.xlarge (general compute) and R5.large (memory-optimized):

    • M5.xlarge: 4 vCPU, 16 GB RAM → handles CPU-heavy workloads like backend processing, transactions.

    • R5.large: 2 vCPU, 16 GB RAM → handles memory-heavy workloads like caching, analytics.

  • Pod scheduling: CPU/memory requests defined in Kubernetes → pods routed to appropriate node type.

  • Scaling: Cluster Autoscaler adjusts node counts based on pod demand.


Pod Example (Production)

MicroservicePod CPUPod MemoryNode Type
Payments1 vCPU4 GBM5.xlarge
Reports1 vCPU8 GBR5.large
Notifications0.5 vCPU2 GBM5.xlarge

2️⃣ EKS Fargate (DR — Singapore)

Cluster & Pods

  • Cluster exists always but pods mostly idle/minimal replicas.

  • CPU / RAM per pod (idle):

    • 0.5–1 vCPU, 1–2 GB RAM (minimal load, e.g., 5% traffic).
  • CPU / RAM per pod (100% failover):

    • Match production pod specs: 1–2 vCPU, 4–8 GB RAM depending on service.
  • Pod scaling:

    • HPA or pipeline triggers scale replicas up to handle full traffic.
  • Fargate provisions compute automatically → serverless, no EC2 nodes.


Pod Example (DR — Idle vs Failover)

MicroserviceCPU (Idle)RAM (Idle)CPU (Failover)RAM (Failover)Notes
Payments0.5 vCPU2 GB1 vCPU4 GBMinimal replica idle → full traffic during failover
Reports0.5 vCPU2 GB1 vCPU8 GBMemory-heavy, scale during failover
Notifications0.25 vCPU1 GB0.5 vCPU2 GBIdle → full failover

3️⃣ Traffic Routing & Failover

  • Normal operation:

    • Route53 directs 95–100% traffic to EKS EC2 nodes, 5% to DR Fargate (warm pods).
  • Failover:

    • Route53 switches 100% traffic to DR Fargate.

    • Pods scale up to full CPU/RAM per production pod spec.

    • Fargate automatically provisions resources; DR pods serve using replicated RDS/S3.


4️⃣ Cost Optimization

  • Production (EKS on EC2): Reserved instances for M5.xlarge/R5.large.

  • DR (Fargate): Idle minimal pods cost very little; only scale during failover.

  • Control plane: $0.10/hr per cluster, applies to both prod and DR.


5️⃣ Interview-Ready Summary

“In production, we use EKS on EC2 with a mix of M5.xlarge for CPU-heavy and R5.large for memory-heavy microservices. Worker nodes autoscale based on pod requests. DR on EKS Fargate has minimal idle pods for validation and scales to match production pod specs during failover, handling 100% of traffic using replicated RDS and S3, with compute automatically provisioned serverlessly.”


Summary — DR setup (fintech wealth-advisory) — short & interview-ready

TL;DR:
We run prod in AUS on EKS (EC2) and a warm DR in SG on EKS Fargate (separate account/VPC/region). Data is replicated (RDS cross-region replicas, S3 CRR, etc.), images live in ECR, Helm/manifests live in CD repo. DR keeps minimal replicas (warm), we split a small % of traffic for validation, and on failover the CD/automation scales DR pods to full capacity and Route53/service-mesh routes 100% traffic to Fargate.


Step-by-step DR flow (pre-failover → failover → failback)

Pre-failover / steady state

  1. Infrastructure: Prod = EKS on EC2 in AUS (M5 / R5 worker mix). DR = EKS Fargate in SG, separate AWS account & VPC.

  2. Shared data: RDS MySQL with cross-region replicas, S3 with CRR, DynamoDB/global tables or other replicated stores.

  3. CI/CD & artifacts: Container images in ECR; Helm charts/manifests in CD repo; secrets in Secrets Manager / Parameter Store; KMS encryption.

  4. DR posture: DR cluster exists and is warm — deployments present. Replicas are minimal (e.g., 1 per service) or selectively pre-warmed. Service mesh or ALB weighted routing sends a small percentage (e.g., 5%) to DR for validation.

  5. Monitoring: Route53/ALB health checks + CloudWatch/Prometheus push alerts to EventBridge.

Failover initiation
6. Detect failure: Route53/ALB health checks or CloudWatch alarms detect primary failure (API or ALB unhealthy).
7. Trigger automation: EventBridge/Lambda or a runbook tool triggers the DR pipeline (or the pipeline may be preconfigured to run automatically on the event). Optionally, autoscalers (HPA/KEDA) can be signaled.

Bringing DR to full capacity
8. CD applies manifests: Pipeline pulls Helm/manifests from CD repo and applies them to the DR cluster (or patches replica counts). Manifests reference stable images in ECR.
9. Fargate provisions pods: Kubernetes schedules pods; Fargate automatically provisions the compute and pods pull images, mount secrets, and initialize.
10. Connect to replicated data: Pods use SG RDS read/replica endpoints or promoted replica as primary, access CRR S3 or local endpoints, and use PrivateLink/VPC endpoints for secure access.
11. Traffic switch: Route53 or service mesh flips routing to 100% DR (DNS failover or weight change). Health checks validate DR endpoints.
12. Operate on DR: Monitoring, logging, and security controls remain active. Runbooks and incident response follow.

Failback
13. Recover prod: When AUS is healthy, ensure data sync (if writes happened in DR), reverse replication or promote original RDS, validate consistency.
14. Switch traffic back: Route53/service mesh routes traffic to prod, scale down DR to warm/minimal replicas, and run post-mortem & DR rehearsal updates.


Key technical pieces & how they connect (one line each)

  • Route53 health checks → detect primary failure and initiate DNS failover.

  • EventBridge / Lambda / PagerDuty → triggers CD pipeline when health fails.

  • CD pipeline (Helm/Kustomize) → applies manifests to DR cluster or scales replicas.

  • ECR → single source for images used by both prod & DR.

  • EKS Fargate → serverless compute that provisions pods on demand.

  • RDS / S3 replication → guarantees data availability and sets RPO.

  • Service mesh / ALB weighted routing → supports warm traffic split (e.g., 95/5) and instant weight shift on failover.

  • Secrets Manager + KMS → secure cross-account secret access.

  • HPA/KEDA or manual scaling → scales DR pods during failover.


Operational considerations (brief)

  • RTO: Minutes (warm pods + automated pipeline); depends on warm strategy.

  • RPO: Depends on replication (async CRR ≈ small; choose aurora/global tables for lower RPO).

  • Testing: Regular DR runbooks & rehearsals, smoke tests, and canary traffic to DR.

  • Security & compliance: Separate accounts/VPCs, IAM least privilege, KMS, audit logging.

  • Cost: DR on Fargate is cost-efficient if mostly idle or minimally warmed.


Interview-ready concise answer you can deliver (one paragraph)

“Our DR is a warm-standby EKS Fargate cluster in a separate AWS account/VPC in Singapore. We replicate primary data (RDS cross-region replicas, S3 CRR), keep minimal replicas in DR for validation (5% traffic split via service mesh), and store images and manifests centrally (ECR + CD repo). On detection of primary failure, monitoring triggers EventBridge → CD pipeline which applies Helm charts to DR, Fargate provisions pods, and Route53/service-mesh shifts 100% traffic to DR. This gives fast failover with minimal operational overhead while meeting fintech RTO/RPO requirements.”


Q. Which deployment strategy you use and why ?

Good question 👌 — here’s the interview-ready way to explain deployment strategy for your fintech wealth advisory platform (Prod = EKS EC2, DR = EKS Fargate):


Deployment Strategy Used

We use a Blue-Green + Weighted Traffic (service mesh / ALB routing) strategy.


Why Blue-Green?

  • Zero downtime: We deploy new versions to a “green” environment while “blue” serves production traffic.

  • Instant rollback: If issues occur, traffic can be switched back quickly.

  • Regulatory compliance: In fintech, outages are unacceptable, so rollback safety is critical.


Why Weighted Traffic / Canary (5–10% to DR)?

  • Continuous DR validation: By sending a small % of live traffic to DR (Fargate), we ensure DR is healthy and up to date.

  • Real traffic testing: Ensures DR is not just theoretical — it’s continuously tested under production-like conditions.

  • Smooth failover: During failover, traffic shifts from 95/5 → 0/100 seamlessly.


Why not Rolling or Recreate?

  • Rolling updates: Good for stateless apps, but rollback takes time if multiple microservices are involved.

  • Recreate strategy: Causes downtime, not acceptable for fintech workloads.


Interview-Ready Answer

“For our fintech wealth advisory platform, we use a Blue-Green strategy combined with weighted traffic routing. Production runs on EKS EC2 nodes, while a warm DR runs on EKS Fargate with minimal replicas. During deployments, we route a small percentage of traffic (5%) to DR to validate, and if healthy, we gradually move 100%. This approach gives us zero downtime, instant rollback capability, and continuous DR validation, which is critical in a regulated fintech environment.”


Q. Compare Blue-Green vs Canary vs Rolling in a small table so you can mention pros/cons if the interviewer pushes further?

Here’s a short comparison you can use in interviews 👇

StrategyProsConsFit for Fintech
RollingGradual rollout, no extra infraSlow rollback, partial downtime risk❌ Not ideal
RecreateSimple, full refreshDowntime, risky❌ Not acceptable
CanaryTest with small % of usersMore complex routing✅ Good for new features
Blue-GreenZero downtime, instant rollbackNeeds extra infra✅ Best for regulated workloads

✅ Answer (1-liner):

“We use Blue-Green with weighted traffic — zero downtime, instant rollback, and continuous DR validation, which is safest for fintech.”


Here’s a tight, interview-ready answer 👇

“For our fintech platform we use a Blue-Green strategy with weighted traffic routing. Production runs on EKS EC2 nodes while the DR EKS Fargate cluster stays warm with minimal pods. We send ~5% of traffic to DR to continuously validate it. During deployments or failover, we can shift 100% traffic instantly. This gives us zero downtime, quick rollback, and continuous DR readiness, which is critical in a regulated fintech environment.”


Q. So how is the traffic flow to pod now when you are using service mesh ?

Here’s the interview-ready way to explain traffic flow with a service mesh (Istio/Linkerd) in your setup 👇


Traffic Flow with Service Mesh

  1. Ingress Gateway:

    • All external requests first hit the service mesh ingress gateway (instead of directly hitting pods via ALB/Ingress).
  2. Sidecar Proxy (Envoy):

    • Every pod has a sidecar proxy injected by the service mesh.

    • Traffic flows Gateway → Sidecar Proxy → Pod.

    • Outbound traffic from pod also flows through the sidecar (for observability, security, retries).

  3. Traffic Splitting:

    • Service mesh enforces routing rules (e.g., 95% to Prod pods on EC2, 5% to DR pods on Fargate).

    • This split happens at the ingress gateway or via VirtualService rules.

  4. Failover:

    • If primary pods are unhealthy, the mesh automatically reroutes 100% traffic to DR pods.

    • Sidecars handle retries, circuit-breaking, and mTLS between services.


Interview-Ready Answer

“With service mesh, all traffic first enters through the ingress gateway, then flows to sidecar proxies attached to each pod. This allows us to split traffic — for example 95% to EKS EC2 pods and 5% to EKS Fargate pods. During failover, the mesh simply shifts 100% traffic to DR automatically. The sidecars also give us observability, retries, and mTLS, which are critical for fintech workloads.”


short interview-ready flow from user → pod with service mesh:


Step-by-Step Traffic Flow

  1. User hits website → DNS routes to CloudFront → ALB/NLB.

  2. ALB/NLB forwards to Service Mesh Ingress Gateway.

  3. Ingress Gateway applies routing rules (95% Prod EC2 pods, 5% DR Fargate pods).

  4. Request goes to sidecar proxy (Envoy) of the target pod.

  5. Sidecar proxy → Application container (pod).

  6. Response flows back the same way (pod → sidecar → ingress → ALB → CloudFront → user).


✅ Interview-ready one-liner:

“When a user hits our fintech platform, traffic flows through CloudFront → ALB → Service Mesh Ingress Gateway → sidecar proxy → pod. The mesh handles routing, observability, and failover, so traffic shifts seamlessly from EC2 pods to Fargate pods when needed.”


Q. So don’t we have here ingress ? to service to deployment then to pod include all that flow ?

Exactly 👍 let me stitch the full short flow with all K8s objects so you can say it in interviews:


Traffic Flow (with Service Mesh + K8s objects):

  1. User → DNS → CloudFront → ALB/NLB.

  2. ALB → Ingress Gateway (service mesh entrypoint).

  3. Ingress → K8s Service (ClusterIP/VirtualService in mesh).

  4. Service → Deployment (load balances across replicas).

  5. Deployment → Pod (sidecar proxy → app container).

  6. Response flows back the same chain → User.


✅ Interview one-liner:

“In our setup, user traffic goes CloudFront → ALB → Mesh Ingress → K8s Service → Deployment → Pod (sidecar → app). The mesh controls routing, observability, and failover, so during DR the same path just shifts traffic to the Fargate pods.”


Do you want me to add DR failover into the same flow (95% Prod → 100% Fargate) so you can answer both normal + DR cases in one shot?

You said:

explain the detailed flow now point to point

ChatGPT said:

Perfect 👍 let’s break it step by step in detail so you can directly explain in an interview. I’ll cover both normal (Prod EKS nodes) and DR (EKS Fargate) traffic flow.


Detailed Traffic Flow (Normal Production – EKS Nodes)

  1. User Access

  2. CloudFront → Load Balancer

    • CloudFront forwards requests to an AWS ALB (Application Load Balancer) in the VPC.

    • ALB terminates SSL and routes traffic to EKS Ingress Controller (NGINX/ALB Ingress).

  3. Ingress Controller → Service Mesh Gateway

    • The Ingress Gateway (Istio/Linkerd) acts as the entry point into the mesh.

    • Applies mTLS, auth policies, routing rules (95% prod / 5% DR if split enabled).

  4. Ingress Gateway → Kubernetes Service

    • For example, the React frontend Service forwards traffic to its Deployment.

    • K8s Service is a ClusterIP/VirtualService object in the mesh.

  5. Service → Deployment → Pod

    • The Service load balances traffic across frontend pods managed by Deployment.

    • Each Pod has 2 containers:

      • App container (React/Node, Python, Java, etc.)

      • Sidecar proxy (Envoy for Istio) → handles routing, metrics, retries.

  6. Pod → Data Layer

    • Pods access shared data sources (RDS MySQL, S3, Secrets Manager).

    • All apps use same RDS cluster with Multi-AZ enabled.

  7. Response

    • Data retrieved → Pod sidecar → Service mesh → Ingress Gateway → ALB → CloudFront → User.

Traffic Flow During Failover (DR – EKS Fargate, Singapore)

  1. Primary outage detected (via Route53 health checks or mesh monitoring).

  2. Route53 DNS failover sends traffic to ALB in Singapore region.

  3. Traffic enters EKS Fargate Ingress Gateway.

  4. Services already deployed with minimal replicas (idle) auto-scale via HPA when load increases.

  5. Fargate pods (frontend, backend, Python APIs, etc.) connect to same RDS (multi-region read replica / Aurora Global DB) and shared storages (S3, Secrets Manager).

  6. Response flows back same path: Pod → Service → Mesh → Ingress → ALB → CloudFront → User.


✅ Interview-ready one-liner summary:

“Traffic enters via Route53 → CloudFront → ALB → Ingress Gateway → K8s Service → Deployment → Pod (sidecar → app → DB). In DR, Route53 shifts traffic to Singapore Fargate cluster, where minimal pods autoscale and connect to the same shared data sources.”

More from this blog

P

Projects

8 posts