Reliability · October 2026

Surviving peak traffic on EKS: architecture, KEDA, Karpenter and DR drills

Diwali sales in India and Black Friday in Germany and the rest of Europe all land in November. This is the playbook I'd use to get a Kubernetes platform through its biggest day, built on what I learned running the FIFA World Cup 2026 platform and on how the industry handles these events.

· updated 9 Oct · 16 min read · Arjun Kanojia

Earlier this year I was on the team at Deutsche Telekom Digital Labs, spread across India and Germany, that ran the infrastructure behind the FIFA World Cup 2026 digital services. We peaked above 10,000 requests per second and had zero downtime through the tournament. The work behind that was not clever. It was a lot of boring preparation done early: multi-region DR with automated backup and restore, distributed k6 load tests, EKS and non-EKS tiers scaled up before every match, autoscaling on top, and an on-call rota that knew exactly what to do.

The stack was a typical mix you see at telecom and media scale: Akamai at the edge, stateless services on EKS, and a stateful tier we ran ourselves on EC2, with MongoDB, Kafka and Aerospike, plus Amazon RDS. That mix matters, because the managed and containerised parts scale in seconds while self-managed databases and brokers on EC2 do not. Most of the real preparation went into the second group.

Teams in India are about to face the same shape of problem with the festive sales, and teams in Germany and across Europe with Black Friday and Cyber Monday. So here is the full playbook, with the configs, not just the principles.

1. What event traffic actually looks like

Normal daily traffic rises and falls over hours. Event traffic doesn't. Before a kickoff or a sale opens, it sits near baseline, then jumps several times over in a few minutes as everyone opens the app together. Three things make that spike nastier than the raw number suggests:

  • Cold caches. The first wave hits endpoints that haven't been called in hours, so cache hit rates drop exactly when load peaks and the database takes the full force.
  • Retry storms. When latency creeps up, clients and SDKs retry. A 5% error rate can quietly double the real request volume if retries have no backoff and jitter.
  • Login and token bursts. Auth, session and token-refresh services are usually sized for steady state, and they become the first bottleneck because every client hits them at once.

The lesson: you can't react your way through a step change. Capacity has to be in place before the spike, and the reactive parts of the system are there to absorb what you got wrong.

2. The reference architecture

This is the reference layout I'd draw for this kind of stack: a CDN edge absorbing as much as possible, two AWS regions, stateless services on EKS, and a stateful tier on EC2 and RDS whose data is copied to the second region so it can take over.

Users / mobile apps Akamai edge: CDN + WAF + bot and rate controlscache static + short-TTL API · absorb DDoS · shed abusive clients DNS failover + health checks AWS REGION A · ACTIVE AWS REGION B · WARM STANDBY ALB (multi-AZ) ALB (multi-AZ) EKS · stateless servicesHPA / KEDA · Karpenter (on-demand + Spot) EKS · scaled downsame manifests via GitOps EC2 + RDS DATA TIER (pre-scaled, not autoscaled) MongoDB replica set Kafka brokers Aerospike cluster Amazon RDS (primary) REPLICAS + RESTORED BACKUPS MongoDB secondaries Kafka (mirrored) Aerospike (XDR) RDS read replica cross-region: replica members · MirrorMaker 2 · Aerospike XDR · RDS replica · snapshots and backups copied
Fig 1. Reference active / warm-standby layout for an Akamai + EKS + EC2 + RDS stack. Stateless services fail over in minutes; the data tier decides your real RTO and RPO.

A few decisions in that picture matter more than they look:

  • The edge is your first autoscaler. Every request Akamai answers from cache never reaches your origin. For event traffic, put short TTLs (even 5 to 30 seconds) on read-heavy API responses like scores, prices and catalogue pages. A 5-second cache on a hot endpoint can turn tens of thousands of requests into one per edge location.
  • Edge rate controls protect you from yourself. Rate and bot rules at the CDN stop one buggy client version or a scraper from eating the headroom you planned for real users.
  • Same manifests in both regions. With GitOps (Argo CD in our case), the standby cluster runs the exact same versions. The only difference between regions should be replica counts and which database is the writer.
  • Self-managed data stores don't autoscale. You can't add a MongoDB shard, rebalance Kafka partitions or grow an Aerospike cluster in the minutes before kickoff. Those tiers are sized from the load test and scaled up days in advance, with headroom for the worst case.
  • Data is the hard part. Stateless services fail over in minutes. Databases and brokers decide your real RTO and RPO, which is why the DR section below is mostly about data.

3. Load testing that finds the knee

We used distributed k6 runs before match days. The goal of a load test is not a green report. It's finding the knee: the point where latency stops rising gently and starts climbing steeply, because that's your real capacity, and it's always lower than the sum of your pod limits.

The executor matters. Most first attempts use a fixed number of virtual users, which accidentally slows down when the system slows down (closed model). Real users don't wait politely, so use an arrival-rate executor (open model) that keeps sending requests whatever the latency. The k6 docs explain the difference well.

// kickoff.js: model a step change, not a gentle ramp
import http from 'k6/http';
import { check } from 'k6';

export const options = {
  scenarios: {
    kickoff: {
      executor: 'ramping-arrival-rate',
      startRate: 800, timeUnit: '1s',
      preAllocatedVUs: 2000, maxVUs: 12000,
      stages: [
        { target: 800,   duration: '10m' },  // baseline
        { target: 12000, duration: '3m'  },  // the spike, steeper than real life
        { target: 12000, duration: '20m' },  // hold: leaks, pool exhaustion, GC
        { target: 3000,  duration: '5m'  },
      ],
    },
  },
  thresholds: {
    http_req_failed:   ['rate<0.005'],
    http_req_duration: ['p(95)<400', 'p(99)<900'],
  },
};

export default function () {
  // weight endpoints by real traffic mix from your access logs
  const r = Math.random();
  const path = r < 0.6 ? '/api/live' : r < 0.85 ? '/api/feed' : '/api/auth/refresh';
  check(http.get(`${__ENV.BASE}${path}`), { '2xx': (res) => res.status < 300 });
}

What I'd insist on for any serious event test:

  • Test above the forecast. Industry practice is to test at 1.5x to 2x the expected peak, because forecasts for big events are usually wrong in the same direction.
  • Run it from outside, through the real edge. Testing the ALB directly skips Akamai, its caching and rate rules, and DNS, which are part of the path users take. Agree the test with your CDN provider first so it isn't blocked as an attack.
  • Hold the peak. A 2-minute spike tells you nothing about connection pool exhaustion, memory growth or database replication lag. Hold for 20 minutes or more.
  • Watch the dependencies, not just the pods. The knee is usually in the data tier: MongoDB connection counts and replication lag, Kafka consumer lag and broker disk throughput, Aerospike latency histograms, RDS connections. Or it's an auth service or a third-party API quota.
  • Test, fix, test again. Leave enough calendar time for at least two rounds. A test the week before launch can only give you bad news.

4. Scaling in layers: HPA, KEDA, Karpenter, headroom

Autoscaling on Kubernetes is really four separate loops, and each has a different reaction time. Understanding that is the whole game.

requests / capacity T−60T−30kickoffT+30T+60 KEDA cron pre-scale (T−45) + headroom pods hold spare nodesheadroom Karpenter + KEDA react to real load trafficcapacity
Fig 2. Capacity has to lead traffic. Scheduled pre-scaling covers the step; reactive scaling covers the error in your forecast.
LayerWhat it scalesTypical reaction timeRole on event day
HPAPods, on CPU or memory15 s to a few minutesBaseline elasticity. Too slow and too indirect for a step change on its own.
KEDAPods, on events, queues, Prometheus queries or a scheduleSeconds after the metric movesScale on the real signal (RPS, queue lag) and pre-scale on a cron schedule.
KarpenterNodesRoughly 45 to 90 s for a new node to be ReadyAdds capacity when pods are pending. Fast, but image pulls and readiness add more.
Headroom podsSpare node capacityInstantLow-priority placeholder pods that real pods evict, so nodes already exist when the spike lands.

KEDA: scale on the schedule and on the real signal

I've used KEDA in production for event-driven scaling, and its most underrated feature for events is combining a cron scaler with a metric scaler. KEDA takes the highest value across triggers, so the cron trigger sets a floor before kickoff and the Prometheus trigger takes over if real traffic goes beyond it.

apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: live-api
spec:
  scaleTargetRef: { name: live-api }
  minReplicaCount: 6
  maxReplicaCount: 300
  advanced:
    horizontalPodAutoscalerConfig:
      behavior:
        scaleUp:   { stabilizationWindowSeconds: 0,
                     policies: [{ type: Percent, value: 100, periodSeconds: 15 }] }
        scaleDown: { stabilizationWindowSeconds: 900 }   # don't drop capacity at half-time
  triggers:
    - type: cron                          # floor from 45 min before kickoff
      metadata:
        timezone: Asia/Kolkata
        start: "15 19 * * *"
        end:   "30 22 * * *"
        desiredReplicas: "120"
    - type: prometheus                    # real demand: RPS per pod
      metadata:
        serverAddress: http://prometheus.monitoring:9090
        query: sum(rate(http_requests_total{service="live-api"}[1m]))
        threshold: "80"                  # target RPS per pod, from the load test knee

The threshold isn't a guess. Take the per-pod RPS where your load test showed p99 latency starting to bend, then set the target at about 70 to 80% of it. That margin is what absorbs the minute before new pods are ready.

Karpenter: the right nodes, and no surprises mid-event

Karpenter is excellent at provisioning fast and at consolidating nodes to save money. That second part is exactly what you don't want during a match. A disruption budget with a schedule freezes voluntary disruption for the event window, while still letting Karpenter add nodes.

apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
  name: event-burst
spec:
  template:
    spec:
      nodeClassRef: { group: karpenter.k8s.aws, kind: EC2NodeClass, name: default }
      requirements:
        - { key: karpenter.sh/capacity-type, operator: In, values: [spot, on-demand] }
        - { key: karpenter.k8s.aws/instance-category, operator: In, values: [c, m] }
        - { key: karpenter.k8s.aws/instance-generation, operator: Gt, values: ["5"] }
        - { key: kubernetes.io/arch, operator: In, values: [amd64, arm64] }
  limits:
    cpu: "4000"                           # a ceiling, so a bug can't scale you into a huge bill
  disruption:
    consolidationPolicy: WhenEmptyOrUnderutilized
    consolidateAfter: 10m
    budgets:
      - nodes: "10%"                    # normal days: gentle churn
      - nodes: "0"                      # event window: no voluntary disruption
        schedule: "0 13 * * *"           # UTC; 18:30 IST
        duration: 5h

Two details people miss. First, allow many instance types and both architectures if your images are multi-arch, because a narrow list is how you hit "insufficient capacity" for one instance type in one AZ at the worst moment. Second, keep the critical path on on-demand capacity with a separate NodePool and use Spot for the burst, so a Spot interruption wave can only ever cost you extra capacity, never your baseline.

Headroom pods: buy the 90 seconds back

Even with Karpenter, a pending pod waits for a node to launch, join, pull images and pass readiness. The standard trick is overprovisioning: run placeholder pods at a negative priority. They reserve real nodes. When a real pod needs room, the scheduler evicts a placeholder instantly, the real pod starts on a node that already exists, and Karpenter replaces the evicted placeholder's capacity in the background.

apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata: { name: headroom }
value: -10
preemptionPolicy: Never
globalDefault: false
---
apiVersion: apps/v1
kind: Deployment
metadata: { name: headroom, namespace: kube-system }
spec:
  replicas: 20                           # scale this up with a KEDA cron before the event
  selector: { matchLabels: { app: headroom } }
  template:
    metadata: { labels: { app: headroom } }
    spec:
      priorityClassName: headroom
      terminationGracePeriodSeconds: 0
      containers:
        - name: pause
          image: registry.k8s.io/pause:3.9
          resources: { requests: { cpu: "2", memory: 4Gi } }

One more thing that matters as much as any of this: pre-pull your images. A DaemonSet that pulls the big application images onto every new node, or images baked into the AMI, removes the slowest step from node startup.

The stateful tier on EC2: scale it days ahead

This is the part most autoscaling articles skip. Kubernetes autoscalers do nothing for MongoDB, Kafka or Aerospike running on EC2. Capacity there is planned, not reactive, and the industry practice looks like this:

StoreWhat limits you at peakHow to prepare
MongoDB (replica set / shards)Connections, working set vs RAM, write concern latency, replication lagSize the working set to fit in memory, cap and pool connections per service, scale instance size or add shards well before the event, and keep secondaries in another AZ and region.
Kafka on EC2Partition count, broker disk and network throughput, consumer lagSet partitions for peak consumer parallelism in advance (you can't easily shrink them later), add brokers and rebalance days before, and use KEDA on consumer lag for the consumers.
AerospikeMemory and index headroom, device throughput, migrationsAdd nodes early so data migrations finish long before the event, keep headroom below the high-water marks, and never trigger a migration on match day.
Amazon RDSConnections, IOPS, CPU, replica lagScale the instance class ahead (it causes a failover), add read replicas, put a connection pooler or RDS Proxy in front, and move hot reads to cache.

The golden rule for stateful systems on event day: no topology changes. No rebalances, no migrations, no instance resizes, no version upgrades. Everything that moves data around happens at least a few days before, while you still have time to recover if it goes wrong.

5. DR: picking a strategy by RTO and RPO

Disaster recovery starts with two numbers the business has to agree to, not engineering. RTO (recovery time objective) is how long you can be down. RPO (recovery point objective) is how much data you can afford to lose. AWS describes four standard strategies in its DR whitepaper, and every serious plan maps to one of them:

StrategyWhat runs in the DR regionTypical RTOTypical RPOCost
Backup and restoreNothing. Backups are copied there.HoursHours (last backup)Lowest
Pilot lightData replicated; compute offTens of minutesSeconds to minutesLow
Warm standbyA scaled-down but working copyMinutesSecondsMedium
Multi-site active/activeFull capacity serving trafficNear zeroNear zeroHighest

For a platform with a hard event date, warm standby is the sweet spot most teams land on. The pieces that make it work on AWS:

  • MongoDB: place replica set members in the second region (with lower priority so they don't become primary by accident), plus scheduled snapshots or mongodump/PITR backups copied cross-region. Failover is a controlled reconfig or an election, not a restore.
  • Kafka: MirrorMaker 2 replicates topics and consumer group offsets to a cluster in the second region. Decide upfront which topics need replication at all; many event streams can simply be rebuilt.
  • Aerospike: Cross-Datacenter Replication (XDR) ships writes asynchronously to the remote cluster, so RPO is the shipping lag. Back it with periodic asbackup to S3.
  • Amazon RDS: a cross-region read replica you can promote, plus automated snapshots copied cross-region with AWS Backup. Promotion is one-way, so plan the fail-back.
  • Kubernetes state: the cluster itself should be rebuildable from Terraform and GitOps in minutes. Anything that isn't in Git, like PersistentVolumes, gets backed up with Velero or AWS Backup and copied cross-region.
  • Objects and secrets: S3 cross-region replication, and secrets replicated to the second region (Secrets Manager supports replica secrets), so a failover doesn't stall on a missing credential.
  • Traffic: health-checked DNS failover, either at the CDN or DNS provider (Akamai has its own traffic management) or with Route 53. Many teams use a manual, audited switch, such as Route 53 Application Recovery Controller, rather than a fully automatic one.

Automatic regional failover sounds attractive, but many teams deliberately keep the final switch manual. A flapping health check that moves traffic between regions during a match can cause more damage than the original fault. Automate everything up to the decision, then have one person press the button.

6. DR drills: backup-restore, failover and game days

A DR plan you haven't run is a document, not a capability. At Avizva, where we ran multi-region HA for EKS, MongoDB and Elasticsearch for SOC 2 healthcare clients, we got recovery in DR scenarios down to 30 minutes, and the only reason we knew that number was that we kept drilling it. Here is how I'd structure drills before any big event:

Four kinds of drill

  1. Backup and restore drill. Take last night's backups and restore them into a clean environment, on a schedule, without warning the team in advance. Time it, then check the data is actually usable: row and document counts, a few known records, and the application booting against it. This is the drill most teams skip and the one that most often fails, usually because of a missing permission, an expired key or a backup that was silently incomplete.
  2. Tabletop. The on-call team walks through the runbook in a room: "the primary database is unreachable, what do we do, who decides, how do we tell the business?" It's cheap, and it finds missing access, missing steps and unclear ownership.
  3. Component failover. Kill something real in staging, then in production during low traffic: terminate nodes, fail over the database writer, block one AZ with security group or NACL changes, or use AWS Fault Injection Service to inject failures in a controlled way.
  4. Region evacuation. Move real production traffic to the standby region and back, on a schedule, with the business informed. This is the only test that proves your RTO.

Backup and restore drill runbook

# Restore drill: prove backups are usable (run monthly, and 2 weeks before any big event)
T-0    Pick a random recent backup set; start the clock in #dr-drill
T+5m   Restore RDS snapshot to a new instance (copy from the DR region vault)
T+5m   Restore MongoDB snapshot / PITR to a scratch replica set on EC2
T+5m   Restore Aerospike from asbackup files in S3 to a test cluster
T+5m   Recreate Kafka topics from IaC; replay from MirrorMaker 2 copy if needed
T+40m  Validate: counts vs source, checksums on sample records, schema version
T+50m  Boot the app against restored data; run smoke tests
T+60m  Record actual restore time per store and any data gap (= real RPO)
After  Tear down; ticket every manual or slow step; update the runbook

Two things make restore drills honest. Restore from the copy in the DR region, not the local one, because that's what you'll have in a real regional outage. And restore with the permissions the on-call engineer actually has, not an admin role, so you find access problems in a drill instead of during an incident.

Region failover drill runbook

# DR drill: region evacuation (target RTO 15m, RPO < 1m)
T-0    Announce start in #incident-drill, start the clock
T+1m   Freeze deploys (Argo CD sync windows), confirm replication lag < 1s
T+2m   Scale Region B: KEDA cron override + Karpenter limits up
T+6m   Verify Region B pods Ready, synthetic checks green
T+7m   Promote RDS replica; step up MongoDB DR members; repoint consumers to mirrored Kafka
T+10m  Flip DNS / traffic routing to Region B behind the edge
T+12m  Watch error rate, p99, login success, payment success
T+15m  Declare recovered, record actual RTO and data loss
After  Fail back the same way; write up every step that was slow or manual

Three rules that make drills worth the effort:

  • Measure, don't estimate. Record actual RTO and RPO every time and track the trend. If a drill takes 40 minutes against a 15-minute target, that's the most useful thing you learned all month.
  • Rotate who runs it. If only one engineer can fail over the platform, you don't have DR, you have a person.
  • Every manual step becomes a ticket. The goal over time is a runbook where each step is a script or a pipeline job that someone triggers, not a sequence of console clicks.

7. Alerting and on-call for the big day

I carried on-call through live matches, and what made those shifts calm was preparation, not heroics. On event day you want dashboards that answer "are users OK?" in five seconds, and alerts on symptoms users feel, not on causes.

The best practice here comes from the Google SRE Workbook: alert on SLO burn rate over two windows. A fast burn (for example 14.4x your error budget over 1 hour, confirmed over 5 minutes) pages someone. A slow burn opens a ticket. This catches real problems quickly and ignores short blips.

groups:
- name: live-api-slo
  rules:
  - alert: LiveApiFastBurn
    expr: |
      (
        sum(rate(http_requests_total{service="live-api",code=~"5.."}[1h]))
        / sum(rate(http_requests_total{service="live-api"}[1h]))
      ) > (14.4 * 0.001)
      and
      (
        sum(rate(http_requests_total{service="live-api",code=~"5.."}[5m]))
        / sum(rate(http_requests_total{service="live-api"}[5m]))
      ) > (14.4 * 0.001)
    labels:  { severity: page }
    annotations:
      summary: "live-api burning 99.9% SLO budget fast"
      runbook: "https://runbooks.internal/live-api#fast-burn"

For the event itself, add a few things you wouldn't run day to day: a staffed war room with clear roles (incident commander, communications, one engineer per critical dependency), with hand-overs agreed if your team spans time zones like India and Germany, a pre-written status-page message, and a short list of "break glass" switches like feature flags to turn off expensive features (recommendations, heavy personalisation) if the database starts to struggle. Shedding a nice-to-have feature for 10 minutes is far better than an outage. Incident tooling helps here too. At Avizva, adopting Zenduty for routing and escalation kept our MTTR to about 30 minutes at most.

8. Keeping the bill sane

Headroom for a peak gets expensive quickly if you buy it the wrong way. What has worked for me is putting stateless workloads on Spot and covering the steady, stateful baseline with Reserved Instances (or Savings Plans). In a previous role, that mix cut our EC2 cost by roughly 25 to 30%. It also made the extra capacity for big days cheap enough that nobody argued about pre-scaling.

For Spot, the industry rules are well established: diversify across many instance types and every AZ, handle the two-minute interruption notice (Karpenter does this natively when you configure its interruption queue), set PodDisruptionBudgets so evictions are gradual, and never put anything stateful or singleton on it.

9. The event-day runbook

WhenWhat
T−4 weeksAgree on RTO/RPO and the traffic forecast with the business. First full load test through the real edge.
T−3 weeksFix the first bottleneck, re-test at 1.5x to 2x forecast. Raise AWS service quotas (EC2 vCPU, ALB, NAT, RDS connections) now, not on the day. Scale up MongoDB, Kafka and Aerospike so rebalances and migrations finish early.
T−2 weeksBackup and restore drill, then a region evacuation drill, both timed. Tabletop with the on-call rota across India and Germany. Fix every slow manual step.
T−1 weekChange freeze for platform components and no topology changes in the data tier. Karpenter disruption budgets and KEDA cron floors merged via GitOps. Final load test.
T−1 dayConfirm last backups completed and copied cross-region, and check replication lag on MongoDB, Kafka mirroring, Aerospike XDR and RDS. Warm caches and the CDN. Pre-pull images. Status-page draft ready.
T−2 hoursWar room open. Dashboards on screen. Headroom pods scaled up, standby region checked.
T−45 minCron pre-scale kicks in. Verify replica counts and node counts against the plan.
EventWatch SLO burn, p99, database connections and replication lag. Feature flags ready. One person makes decisions.
T+1 dayScale down deliberately. Blameless review: forecast vs actual, what scaled, what didn't, and what to automate next.

The short version

None of this is exotic. Big events rarely fail because of a clever new failure mode. They fail because a load test was skipped, a quota wasn't raised, a failover was never practised, or capacity was left to react to a spike it could never catch. If your biggest day is a few weeks out, there's still time to do every step above. Start with the load test and the DR drill, because those two will tell you everything else you need to fix.

← All posts