Autoscaling
Also called: auto scaling, HPA, horizontal pod autoscaler, KEDA.
A rule that adds and removes copies of your app for you, based on a number you watch, such as CPU or the lag of a queue. It checks the number again and again, adds copies while it is too high, and removes them after it has stayed low for a while. A minimum and a maximum keep the count in a safe range.
Consumer pods read orders from Kafka. The rule: while the lag (messages waiting) is over 800, add a pod. Start a traffic spike, then wait for the cool-down.
Pods: 1 of 6 · Lag: 0 · Time: 0:00
Normal traffic: 1 pod keeps up and the lag is 0.
Say it in a prompt
Autoscale the order-consumer Deployment on Kafka consumer lag with a HorizontalPodAutoscaler (or KEDA): min 1, max 6 pods, because the orders topic has 6 partitions. Add at most 1 pod every 15 seconds while lag is over 800 messages. Scale down only after lag has stayed low for 5 minutes (stabilizationWindowSeconds: 300), 1 pod every 15 seconds. Vague vs precise prompt
Vague prompt
make the consumers scale automatically Typical resultAdds an autoscaler on CPU with a maximum of 20 pods. The consumers mostly wait on Kafka, not the CPU, so the lag grows and nothing scales; and when it does scale, a 6-partition topic leaves up to 14 of the 20 pods idle.
Precise prompt
Autoscale the order consumers on Kafka lag: min 1, max 6 (6 partitions). Add 1 pod every 15 s while lag > 800; scale down only after 5 minutes of low lag. Typical resultA spike adds pods one by one until the lag falls, never more pods than partitions, and the extra pods go away 5 minutes after it is calm, so a short dip doesn't make them flap.
Seen on
- Kubernetes docs: The HorizontalPodAutoscaler checks a metric every 15 seconds by default and changes the number of replicas; scale-down waits for a 300-second stabilization window, and policies can limit how many pods are added or removed per period.
- KEDA docs: Its Apache Kafka scaler scales a consumer on consumer-group lag (lagThreshold), and by default never runs more replicas than the topic has partitions.
You might describe it as
- add servers automatically when the queue backs up
- scale with the traffic, then shrink back when it is quiet
- let a rule decide how many copies to run
Not to be confused with
- Horizontal scaling
Horizontal scaling is adding copies; autoscaling is a rule that adds and removes them for you.