Skip to content
Prompt Words
History

Autoscaling

Also called: auto scaling, HPA, horizontal pod autoscaler, KEDA.

A rule that adds and removes copies of your app for you, based on a number you watch, such as CPU or the lag of a queue. It checks the number again and again, adds copies while it is too high, and removes them after it has stayed low for a while. A minimum and a maximum keep the count in a safe range.

Consumer pods read orders from Kafka. The rule: while the lag (messages waiting) is over 800, add a pod. Start a traffic spike, then wait for the cool-down.

Pods: 1 of 6 · Lag: 0 · Time: 0:00

Normal traffic: 1 pod keeps up and the lag is 0.

    1 pod

    Say it in a prompt

    Autoscale the order-consumer Deployment on Kafka consumer lag with a HorizontalPodAutoscaler (or KEDA): min 1, max 6 pods, because the orders topic has 6 partitions. Add at most 1 pod every 15 seconds while lag is over 800 messages. Scale down only after lag has stayed low for 5 minutes (stabilizationWindowSeconds: 300), 1 pod every 15 seconds.

    Vague vs precise prompt

    Vague prompt

    make the consumers scale automatically

    Typical resultAdds an autoscaler on CPU with a maximum of 20 pods. The consumers mostly wait on Kafka, not the CPU, so the lag grows and nothing scales; and when it does scale, a 6-partition topic leaves up to 14 of the 20 pods idle.

    Precise prompt

    Autoscale the order consumers on Kafka lag: min 1, max 6 (6 partitions). Add 1 pod every 15 s while lag > 800; scale down only after 5 minutes of low lag.

    Typical resultA spike adds pods one by one until the lag falls, never more pods than partitions, and the extra pods go away 5 minutes after it is calm, so a short dip doesn't make them flap.

    Seen on

    • Kubernetes docs: The HorizontalPodAutoscaler checks a metric every 15 seconds by default and changes the number of replicas; scale-down waits for a 300-second stabilization window, and policies can limit how many pods are added or removed per period.
    • KEDA docs: Its Apache Kafka scaler scales a consumer on consumer-group lag (lagThreshold), and by default never runs more replicas than the topic has partitions.

    You might describe it as

    • add servers automatically when the queue backs up
    • scale with the traffic, then shrink back when it is quiet
    • let a rule decide how many copies to run

    Not to be confused with

    • Horizontal scaling

      Horizontal scaling is adding copies; autoscaling is a rule that adds and removes them for you.