Skip to content
Prompt Words
History

Horizontal scaling

Also called: scale out, add replicas.

Handling more load by running more copies of the app, with a load balancer spreading requests across them. The other way, vertical scaling, makes one machine bigger: it often needs a restart, and there is a biggest size. Copies only work if any copy can serve any request, so state like sessions lives outside the app.

One server with 4 cores is at 100% CPU and requests are slow. Each button starts from there: make the one server bigger, or run 3 copies of it.

Servers: 1 · Cores each: 4 · CPU: 100% · Response: 2.1 s

One server with 4 cores is at 100% CPU. Requests wait in line: 2.1 s each.

    1 server at 100%

    Say it in a prompt

    Scale the web app horizontally: run it as 3 replicas of the same container image behind the load balancer, each with 2 CPU cores and 4 GiB of memory. Keep the app stateless (sessions in Redis, uploads in object storage), so any replica can serve any request, and make sure a replica can be added or removed without a restart of the others.

    Vague vs precise prompt

    Vague prompt

    the server is at 100% CPU, make it handle more

    Typical resultMoves the app to a bigger machine with twice the cores. It goes down for the switch, it is still one machine that can crash, and next year it needs an even bigger one.

    Precise prompt

    Run the app as 3 replicas of the same image behind the load balancer, 2 cores each. Move sessions to Redis and uploads to object storage so any replica can serve any request.

    Typical resultLoad spreads over 3 copies at about a third of the CPU each. Adding a 4th copy is one command, and one copy crashing leaves 2 that keep serving.

    Seen on

    • Kubernetes docs: Defines horizontal scaling as answering more load by running more Pods, and vertical scaling as giving the running Pods more CPU or memory.
    • Kubernetes docs: kubectl scale sets a new number of replicas for a Deployment or ReplicaSet, for example --replicas=3.
    • Azure Architecture Center: Design to scale out: add or remove instances to match demand, make sure any instance can handle any request (no session stickiness), and find bottlenecks such as the database first.

    You might describe it as

    • run more copies instead of a bigger server
    • add servers when one is not enough
    • spread the load over identical copies of the app

    Not to be confused with

    • Load balancer

      Horizontal scaling adds more copies of the app; a load balancer spreads the requests across those copies.

    • Autoscaling

      Horizontal scaling is adding copies of the app, by hand or by a rule; autoscaling is the rule that adds and removes them for you as the load changes.