Skip to content
Prompt Words
History

System design

How big systems fit together.

0 of 14 learned

Scaling 0 / 6

Handling more users and more data.

Cache-aside

Also called: lazy loading, look-aside cache, lazy caching.

Read a product twice: the first read misses the cache and goes to the database, the second is served from Redis.

Cache hits: 0 · Database reads: 0

Nothing read yet. The cache starts empty.

    Hits 0 misses 0

    The app reads from a cache such as Redis first; on a miss it reads the database, stores the result in the cache and returns it. On a write it updates the database and then deletes the cached key, so the next read loads fresh data. The cache only ever holds data someone asked for.

    Say it in a prompt

    Add a cache-aside layer with Redis for GET /products/:id: read the key product:{id} first; on a miss, load the row from Postgres, store it with a 300-second TTL and return it. On update or delete, write to Postgres first, then delete the key. If Redis is down, read straight from Postgres.

    Vague vs precise prompt

    Vague prompt

    make the product page faster when lots of people open it

    Typical resultAdds an in-memory Map inside one server with no expiry. The other servers still hit the database, and edits don't show until a restart.

    Precise prompt

    Add cache-aside with Redis for GET /products/:id: read product:{id} first; on a miss load it from Postgres, cache it for 300s (TTL) and return it. On update, write Postgres, then delete the key.

    Typical resultRepeat reads skip Postgres, the database sees at most one query per product every 5 minutes plus one after each edit, and an edit shows on the next read.

    Seen on

    • Azure Architecture Center: Documents the Cache-Aside pattern: try the cache, load from the data store on a miss and add the item to the cache, and invalidate the cached item when the data changes.
    • Amazon ElastiCache: Calls the same strategy lazy loading: data goes into the cache only when it is requested, and adding a TTL keeps it from getting too stale.

    You might describe it as

    • check Redis first, then the database
    • only cache what people actually ask for
    • keep a copy of hot rows so the database gets fewer reads

    Not to be confused with

    • TTL (time to live)

      Cache-aside decides when data goes into the cache (on a miss); a TTL decides when it drops out.

    • CDN

      Cache-aside is a cache your app code fills from the database; a CDN caches whole HTTP responses on servers near the user.

    CDN

    Also called: content delivery network, edge cache, edge network.

    Users in Tokyo open a page with an image, one after another. The first one misses: the Tokyo edge fetches the image from the origin in Virginia. Everyone after that gets the edge's copy.

    Edge hits: 0 · Trips to the origin: 0

    Nothing loaded yet. The edge server in Tokyo has no copy of the image.

      V1 edge empty hits 0 misses 0

      A network of servers around the world that keeps cached copies of your files and responses close to users. Each request goes to a nearby edge server, and only a miss travels on to your own server (the origin). Cache-Control headers decide what it caches and for how long.

      Say it in a prompt

      Serve /assets/* through a CDN: put a content hash in every file name (app.3f9a1c.js) and send Cache-Control: public, max-age=31536000, immutable. For HTML pages send Cache-Control: public, max-age=0, s-maxage=60, stale-while-revalidate=300, and purge the HTML paths on every deploy.

      Vague vs precise prompt

      Vague prompt

      the site is slow for users in Asia, fix it

      Typical resultCompresses the images and minifies the JavaScript, but every request still crosses the ocean to one server in Virginia.

      Precise prompt

      Put /assets/* behind a CDN with hashed file names and Cache-Control: public, max-age=31536000, immutable. Give HTML Cache-Control: public, max-age=0, s-maxage=60, stale-while-revalidate=300, and purge the HTML paths on deploy.

      Typical resultScripts, styles and images come from an edge server near each user. Browsers ask the edge again for HTML; the edge refreshes its copy in the background after a minute and drops it on every deploy. The origin sees far fewer requests.

      Seen on

      • Cloudflare: Caches static files such as images, CSS and JavaScript by default, but not HTML or JSON, and follows the Cache-Control headers from your origin.
      • Amazon CloudFront: Routes each request to the edge location with the lowest latency; if the file isn't cached there, it fetches it from your origin.

      You might describe it as

      • put the images on servers closer to the users
      • the site loads slowly for people in other countries
      • serve static files from the edge

      Not to be confused with

      • Cache-aside

        A CDN caches whole HTTP responses near the user; cache-aside is your app caching database results next to the database.

      • Load balancer

        A CDN answers from its own cache when it can; a load balancer always passes the request on to one of your servers.

      Load balancer

      Also called: LB, application load balancer, ALB.

      The balancer hands requests to the 3 servers in turn: 1, 2, 3, then 1 again (round robin). Take server 2 down to see health checks find it and take it out.

      Server 1: 0 · Server 2: 0 · Server 3: 0 · Failed: 0 · Balancer sends to: 1, 2, 3

      No requests yet. All 3 servers are healthy.

        Served 0 0 0 failed 0 healthy 1 2 3

        A server in front of several copies of your app that passes each request to one of them, using a rule such as round robin or least connections. It runs health checks and stops sending traffic to a copy that fails them.

        Say it in a prompt

        Put an application load balancer in front of 3 app instances with round-robin routing. Health-check GET /healthz every 10 seconds; take an instance out after 2 failures in a row and put it back after 3 passes. Keep the app stateless: store sessions in Redis, not in server memory.

        Vague vs precise prompt

        Vague prompt

        we have more users now, make the server handle it

        Typical resultRaises the worker or thread count on the one server. When that server goes down, the whole site goes down with it.

        Precise prompt

        Run 3 app instances behind an application load balancer with round-robin routing. Health-check GET /healthz every 10s; remove after 2 failures, re-add after 3 passes. Move sessions to Redis so any instance can serve any user.

        Typical resultTraffic splits across 3 instances, a crashed instance stops getting requests within about 20 seconds, and users stay logged in whichever instance answers.

        Seen on

        • NGINX: Its upstream module spreads requests round-robin by default, with least-connected and ip-hash as options, and marks a server as failed after max_fails failed attempts.
        • AWS Application Load Balancer: Sends health-check requests to every target and routes only to healthy ones; by default a target is taken out after 2 failed checks in a row.

        You might describe it as

        • spread the traffic over several servers
        • one address in front of many app copies
        • stop sending users to the server that crashed

        Not to be confused with

        • CDN

          A load balancer forwards every request to one of your servers; a CDN answers from its own cached copy when it can and only asks your server on a miss.

        Read replica

        Also called: replica, read-only replica, secondary, follower.

        Writes go to the primary; the replicas copy each change 2 s later (slow on purpose). A read in that gap gets old data.

        Writes: 0 · Reads: 0 (0 stale) · Replicas are up to date

        The primary and both replicas have price $10.

          Writes 0 reads 0 stale 0

          A read-only copy of a database that the primary keeps updated, usually asynchronously. You send heavy reads such as listings and reports to it, so the primary has room for writes. Because copying takes time, a replica can be a little behind the primary (replica lag).

          Say it in a prompt

          Add a Postgres read replica and route read-only queries (search, product listings, admin reports) to it through a second connection pool. Keep all writes on the primary. For read-your-writes, send a user's reads to the primary for 10 seconds after they write, so the page that loads right after saving shows the change. Alert when replica lag goes over 5 seconds.

          Vague vs precise prompt

          Vague prompt

          the database is slow, make it handle more traffic

          Typical resultSuggests a bigger database server or adds indexes here and there. Reads and writes still compete on one machine.

          Precise prompt

          Add a Postgres read replica. Send read-only queries (search, listings, admin reports) to it through a separate pool. Keep writes on the primary, and for 10s after a user writes, send that user's reads to the primary too (read-your-writes). Alert if replica lag passes 5s.

          Typical resultReports and listings move off the primary, writes get its full capacity, and a user who just saved still sees their change because that read goes to the primary.

          Seen on

          • Amazon RDS: A read replica is a read-only copy of a DB instance. RDS copies the primary's updates to it asynchronously, and you route read queries to it to take load off the primary.
          • A reporting dashboard runs its slow queries against a replica, so checkout writes on the primary stay fast.

          You might describe it as

          • a second copy of the database just for reading
          • send the heavy report queries somewhere else
          • database that follows the main one

          Not to be confused with

          • Sharding

            A read replica is a full copy of the data that only serves reads; sharding splits the data so each database holds a different part, which spreads writes too.

          Sharding

          Also called: horizontal partitioning, partitioning, shards.

          hash(user_id) % 6 picks a logical shard, and a lookup table says which database (shard) holds it.

          Shard 0: 3 users · Shard 1: 3 users · Shard 2: 3 users

          9 users on 3 shards, placed by hash(user_id) % 6 and a lookup table.

            Rows 3 3 3

            Splitting one big dataset across several databases by a shard key, so each database (a shard) holds only its part. All reads and writes for one key go to one shard, which spreads both the data and the write load. Queries that need many shards get slower and harder to write.

            Say it in a prompt

            Shard the orders table by tenant_id: logical_shard = hash(tenant_id) % 64, and a lookup table maps the 64 logical shards onto 8 Postgres databases. Every query must include tenant_id and go through one shard-router module. Moving a shard means copying it and updating the lookup table, with no rehashing.

            Vague vs precise prompt

            Vague prompt

            our orders table is huge, split it up so it scales

            Typical resultCreates orders_2024 and orders_2025 tables on the same server. The disk and the write load stay on one machine.

            Precise prompt

            Shard orders by tenant_id: logical_shard = hash(tenant_id) % 64, with a lookup table mapping the 64 logical shards onto 8 Postgres databases. All queries include tenant_id and go through one shard-router module.

            Typical resultEach database holds about an eighth of the tenants and their writes, one tenant's queries touch one database, and a shard can be moved by editing the lookup table.

            Seen on

            • Notion: Split its Postgres data into 480 logical shards spread over 32 physical databases, with the workspace ID as the shard key.
            • MongoDB: A sharded cluster spreads a collection's documents across shards by a shard key, and the mongos router sends each query to the right shard.

            You might describe it as

            • split the customers across several databases
            • each database only holds some of the users
            • one table is too big for one server

            Not to be confused with

            • Read replica

              Sharding gives each database a different slice of the data; a read replica is a full copy that only serves reads.

            TTL (time to live)

            Also called: time to live, expiry, expiration time, max-age.

            The key lives 10 s here (the prompt uses 300). Until then, reads come from Redis, even if Postgres changed.

            product:1 is not in Redis · Cache hits: 0 · Database reads: 0

            Redis is empty. Read the key to load it from Postgres.

              Empty

              A lifetime in seconds attached to a cached value, a DNS record or a message; when it runs out, the copy is deleted or treated as missing. A short TTL means fresher data and more load on the source; a long TTL means less load and staler data.

              Say it in a prompt

              Set a TTL on every cache write: 300 seconds for product pages, 30 seconds for stock counts, 24 hours for the country list. Add ±10% random jitter to each TTL so keys written together don't all expire in the same second, and delete a key right away when its row changes.

              Vague vs precise prompt

              Vague prompt

              make sure the cache doesn't show old data

              Typical resultClears the whole cache whenever anything changes, or turns caching off for that page, so the database load comes straight back.

              Precise prompt

              Give each cache entry a TTL: 300s for product pages, 30s for stock counts. Add ±10% jitter so keys don't expire together, and delete a key as soon as its row is updated.

              Typical resultStock is never more than 30 seconds old and product pages at most 5 minutes, and expiries spread out instead of hitting the database all at once.

              Seen on

              • Redis: EXPIRE key seconds sets a timeout, after which the key is deleted automatically; the TTL command shows how many seconds are left.
              • Cloudflare DNS: Every DNS record has a TTL that controls how long it is cached; the Auto setting is 300 seconds.
              • Amazon CloudFront: By default a file stays in an edge location for 24 hours before it expires, unless the origin's headers say otherwise.

              You might describe it as

              • how long before the saved copy counts as old
              • auto-delete the cached value after a while
              • expire the key after five minutes

              Not to be confused with

              • Cache-aside

                A TTL is only the expiry timer on an entry; cache-aside is the whole read, miss, load and store routine that usually sets one.

              Messaging 0 / 5

              Passing work between services.

              Backpressure

              Also called: back pressure, flow control.

              The producer reads a CSV file at 10 rows a second, the consumer inserts 2 rows a second. Import with backpressure off, then turn it on and import again.

              In the buffer: 0 · Inserted: 0 · Peak: 0

              Nothing imported yet. Backpressure is off.

                Idle bp off peak 0

                A way for a slow consumer to make a fast producer slow down, instead of letting work pile up until memory runs out. The consumer only accepts what it can handle, and the producer waits, keeps a bounded buffer or drops the extra.

                Say it in a prompt

                Add backpressure to the CSV import: read rows as a stream into a bounded buffer of 1,000 rows, pause reading while the buffer is full and resume when it drains below 500. Insert in batches of 200 with at most 4 inserts in flight.

                Vague vs precise prompt

                Vague prompt

                the import crashes on big files, fix it

                Typical resultRaises the Node memory limit to 8 GB. A bigger file crashes it again, because the whole file still loads before the first insert.

                Precise prompt

                Stream the CSV with backpressure: a bounded buffer of 1,000 rows, pause reading when it is full, resume below 500. Insert in batches of 200 with at most 4 in flight.

                Typical resultMemory stays flat whatever the file size, reading pauses whenever the database falls behind, and a 10 GB file imports like a small one, only slower.

                Seen on

                • Node.js: Its streams guide explains backpressure: write() returns false once the buffer passes highWaterMark (16 KB by default), and the source should wait for the 'drain' event.
                • Reactive Streams: A standard for asynchronous stream processing with non-blocking back pressure, so the queues between threads can stay bounded.

                You might describe it as

                • the importer runs out of memory on big files
                • slow down the sender when the receiver can't keep up
                • work piles up faster than we can process it

                Not to be confused with

                • Rate limiting

                  Backpressure slows the sender down based on how busy the receiver is right now; rate limiting is a fixed allowance per client, and extra requests are refused.

                Dead-letter queue

                Also called: DLQ, dead letter exchange, poison message queue.

                Send a few emails and one bad one. The bad one fails 3 times, then moves to the dead-letter queue, so it stops looping and the others keep going. The prompt says 5 tries; the demo uses 3 to keep it short.

                Sent: 0 · In the queue: 0 · In the DLQ: 0

                The queue is empty. Each message gets up to 3 tries (maxReceiveCount 3).

                  Sent 0 dlq 0

                  A side queue where a message goes after it has failed a set number of times, so one broken message stops looping in the main queue. You inspect it, fix the bug and send it back (redrive).

                  Say it in a prompt

                  Give the emails queue a dead-letter queue: after 5 failed receives (maxReceiveCount 5), move the message to emails-dlq. Keep DLQ messages for 14 days, alert as soon as the DLQ is not empty, and add an admin command that redrives its messages to the main queue after a fix.

                  Vague vs precise prompt

                  Vague prompt

                  some jobs fail and retry forever, stop that

                  Typical resultWraps the handler in try/catch and quietly drops the job on any error. The failed emails are gone and nobody knows.

                  Precise prompt

                  Add a dead-letter queue: after 5 failed receives (maxReceiveCount 5), move the message to emails-dlq. Keep it 14 days, alert when the DLQ is not empty, and add a command to redrive messages after a fix.

                  Typical resultA broken message stops after 5 tries instead of looping, nothing is lost, someone gets alerted, and the fixed messages replay with one command.

                  Seen on

                  • Amazon SQS: A redrive policy's maxReceiveCount sets how many times a message can be received before SQS moves it to the dead-letter queue; redrive moves messages back out.
                  • RabbitMQ: Dead-letters a message that is rejected without requeue, expires by TTL, overflows the queue's length limit, or passes a quorum queue's delivery limit.

                  You might describe it as

                  • a bad message keeps crashing the worker over and over
                  • put the failed jobs aside to look at later
                  • where messages go after too many retries

                  Not to be confused with

                  • Queue

                    The dead-letter queue only receives messages that failed too many times; the main queue holds the normal work.

                  Fan-out

                  Also called: fanout, fan-out on write, scatter.

                  When Ana posts, the API saves it once and answers her. A background worker then writes the post into each follower's feed. More followers means more writes, but opening a feed stays one quick read.

                  Followers: 4 · No posts yet

                  Ana hasn't posted yet. Each follower has an empty feed.

                    Posts 0 writes 0 followers 4

                    Sending one incoming message or request out to many receivers at once. Fan-out on write copies a new post into every follower's feed when it is created, so reading a feed later is one quick lookup.

                    Say it in a prompt

                    When a user posts, fan out on write: a background worker pushes the post ID onto each follower's Redis list feed:{followerId} and trims it to the newest 800. For accounts with more than 10,000 followers, skip the fan-out and merge their latest posts in at read time.

                    Vague vs precise prompt

                    Vague prompt

                    build a home feed that shows posts from people I follow

                    Typical resultRuns one big query joining follows and posts on every page load, and the feed gets slower with every user who signs up.

                    Precise prompt

                    Fan out on write: when a user posts, a worker pushes the post ID to feed:{followerId} in Redis for each follower (trim to 800). For accounts with over 10,000 followers, skip fan-out and merge their posts at read time.

                    Typical resultOpening the feed is one Redis read, posting costs one write per follower in the background, and a celebrity post doesn't trigger millions of writes.

                    Seen on

                    • Amazon SNS: In its fanout scenario, a message published to a topic is copied to many endpoints, such as several SQS queues, for parallel processing.
                    • Twitter: Its Timelines at Scale talk, summarized by High Scalability, describes a fanout service inserting each new tweet ID into every follower's home timeline list in Redis, capped at 800 entries.

                    You might describe it as

                    • copy the new post into every follower's feed
                    • one event goes out to many places at once
                    • send the same message to all the queues

                    Not to be confused with

                    • Pub/sub

                      Fan-out is the copying of one message to many receivers; pub/sub is a messaging setup that does that copying through a topic.

                    Pub/sub

                    Also called: publish/subscribe, topic, event bus.

                    Checkout sends each event once, to a topic, and doesn't know who listens. Every subscriber gets its own copy in its own queue, so one that is down only falls behind.

                    Events published: 0 · Copies delivered: 0

                    3 services subscribe to order.placed. Nothing published yet.

                      Published 0 subscribers 3

                      A publisher sends an event to a topic without knowing who listens, and every subscriber to that topic gets its own copy. You can add a new consumer without changing the publisher.

                      Say it in a prompt

                      When an order is placed, publish an order.placed event with {orderId, userId, total} to a topic. Subscribe three services: email sends the receipt, inventory reserves stock, analytics records the sale. Each subscriber gets its own queue and retries on its own, and ignores an orderId it has already handled.

                      Vague vs precise prompt

                      Vague prompt

                      after an order, send an email and update the stock and analytics

                      Typical resultCalls the three services one after another inside the checkout request. If analytics is down, checkout fails.

                      Precise prompt

                      Publish order.placed {orderId, userId, total} to a topic. Subscribe email, inventory and analytics, each with its own queue and retries, each ignoring orderIds it has seen. Checkout only waits for the publish.

                      Typical resultCheckout returns as soon as the event is published, and a slow or broken analytics service falls behind on its own without blocking orders or emails.

                      Seen on

                      • Google Cloud Pub/Sub: Publishers send events to a topic without regard to how they will be processed, and Pub/Sub delivers them to all the services that subscribe.
                      • Redis: PUBLISH pushes a message to every client subscribed to a channel; delivery is at-most-once, so a subscriber that is disconnected misses it.

                      You might describe it as

                      • tell every service when something happens
                      • announce an event and whoever cares listens
                      • add a new listener without touching the sender

                      Not to be confused with

                      • Queue

                        Pub/sub gives every subscriber a copy of each event; a queue gives each message to only one worker.

                      • Fan-out

                        Pub/sub is the setup of topics and subscribers; fan-out is the copying of one message to many receivers, which a topic does for you.

                      Queue

                      Also called: message queue, job queue, work queue, task queue.

                      The web server answers each upload at once with 202 Accepted. A worker resizes one photo every 2 seconds and deletes the job only when it's done. If the worker crashes first, the job comes back after the visibility timeout (5 seconds here).

                      Waiting: 0 · In flight: 0 · Done: 0

                      No jobs yet. Worker 1 is waiting for work.

                        Waiting 0 flight 0 done 0 workers 1

                        A list of jobs that producers add to and workers take from, so slow work happens in the background instead of inside the request. Each message goes to one worker, and a message that isn't confirmed in time goes back on the queue for another try.

                        Say it in a prompt

                        Move image resizing out of the upload request into an SQS queue. The API saves the file, enqueues {imageId} and returns 202 right away. Run 4 workers with a 60-second visibility timeout, and delete a message only after the resize succeeds. After 5 failed receives, move it to a dead-letter queue.

                        Vague vs precise prompt

                        Vague prompt

                        uploading a photo takes forever, make it faster

                        Typical resultSwaps in a faster image library but still resizes inside the request, so a big upload still times out after 30 seconds.

                        Precise prompt

                        Put image resizing on an SQS queue: the API stores the file, enqueues {imageId} and returns 202. Run 4 workers with a 60s visibility timeout; delete the message only after success; move it to a DLQ after 5 receives.

                        Typical resultThe upload answers in well under a second, 4 workers resize in the background, and a crashed worker's job reappears for another worker after 60 seconds.

                        Seen on

                        • Amazon SQS: A received message stays hidden from other consumers for the visibility timeout (30 seconds by default); if it isn't deleted in time, it becomes visible again for another try.
                        • RabbitMQ: Its work queues tutorial hands each task to the next worker in turn, and redelivers a task to another worker if the first one dies before acknowledging it.

                        You might describe it as

                        • do the slow part in the background
                        • a to-do list that the workers pick jobs from
                        • don't make the user wait for the email to send

                        Not to be confused with

                        • Pub/sub

                          In a queue each message goes to one worker; in pub/sub every subscriber gets its own copy.

                        • Dead-letter queue

                          The queue holds work waiting to be done; the dead-letter queue holds the messages that kept failing.

                        Reliability 0 / 3

                        Staying up when parts fail.

                        Circuit breaker

                        Also called: breaker, fail fast.

                        The payments API is down. After 3 failed calls in a row the breaker opens: calls fail fast for 5 seconds, then one test call goes through.

                        Closed Calls go through. 0 of 3 failures.

                          Closed 0 api down

                          A wrapper around calls to another service that counts failures and, past a threshold, stops calling it for a while and fails fast instead (open). After a wait it lets a few test calls through (half-open), and if they succeed, normal calls resume (closed).

                          Say it in a prompt

                          Wrap calls to the payments API in a circuit breaker: 2-second timeout; open when 50% of the last 20 calls fail or time out; while open, fail fast for 30 seconds with a 'try again soon' error; then go half-open and allow 3 test calls, closing if all succeed and opening again if any fails.

                          Seen on

                          • Netflix Hystrix: Opens the circuit when request volume and error percentage pass their thresholds, short-circuits calls for a sleep window, then lets a single test request through (half-open).
                          • Resilience4j: Its CircuitBreaker has CLOSED, OPEN and HALF_OPEN states; by default it opens at a 50% failure rate, waits 60 seconds, then allows 10 calls in half-open.

                          You might describe it as

                          • stop calling the service that keeps timing out
                          • fail fast when the other API is down
                          • one broken service drags everything down with it

                          Not to be confused with

                          • Graceful degradation

                            A circuit breaker decides to stop calling a failing service; graceful degradation decides what the user sees instead.

                          • Rate limiting

                            A circuit breaker protects you from a dependency that is failing; rate limiting protects your service from callers sending too much.

                          Graceful degradation

                          Also called: degraded mode, fail soft, fallback.

                          The product is the core of the page; recommendations are an extra. Take Recommendations down and load the page again: the product still shows, with Popular items in place of personal picks.

                          Pages loaded: 0 · With a fallback: 0 · Error pages: 0

                          Not loaded yet. Both services are up.

                            Idle recs up

                            Designing a service so that when a dependency fails, its core job keeps working and only the extras are dropped or simplified. The page still loads with stale, default or missing parts instead of showing an error.

                            Say it in a prompt

                            Make the product page degrade gracefully: give the recommendations, reviews and stock services a 300 ms timeout each. If recommendations fail, show the cached bestseller list; if reviews fail, hide the reviews block; if stock fails, show 'Check availability at checkout'. Only a missing product returns an error page.

                            Vague vs precise prompt

                            Vague prompt

                            the product page shows an error when the reviews service is down, fix it

                            Typical resultAdds a retry loop around the reviews call. The page now takes 10 seconds and then shows the same error.

                            Precise prompt

                            Degrade gracefully: give recommendations, reviews and stock a 300ms timeout each. On failure show cached bestsellers, hide reviews, and show 'Check availability at checkout'. Only a missing product returns an error page.

                            Typical resultWhen reviews or recommendations are down, the page still loads quickly with those parts hidden or cached, and people can keep buying.

                            Seen on

                            • AWS Well-Architected Framework: Best practice REL05-BP01 asks components to keep their core function when dependencies fail, for example a store's landing page still showing everything else when the recommendations system is down.
                            • A search page still shows its results, without the 'people also bought' box, when the recommendations service times out.

                            You might describe it as

                            • show the page even if one part of it is broken
                            • hide the reviews box instead of erroring when that service is down
                            • keep checkout working when the extras fail

                            Not to be confused with

                            • Circuit breaker

                              Graceful degradation is what the user gets while a part is down; a circuit breaker is one way to notice the failure and stop calling that part.

                            Rate limiting

                            Also called: throttling, token bucket, request quota, API rate limit.

                            The bucket holds 5 tokens and gets 1 back every second. Each request spends one; with none left, the answer is 429.

                            No requests yet.

                              Tokens 5

                              Capping how many requests each client may make in a period and refusing the rest, usually with HTTP 429 Too Many Requests. A token bucket allows short bursts: every request spends a token, and tokens refill at a steady rate up to a fixed capacity.

                              Say it in a prompt

                              Rate-limit the public API per API key with a token bucket in Redis: capacity 20 tokens, refilled at 10 tokens per second. When the bucket is empty, return 429 Too Many Requests with a Retry-After header in seconds, and send X-RateLimit-Remaining on every response.

                              Seen on

                              • Stripe: Allows 100 API requests per second per account in live mode and 25 in a sandbox, answers extra requests with 429 Too Many Requests, and suggests a client-side token bucket.
                              • Amazon API Gateway: Throttles with the token bucket algorithm: the rate is how fast tokens are added and the burst is the bucket's capacity; throttled clients get 429.
                              • GitHub REST API: Authenticated users get 5,000 requests per hour; the x-ratelimit-remaining and x-ratelimit-reset headers show what is left and when it resets.

                              You might describe it as

                              • stop one client from flooding the API
                              • only let each key call it 10 times a second
                              • a bot is hammering our endpoint

                              Not to be confused with

                              • Backpressure

                                Rate limiting is a fixed allowance per client, and requests over it are refused; backpressure slows the sender down based on how busy the receiver is right now.

                              • Circuit breaker

                                Rate limiting protects your service from callers sending too much; a circuit breaker protects you from a dependency that is failing.