# EP 72: How Netflix Ensures Reliability with Prioritized Load Shedding
*Netflix ensures reliability by shedding low-priority requests during stress while keeping streaming smooth, validated through Chaos Engineering.*
By [Rohit Lakhotia](https://scaleengineer.com/authors/rohit-lakhotia)
Published: 2025-03-31
Canonical: https://scaleengineer.com/blog/how-netflix-ensures-reliability-with-prioritized-load-shedding
---
![](https://www.vpdae.com/open/eac649b4.gif?opens=1)

![](https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/b98cd2d7-0e8e-4d97-9e23-146208304a19/ist3zpxjq5g7nb9zgyl6y09304qf?t=1775937244)

### Seamlessly integrate your tools and services on HubSpot

With flexible UI and extensibility tools, HubSpot’s developer platform allows you to build apps that help teams unify their tech stack across our platforms\. From listing your app in the HubSpot Marketplace to building custom solutions, there’s something for everyone to grow better\.

[ Learn More](https://www.vpdae.com/redirect/5e6kkzkdi2ceqd3vdeatfhz3um)

Discover how Netflix set out to be more resilient by consistently prioritizing requests across device types, progressively throttling requests based on priority, and validating assumptions for requests of specific priorities\.

Checkout this video from **Amazon Web Services \(AWS\) re:Invent session**\.

[Embedded content](https://youtube.com/embed/TmNiHbh-6Wg)

Imagine you're driving on a highway and suddenly all the vehicles slow down as if they are crawling and apparently you are stuck in a long traffic jam\. Sometimes it's because of an accident but other times there's no clear reason, it’s just congestion\. Now, what if traffic management could prioritize cars** based on urgency**? Emergency vehicles, buses, or high\-priority travelers would move through, while others might wait\.

This is how **Netflix handles traffic on its backend services** by prioritizing critical requests when the system is under stress which ensures that viewers can always watch their favorite shows, even when failures occur in the background\.

## Why Netflix needed Load Shedding

Netflix’s infrastructure handles **millions of requests per second** from users on different devices \(TVs, phones, browsers\)\. However, it cannot run flawlessly all the time right, failures can happen due to:

- **Misbehaving clients**: When devices \(like a user’s mobile app or smart TV\) repeatedly retry failed requests too aggressively, they can overload the system\.
- **Under\-scaled backend services**: If a service doesn’t have enough resources to handle traffic, it can slow down or crash\.
- **Bad deployments**: A faulty software update can introduce bugs that disrupt services\.
- **Network issues**: Unstable network connections can delay responses or cause errors\.
- **Cloud provider issues**: Problems with infrastructure providers \(like AWS\) can impact Netflix’s ability to serve content\.

Any of these failures can create unexpected load on Netflix’s systems, sometimes making it impossible for users to stream content\. To prevent such disruptions, Netflix engineers set three key goals:

1. **Prioritize Requests Consistently**: No matter what device a user is on \(mobile, web browser, or TV\), requests should be prioritized properly\.
2. **Throttling Based on Priority**: When the system is under stress, it should drop lower\-priority requests first while keeping critical ones intact\.
3. **Validate Assumptions with Chaos Testing**: Netflix deliberately simulates failures \(known as Chaos Testing\) to ensure their system behaves as expected under stress\.

To explain you in short,

Previously, Netflix used **[circuit breakers](https://netflixtechblog.com/introducing-hystrix-for-resilience-engineering-13531c1ab362)** which acted like an on/off switch, cutting off requests when a service was overloaded\. But this approach was **too harsh** because it didn’t **prioritize important traffic** over non\-essential requests\.

To fix this, Netflix engineers introduced **priority\-based progressive load shedding**, which ensures that only the least important requests are dropped, keeping streaming uninterrupted\.

*An example of a request that can be dropped is requests for show trailers\. When a user is scrolling through Netflix, the website will autoplay the trailer of whatever movie/TV show the user currently has selected\.*

*When Netflix's servers are under high stress, they make a very clever trade\-off\. While browsing, users normally see autoplay trailers for shows they hover over\. But during system strain, Netflix simply drops these trailer requests\. And this small compromise prevented system overload, maintained core user experience while ensuring critical streaming functions remain smooth\.*

## How Netflix implemented prioritized load shedding?

In order to implement prioritized load shedding, Netflix engineers had these 3 steps in their minds,

1. **Define a Request Taxonomy: **Find a way to categorize requests by priority and assign a score to each request that describes how critical the request is to the user streaming experience\.
2. **Implement the Load Shedding Algorithm: **Netflix decided to implement the algorithm in their API Gateway, Zuul
3. **Validate Assumptions using Fault Injection: **Use Chaos Engineering principles to test and refine system resilience through controlled failure simulations\.

![](https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/04790b62-e58e-4d01-81e3-f836465e13cb/image.png?t=1742930366)

### Define a Request Taxonomy

To decide which requests should be prioritized, Netflix classified them based on three key factors, **Throughput** \(How many requests are being made\), **Functionality** \(What the request is doing \(e\.g\., fetching a movie vs\. logging an event\)\), **Criticality** \(How much the request impacts playback\)

Further, Netflix categorized all incoming requests into three buckets based on **how important they were**:

1. **NON\_CRITICAL**: These requests don’t affect video playback at all\. They **consume a lot of resources** but can be dropped without users noticing\.
  
  *Examples: Logging, background requests, telemetry data\.*
2. **DEGRADED\_EXPERIENCE**: These requests impact user experience but **won’t stop playback**\. If dropped, users may experience **inconvenience** but can still watch\.
  
  *Examples: Pause/resume markers, language selection, viewing history\.*
3. **CRITICAL**: These requests **must** be served; otherwise, users **can’t stream content**\.
  
  *Example: Pressing "Play" if this request fails, playback won’t start\.*

To manage traffic efficiently, Netflix's **API gateway service \(Zuul\)** assigns **a priority score \(1–100\) to each request** based on these categories\. This priority is then used to **decide which requests get dropped first** when the system is overloaded\.

### Load Shedding Algorithm

The first task was to find the best place to throttle traffic\. When Netflix experiences high traffic, it needs to decide **where** to reduce the load to keep streaming smooth\. This can happen at two levels:

1. **Service Throttling**: Limiting traffic to a specific backend service\.
2. **Global Throttling**: Reducing traffic across all backend services\.

**Service\-Level Throttling**

If a **specific backend service** \(e\.g\., the video catalog service\) is struggling, Netflix **reduces the number of requests** being sent to it\.

Zuul detects issues using **error rates and concurrent request counts** and then limits traffic **only for that service**, rather than cutting off everything\.

**Global Throttling**

If **Zuul itself** is overloaded, it **throttles traffic globally** across all services\.

Zuul monitors three key signals to check if it’s in trouble: **CPU Usage**, **Concurrent Requests, Connection Count**\. If any of these numbers **cross a critical threshold**, Zuul takes action by **dropping low\-priority requests** to stay alive\.

This is crucial because if Zuul **goes down**, no requests can reach backend services, **all of Netflix goes down **leading to a **total outage**\.

![](https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/58037bab-db2a-4d21-81a2-e0805c111040/image.png?t=1742932599)

**Priority\-based progressive load shedding**

Once Netflix categorized requests by priority, it combines this with **load shedding** to make streaming even more reliable\. When Netflix faces high traffic or system strain, **it doesn’t just drop requests randomly**, it follows a **progressive approach**:

1. **Start with the lowest priority requests**: Features like logging, background updates, or non\-essential UI elements are the first to be dropped\.
2. **Gradually increase the priority threshold**: As system strain worsens, Netflix **sheds more traffic step by step**, ensuring that critical services keep running\.
3. **Use a cubic function to control throttling:** Instead of a sudden drop in service, Netflix slowly increases the number of requests it blocks\. But if things get really bad, the system **quickly escalates throttling** to prevent a total failure\.

You might be wondering, but why this works?

Most Netflix requests **don’t directly impact streaming availability**, so by progressively shedding lower\-priority requests, Netflix **keeps movies and shows running even under heavy system strain**\.

This ensures that, **members can always press "Play"** and watch their favorite shows, **non\-essential features \(like viewing history\) may slow down, but don’t break the experience and most importantly, Netflix’s servers stay stable** instead of crashing under pressure\.

### Retry Storms

A big issue with shedding traffic is that **clients \(devices\) might automatically retry failed requests** which can cause a **retry storm** and eventually overwhelming the system even more\.

To prevent this, **Zuul sends a signal** to devices, telling them:

- **How many times they can retry**
- **How long they should wait before retrying**

For example:

```
{
  "maxRetries": <max-retries>,
  "retryAfterSeconds": <seconds>
}
```

This resulted in **stopping unnecessary retries** \(Instead of devices spamming requests, they follow a structured retry pattern\), **Prioritizes critical traffic** \(High\-priority requests \(like pressing "Play"\) can retry more frequently, while non\-essential ones \(like background updates\) wait longer\)\.

### Validating the Request Taxonomy

To make sure Netflix correctly categorized requests as **NON\_CRITICAL, DEGRADED, or CRITICAL**, they needed a way to test what happens when certain requests are blocked\.

To do this, Netflix engineers used their **[Failure Injection Tool](https://netflixtechblog.com/fit-failure-injection-testing-35d8e2a9bb2)**** \(FIT\)** to **simulate request failures**\. They could **intentionally block certain types of requests** for specific users or devices to see **if shedding them impacted the viewing experience**\.

### Validate Assumptions using Fault Injection

We can see that Netflix is evolving rapidly day by day and what was once a **non\-critical request** might suddenly become essential someday\. Since Netflix runs on **a variety of devices and app versions**, they needed a way to **continuously validate** that shedding certain requests wouldn't disrupt streaming\.

Netflix uses **Fault Injection** experiments to test system resilience by introducing disruptions like traffic spikes, CPU load, and latency\. Netflix engineers didn’t just **assume** their new system worked; they **tested it using Chaos Engineering**:

**Chaos Tools**: Netflix built **Chaos Monkey** and **ChAP \(Chaos Automation Platform\)** to automate failure testing in production\.

1. **Failure Injection Testing \(FIT\)**
  
  They **manually blocked certain priority levels** for specific devices\. If blocking a request caused a noticeable issue, they **reclassified it** \(e\.g\., a "NON\_CRITICAL" request might actually be "DEGRADED"\)\.
  
  
2. **A/B Experimentation with ChAP, ** **[ChAP](https://netflixtechblog.com/chap-chaos-automation-platform-53e6d528371f)**, an **A/B testing platform**
  
  Netflix randomly **throttled requests** for a small group of real users and **monitored the impact**\.
  
    - If users didn’t notice, those requests were confirmed to be **safe to drop**\.
    - If users had issues, Netflix **adjusted the priority classification** to prevent real\-world problems\.

One early test revealed a **race condition in Android and iOS clients**, causing occasional playback failures\. Thankfully the team at Netflix were on to this **continuous experimentation**, they **fixed the bug** before it could impact millions of users\.

**How Netflix did this?**

1. **Created two groups**:
  
    - **Control group**: Normal streaming experience\.
    - **Treatment group**: Some non\-critical requests are throttled\.
2. **Ran the test for 45 minutes** across real users\.
3. **Measure the impact**:
  
    - ChAP tracks **key streaming performance indicators \(KPIs\)** across devices\.
    - If there's a significant difference between groups, it means **a "non\-critical" request was actually important**\.

It looks good in the books but let’s see some** Real\-World Benefits of Load Shedding,**

Netflix had a **major outage in 2019**, preventing a large percentage of users from streaming\.

In **2020**, after implementing **priority\-based load shedding**, a similar issue occurred\. This time, **Zuul automatically shed low\-priority traffic**, **keeping the service stable**\.

**The Impact:**

**Before Load Shedding \(2019\)**: Outage for many users\. **After Load Shedding \(2020\)**: Users kept watching without disruption, while the system self\-recovered\.

After the implementations were over, they did some analysis and based on their analysis,

**High\-priority requests \(like pressing "Play"\) remained unaffected while lower\-priority requests were progressively throttled until stability was restored\.**

Official blog from Netflix: [Keeping Netflix Reliable Using Prioritized Load Shedding \| by Netflix Technology Blog \| Netflix TechBlog](https://netflixtechblog.com/keeping-netflix-reliable-using-prioritized-load-shedding-6cc827b02f94)

By now, you must have had a clear idea of,** How Netflix Ensures Reliability with Prioritized Load Shedding? **In a nutshell, Netflix ensures reliability with **priority\-based load shedding** by throttling low\-priority requests during system strain while keeping critical streaming functions intact\. They validate this using **Chaos Engineering** to continuously test their changes and updates\.

**Congratulations\! You've just advanced another step in your tech journey\. Keep progressing\!**
