# How Spotify Powers Music Streaming for Millions
*Spotify uses Kafka, microservices, and ML to deliver real-time, personalized music to millions, powered by a fast, scalable cloud backend.*
By [Rohit Lakhotia](https://scaleengineer.com/authors/rohit-lakhotia)
Published: 2025-08-11
Canonical: https://scaleengineer.com/blog/how-spotify-powers-music-streaming-for-millions
---
Spotify is more than just a music app; it’s a complex system running behind the scenes to make sure you can listen to your favorite songs anytime, anywhere\. Let’s discover how Spotify's engineering and technology make this possible in this blog today\.

### Spotify's Tech Stack

Spotify's core backend is built on a **[microservices ](/blog/what-are-microservices)****architecture**\. This means instead of one big app, Spotify runs many smaller services, each handling a different job \(like user login, playlist management, or audio streaming\)\. This setup makes scaling easier and avoids one error crashing the entire system\.

- **Programming Languages**: Java and Scala are the go\-to languages, especially for backend services\.
- **Frameworks**: Spotify often uses the **Spring Framework** for Java and **Node\.js** for certain microservices\.

**Real\-time Music Streaming?** That’s where **[Apache Kafka](/blog/what-is-kafka)** comes in\! Kafka is a powerful real\-time data streaming tool that helps Spotify manage real\-time data streams, enabling it to deliver music quickly and smoothly to your device as you hit play\.

For storing all that data, Spotify relies heavily on **Apache Cassandra**, a NoSQL database\. It's known for being super scalable and fault\-tolerant, perfect for handling Spotify's massive user base\.

On the frontend, the web application is built using **React**, with tools like **Redux** and **Sass** helping manage app state and styling\.

Spotify originally hosted on **AWS** but later moved to **Google Cloud** for more scalability and better services\. It uses **Kubernetes** to orchestrate its containers, ensuring its microservices run efficiently\.

## Why They Moved from PostgreSQL to Cassandra

As Spotify grew, so did its data problems\. By 2011, after launching in the US, Spotify's user base exploded and with millions of people streaming music every day, they needed a backend system that could handle all that data reliably\.

Initially, they used **PostgreSQL** as their primary database\. But as user numbers skyrocketed, PostgreSQL couldn’t keep up\. The final push for change came when a transatlantic data cable broke \(yes, possibly from a shark bite\), cutting off data flow between key data centers\.

Yes, you heard it right\!\! In 2015, something bizarre happened, the undersea cable connecting Spotify’s data centers in London and Ashburn got **cut**\. Rumor has it, it might’ve been a **shark**\!

While that may sound wild, the impact was very real\. It forced Spotify to realize:

“If we want to scale and avoid such bottlenecks in the future, we need a better, more resilient database system\.”

That’s when **Apache Cassandra** came into the picture\.

**But WHY ****[Cassandra](/blog/what-is-cassandra)****?**

Spotify needed a database that could scale **horizontally**, meaning it could easily handle more users and more data just by adding more machines, without breaking things\. Cassandra was built for exactly that kind of situation\. It’s highly available, fault\-tolerant, and fast, perfect for a streaming giant\.

At the time, they had **35 million active users**, and the existing setup just couldn’t keep up\.

That’s when Spotify moved to **Cassandra**\. To make the transition smooth, Spotify engineers used a clever technique called **dark loading** where data is silently copied to the new system without affecting the live platform\. This allowed them to **test Cassandra in real\-time** without risking downtime or breaking features for users\.

Spotify even built a dedicated internal system for managing future migrations\. They automate as much as possible, plan priorities, and develop "migration products" to handle them with minimal disruption\.

### Scaling Up and Modernizing

After Spotify made its big move from PostgreSQL to Cassandra, they didn’t just stop there, they’ve kept pushing forward, scaling both their systems and user experience year after year\.

#### Quick Timeline of Spotify’s Client Evolution:

| Year | What Happened |
| --- | --- |
| **2006** | Spotify & desktop app created |
| **2012** | Web Player launched |
| **2013** | Desktop & Web Player split into separate teams |
| **2015** | Desktop UI rebuilt with web tech |
| **2017** | Desktop & Web teams merged |
| **2019** | Shared UI components introduced |
| **2021** | Unified apps released |

#### 2023: Overhauling the iOS Build System with Bazel

One of the biggest behind\-the\-scenes upgrades came in **2023**\. Spotify decided to replace the entire build system for its **iOS app** with **Bazel**, a powerful, open\-source build tool from Google\.

And this wasn’t a small change:

Over **120 engineering teams** were involved\.

Why Bazel? Because it’s **faster**, more **reliable**, and scales way better as teams and codebases grow\. For Spotify, this meant:

- Faster build times
- More stable releases
- Easier onboarding for developers
- And a smoother app experience for users

#### 2021: Unifying Web & Desktop Apps

Before 2023, Spotify had already made a huge step in **2021** by redesigning and releasing **new versions** of their desktop and web apps\.

This wasn’t just a UI change\. It was a full\-on structural shift:

- Both apps were rebuilt to share **one codebase**
- Engineers from both teams were combined into a **single department**
- They adopted a **modular, container\-based UI** system

This allowed Spotify to reuse components, speed up development, and make both apps **more consistent and faster**\.

#### 2024: Tapping Into ML & Generative AI

Spotify is also betting big on **machine learning** and **AI**\.

In **2024**, they launched a new system to **generate annotations** for the millions of songs, podcasts, and videos on the platform\. These annotations:

- Help improve recommendations
- Feed new training data into Spotify’s ML models
- And power future generative AI features

This shift means Spotify isn’t just organizing content anymore, it’s understanding it\.

### How Spotify Knows What You’ll Love

Spotify feels like it *just knows* what song to play next, right?? Every time you play, skip, like, or search a song, Spotify captures it as an **event**\. These events power their **event\-driven architecture**, which updates your preferences in real time\.

Then come the **machine learning models** trained on:

- What you and similar users like
- Song characteristics \(tempo, mood, etc\.\)
- Even text data from reviews or lyrics

Together, this builds your **Discover Weekly**, **AI DJ**, and personalized playlists so your Spotify always feels made *just for you*\.

### How Spotify Uses Event\-Driven Architecture to Personalize Your Music

Spotify wants every user's experience to feel unique and a big part of how they do that is through **event\-driven architecture**\.

#### What does that mean?

Event\-driven architecture is a system design where components **react to events**, basically changes or actions that happen in real time\.

In Spotify’s case, **every user interaction** is an event:

- Playing or pausing a song
- Skipping a track
- Liking a song
- Creating a playlist
- Searching for an artist

Each of these actions triggers an event that’s **captured and sent** through a tool called **Apache Kafka**, a system built for handling millions of real\-time messages\. This ensures Spotify reacts instantly to user behavior\.

![](https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/0a034783-2542-4f66-bffe-835291255819/Mermaid_Chart_-_Create_complex__visual_diagrams_with_text._A_smarter_way_of_creating_diagrams.-2025-07-26-131618.png?t=1753535831)

#### What happens to these events?

1. **Stored in Cassandra**: All this user data is stored in **Cassandra**, a highly scalable NoSQL database\. Cassandra helps Spotify handle massive volumes of data without slowing down\.
2. **Enriched with metadata**: Along with your actions, Spotify also collects metadata about each song like:
  
    - Genre
    - Instruments used
    - Tempo
    - Mood or emotion
    - Artist details
3. **Used for recommendations**: This combined data, what users do \+ what the content is — feeds into Spotify’s **recommendation engines**\. That’s how they create:
  
    - Personalized playlists like “Discover Weekly” or “Daily Mix”
    - The AI DJ that speaks to you
    - Smart search suggestions

This matters a lot for Spotify because, by reacting to events in real time and combining that with rich metadata, Spotify can: Understand what you enjoy at a deeper level, Adapt quickly to your changing tastes, Make the platform feel more “you” every time you use it

### The Magic Behind Spotify’s Music Recommendations

Spotify doesn’t just recommend songs randomly, it deeply analyzes every track and your listening habits to give you music you'll love\.

#### Step 1: Understanding the Sound

Spotify’s audio analysis system breaks down songs into **12 key metrics** like tempo, energy, danceability, etc\. They also use **Natural Language Processing \(NLP\)** to understand:

- Lyrics
- Playlist titles
- Even web content about the song or artist

#### Step 2: Building Your Taste Profile

All your interactions like what you listen to, how often, and which artists you love are collected using **Kafka** and stored in **Cassandra**\. This data is used to build a **“Taste Profile”** that’s unique to you\.

#### Step 3: Making Smart Recommendations

Spotify runs machine learning models on this data to power different features:

- **Discover Weekly**: Suggests new releases based on your taste
- **Daily Mixes**: Groups your favorite songs by genre
- **AI DJ**: Creates a never\-ending playlist based on your habits, moods, and more

This entire system, powered by audio analysis, metadata, and machine learning is what Spotify calls *“the magic behind the music\.”* It keeps the music fresh, relevant, and personalized every time you hit play\.

### Behind the Scenes of Spotify Wrapped

**Spotify Wrapped** is more than just a fun feature, it’s one of Spotify’s most complex engineering projects\.

Every year, Spotify collects your **listening data**, favorite songs, artists, moods, etc\. and turns it into personalized Wrapped stories\. This is powered by massive data processing jobs written in **Scio** and run on **Apache Beam**\.

To optimize this massive load, especially in 2020, they introduced **Sort Merge Bucket \(SMB\)**, a technique that improved performance by:

- Efficient **sharding** and **partitioning**
- Better **parallel processing**
- Lower **costs** and **faster delivery**

Wrapped Isn’t Just Data, It’s a Visual Experience\! Engineers also help bring Wrapped to life visually\.

In 2022, they introduced **Listening Personalities** using your data to assign one of 16 unique music personas\.

In 2023, developers adopted a new animation system called **Lottie**\. This helped create smoother and more personalized visuals that worked well across devices — web, mobile, and more\. They even combined **standard animations** with **user\-specific ones** to create a seamless, tailored experience\.

By optimizing visuals and saving costs, Spotify could invest more in marketing getting Wrapped in front of more users around the world\.

### Key Takeaways:

Here’s what makes Spotify Engineering stand out:

- Seamless database migrations without users noticing
- Scaling effortlessly with Kubernetes and the cloud
- Real\-time data handling with Kafka
- Personalized recommendations powered by machine learning
- Beautiful user experiences thanks to smart frontend and design tools

The result? Your favorite song loads instantly, and your playlists feel like they were made just for you\.

By now, you must have had a clear idea of, **How Spotify Powers Music Streaming for Millions? **In a nutshell, Spotify powers music streaming for millions using microservices, Kafka for real\-time events, and Cassandra for massive data handling at scale\. Machine learning and metadata fuel personalized playlists, while scalable cloud infra ensures fast, resilient, and seamless user experiences\.
