# How Zomato Handles 100 Million Daily Search Queries
*Zomato fixed search scale issues by moving from Field Cache to DocValues and using nested docs, cutting costs, OOM errors & boosting speed.*
By [Rohit Lakhotia](https://scaleengineer.com/authors/rohit-lakhotia)
Published: 2025-09-29
Canonical: https://scaleengineer.com/blog/how-zomato-handles-100-million-daily-search-queries
---
Whether you’re craving a hot pizza or a delicious biryani, finding exactly what you want on Zomato happens in just a few clicks\. But have you ever wondered **what powers the search behind the scenes** on Zomato’s app and website?

Every search query that a customer makes needs to return **accurate and relevant results**, whether it’s for restaurants, cuisines, or reviews\. Zomato’s search system handles **over 100 million queries every day**, and the backbone of this system is **Solr\-Lucene**\.

Solr is an **open\-source, enterprise\-level search platform** built on **Apache Lucene**, a widely trusted library for search applications\. It’s reliable, scalable, and has a strong developer community\. For context, Elasticsearch also uses Lucene under the hood\.

## The Challenge: Scaling Solr

Initially, the search system worked smoothly\. But as Zomato’s traffic grew, **scaling became a problem**\. A single Solr node couldn’t handle the high query load and often ran into **Out of Memory \(OOM\) errors**\. This meant that adding more traffic required a larger cluster, which **increased costs significantly**\.

So why were these OOM errors happening?

### Understanding OOM in Solr

Solr runs as a **Java process**\. Java has a **Garbage Collector \(GC\)** that cleans up unused memory, but in this case, the GC struggled to free memory in the **Old Gen Heap space**, eventually causing the process to crash\.

Zomato runs Solr clusters in a **Master\-Slave setup**:

- **Master node:** Handles indexing \(writing data\)\.
- **Slave nodes:** Handle queries \(reading data\) and periodically sync with the Master\.

Even after tweaking JVM and memory settings, the slaves still went OOM\. The problem was traced to **Solr’s Field Cache**, which grows indefinitely during queries that involve sorting, grouping, or faceting\.

### What is Field Cache?

The Field Cache is a core Solr component that **speeds up sorting, grouping, and faceting** by storing field values of all filtered documents in memory\. Without it, Solr would need to scan every document each time a query is run, which is slow and resource\-intensive\.

Here’s how it works:

**Document\-level storage**:

```
{
  'DocX': {'A':1, 'B':2, 'C':3},
  'DocY': {'A':2, 'B':3, 'C':4},
  'DocZ': {'A':4, 'B':3, 'C':2}
}
```

**Field Cache \(column\-oriented\)**:

```
{
  'A': {'DocX':1, 'DocY':2, 'DocZ':4},
  'B': {'DocX':2, 'DocY':3, 'DocZ':3}
}
```

The **issue:** Field Cache isn’t configurable and **cannot limit its size**, so it grows unbounded over time, especially with **Dynamic Fields**, which allow an unlimited number of fields\. This can cause OOM issues during aggregation queries\.

For comparison, **Elasticsearch limits total fields to 1000** by default to prevent “mapping explosion,” but Solr doesn’t impose such a limit\.

## Temporary Workaround: Mock Documents

In Zomato’s setup, slave nodes only purge their caches when the master index updates\. To force regular cache purging, they **indexed a mock document periodically**, ensuring that the slave nodes always had an updated index\.

**Trade\-offs:**

- This prevented OOM errors\.
- But it increased **CPU usage** and **query latency**, because rebuilding the Field Cache is resource\-intensive\.
- Syncing too often would further reduce throughput and increase memory pressure\.

Clearly, a **better solution** was needed\.

## The Right Fix: DocValues

The permanent solution was to use **DocValues**\.

**What are DocValues?**

- They store field values in a **column\-oriented format** at **index time** instead of query time\.
- Unlike Field Cache, DocValues are memory\-efficient because they rely on the **operating system’s memory management \(MMapDirectory\)** rather than Java’s heap\.
- This eliminates OOM issues while **boosting query speed and throughput**\.

Implementation is simple:

```
<field name="field_name" type="string" indexed="false" stored="false" docValues="true" />
```

**Benefits:**

- Reduces the overhead of un\-inverting data at query time\.
- Handles sorting, grouping, and faceting efficiently\.
- Makes the system more **scalable and resilient** to growing traffic\.
- **10x increase in throughput** per slave node\.
- Reduced cluster size, saving roughly **₹30L/month \(~80% cost reduction\)**\.

### Challenges with DocValues

- **Dynamic Fields:** Using DocValues with sparse Dynamic Fields slowed down indexing because every document had to be checked for each field\.
- **Lucene fix:** This was resolved in **Lucene 7\.0\.0**, which iterated only through non\-empty fields\.
- Zomato also forked Solr\-Lucene for some fixes to suit their specific use cases\.

After these improvements, clusters upgraded to **Solr v7\.6\.0 and later v8\.7\.0** showed:

- Faster indexing
- Lower latency
- Higher cache hits
- Greater resiliency
- Lower operational costs

## Key Takeaways

1. **Every search use case is unique:** Experiment with JVM, cache, and schema settings to optimize performance\.
2. **Use DocValues for sorting/faceting:** It reduces memory usage and avoids OOM errors\.
3. **Avoid frequent syncing:** Solr performs best with relatively static indices\.
4. **Dynamic Fields require care:** Mapping explosion can still occur; it must be managed thoughtfully\.

Till now, we discussed how **dynamic fields** in their Solr schema caused performance bottlenecks and high costs as Zomato scaled\. While fixes like **upgrading Solr versions** and using **docValues** helped improve throughput and memory efficiency, the reliance on dynamic fields still limited their ability to expand the schema for new use cases\.

Now let’s understand how they **restructured their schema** and **rewrote their queries** to make the system more resilient, scalable, and cost\-efficient without sacrificing performance\.

## Dynamic Fields: Pros and Cons

Dynamic fields let you **add custom fields to documents without changing the schema**\. For example, if we want to store popularity scores for a restaurant across multiple delivery areas, we could use fields like popularity\_area1, popularity\_area2, and so on:

```
{
    'id': 'res1',
    'entity': 'restaurant',
    'name': 'Dynamic Pizza Store',
    'location': '89.998900, 28.8389300',
    'popularity_area1': 0.93,
    'popularity_area2': 0.25,
    'popularity_areaN': 0.71
}
```

**Problem:** On Zomato, a single restaurant can deliver to many areas\. Across millions of restaurants, the **number of unique fields explodes**, leading to:

- Out\-of\-memory errors
- Slower indexing
- Difficulty recovering from failures

Even small experiments, like testing a new popularity metric \(popularity\_v2\_\*\), add hundreds of thousands of fields, worsening the problem\.

## Granular Data and Use Cases

Zomato operates in a **hyperlocal environment**, meaning user relevance often depends on **fine\-grained, location\-specific data**\.

For example:

- A dish like **Sushi** may be more popular in Golf Course Gurugram than in Old Gurugram\.
- Popularity or demand for a dish varies **even within a city**\.

This requires storing **localized popularity data** for each entity\. While dynamic fields seem like a logical choice, **high\-cardinality dynamic fields are not scalable**\.

### Alternative Model: Maps

One approach is to store popularity as a **map** instead of multiple dynamic fields:

```
{
    'id': 'res1',
    'entity': 'restaurant',
    'name': 'Dynamic Pizza Store',
    'location': '89.998900, 28.8389300',
    'popularity': {
        'area1': 0.93,
        'area2': 0.25,
        'areaN': 0.71
    }
}
```

**Why not this?**

- Solr cannot filter, sort, or aggregate efficiently on map fields\.
- Queries like “restaurants delivering in areaX sorted by popularity” would be impossible\.

## The Solution: Nested Documents

Solr supports **nested documents**, which allow a **parent document** to have **child documents**, while preserving efficient indexing and querying\.

For the restaurant example:

```
{
    'id': 'res1',
    'entity': 'restaurant',
    'name': 'Dynamic Pizza Store',
    'location': '89.998900, 28.8389300',
    'popularity': [
        {'id': 'res1_area1', 'entity': 'popularity', 'value': 0.93},
        {'id': 'res1_area2', 'entity': 'popularity', 'value': 0.25},
        {'id': 'res1_areaN', 'entity': 'popularity', 'value': 0.71}
    ]
}
```

**Key Solr fields for nested documents**:

- \_root\_ : ID of the root document
- \_nest\_path\_ : path of the document in the hierarchy
- \_nest\_parent\_ : ID of the immediate parent

Example schema:

```
<field name="_root_" stored="false" type="string" indexed="true"/>
<fieldType name="_nest_path_" class="solr.NestPathField"/>
<field name="_nest_path_" type="_nest_path_"/>
<field name="_nest_parent_" stored="true" indexed="true" type="string"/>
```

By defining these fields, Solr can maintain **hierarchical relationships**, allowing an entire restaurant \(with dishes and popularity scores\) to be **indexed, updated, or deleted as a single unit**\.

## Benefits of Nested Documents

- Eliminates **mapping explosion** caused by dynamic fields
- Improves **indexing speed** despite a slight increase in index size
- Allows efficient **hierarchical queries**
- Reduces unique field counts, improving performance and memory usage

### Querying Nested Documents: Block Join Queries \(BJQ\)

To leverage nested documents, we use **Block Join Queries**:

1. **Child Query:** Find child documents for specific parent attributes

```
q={!child of=<blockMask>}<someParents>
```

2. **Parent Query:** Find parent documents for specific child attributes

```
q={!parent which=<blockMask>}<someChildren>
```

**Example:**

- Parent: Restaurant
- Child: Dish or popularity in an area

BJQ allows queries like:

- “Find restaurants serving a specific dish”
- “Find dishes in restaurants popular in a specific area”

This makes searching **flexible and powerful** for Zomato’s hyperlocal data\.

## Impact of Schema Change

- Dynamic fields replaced by **nested child documents**
- Query paradigm switched to **BJQ**
- **Index size slightly increased**, but indexing speed **improved**
- Eliminated memory and performance bottlenecks caused by field explosion

While nested documents solve the dynamic field problem, Zomato’s growing number of restaurants and delivery locations means **child documents are increasing**\. Maintaining the full index on a single machine in a Master\-Slave setup is **no longer viable**\.

![](https://media.beehiiv.com/cdn-cgi/image/fit=scale-down,format=auto,onerror=redirect,quality=80/uploads/asset/file/aaafaf5d-d460-47c7-8ded-d704fcc37e9f/Mermaid_Chart_-_Create_complex__visual_diagrams_with_text._A_smarter_way_of_creating_diagrams.-2025-09-14-135841.png?t=1757858363)

Through **DocValues, nested documents, and query redesign**, Zomato’s search system can now:

- Handle millions of queries daily
- Deliver **hyperlocal, relevant results**
- Reduce memory and cost issues
- Scale efficiently for future growth

## Let’s go beyond Master\-Slave, Shift to SolrCloud

But schema changes alone weren’t enough\. As Zomato grew, the Master\-Slave setup created new headaches:

- Slaves with **120GB RAM** still hit latency due to GC pauses
- New replicas took **10\+ minutes to autoscale**
- Failovers were manual and unreliable

The shift to **SolrCloud** solved this:

- Data split into **shards** across nodes
- Queries distributed across shards → better parallelism
- Smaller nodes \(75% less resources per machine\)
- Faster autoscaling and recovery

## Smart Sharding for Hyperlocal Search

Zomato introduced **custom data sharding** aligned with their hyperlocal business model\. Queries were routed directly to the right shard, turning off distributed search for 95% of cases\.

Result:

- 75% reduction in node load
- Lower latencies
- Cheaper infra bills

## Downtime Scare & Fixing Reliability

One shard failure during peak hours exposed reliability gaps:

- Clusters ran on **Spot Instances**, which can vanish anytime
- All replicas were **NRT**, which overloaded the new leader during recovery
- End result → a shard went fully down, causing downtime

**The Fix**

- Adopted a mix of **TLOG \+ Pull replicas**
- Ensured replicas could survive leader loss and sync incrementally
- Slight trade\-off in real\-time freshness, but **much higher availability**

## Faster Scaling with Snapshot Bootstraps

Another bottleneck was replica spin\-ups\. New nodes downloading entire indexes from leaders overloaded the cluster\.

Zomato’s solution:

- Take **index snapshots** and store them in S3
- New replicas boot from snapshots instead of leaders
- Then fetch only incremental changes

This improved autoscaling speed and added **point\-in\-time recovery**\.

## Key Takeaways

Through continuous iteration, Zomato evolved its search system step by step:

- **Field Cache → DocValues**: fixed OOMs, boosted throughput
- **Dynamic Fields → Nested Documents**: avoided mapping explosion
- **Master\-Slave → SolrCloud**: improved scalability & autoscaling
- **NRT only → TLOG \+ Pull replicas**: ensured high availability
- **Leader syncs → Snapshot bootstraps**: accelerated scaling

Official blog from Zomato: [Explained: How Zomato Handles 100 Million Daily Search Queries\! \(Part One\)](https://blog.zomato.com/explained-how-zomato-handles-100-million-daily-search-queries-p1), [Explained: How Zomato Handles 100 Million Daily Search Queries\! \(Part Two\)](https://blog.zomato.com/explained-how-we-handle-100million-daily-search-queries-pt2), [Explained: How Zomato Handles 100 Million Daily Search Queries\! \(Part Three\)](https://blog.zomato.com/explained-how-zomato-handles-100-million-daily-search-queries-part-three)

By now, you must have had a clear idea of,** How Zomato Handles 100 Million Daily Search Queries? **In a nutshell, Zomato handles **100M\+ daily searches** by redesigning its Solr architecture around **DocValues, nested documents, SolrCloud, and smarter replication strategies**\. This shift boosted **throughput, scalability, and hyperlocal relevance**, ensuring fast and reliable results at massive scale\.

**Congratulations\! You've just advanced another step in your tech journey\. Keep progressing\!**
