# Tag: System Design

Posts tagged System Design.

- [What is DNS, and how does it work?](https://scaleengineer.com/blog/what-is-dns-and-how-does-it-work)
  - Why is DNS so important that Facebook, Instagram and Whatsapp had a outage due to that?
- [What is CI/CD and why is it even needed?](https://scaleengineer.com/blog/what-is-ci-cd-and-why-is-it-even-needed)
  - It's the automation that makes developer life simpler and efficient!
- [What are microservices?](https://scaleengineer.com/blog/what-are-microservices)
  - Netflix uses microservices but Google doesn't. But what exactly is that?
- [What is an ORM?](https://scaleengineer.com/blog/what-is-an-orm)
  - Ever thought of skipping database languages? ORM is for you!
- [EP 22: What is SPF Record? Why is it used?](https://scaleengineer.com/blog/what-is-spf-record-why-is-it-used)
  - An SPF record controls which servers can send emails for a domain, preventing email fraud.
- [What is ⁠PostgreSQL?](https://scaleengineer.com/blog/what-is-postgresql)
  - PostgreSQL is a robust, open-source object-relational database system known for advanced features, scalability, and support for complex queries.
- [EP 26: What is Docker?](https://scaleengineer.com/blog/what-is-docker)
  - Docker simplifies application deployment by packaging software into standardized containers.
- [EP 27: What is Kubernetes?](https://scaleengineer.com/blog/what-is-kubernetes)
  - Kubernetes orchestrates and automates the deployment, scaling, and management of containerized applications.
- [EP 28: What is Kafka?](https://scaleengineer.com/blog/what-is-kafka)
  - Kafka is a distributed streaming platform for real-time data with low latency and high throughput.
- [EP 29: What is Cassandra?](https://scaleengineer.com/blog/what-is-cassandra)
  - Cassandra is a scalable, distributed NoSQL database for handling large data with high availability.
- [EP 45: How Slack Maintains Reliability and Uptime](https://scaleengineer.com/blog/how-slack-maintains-reliability-and-uptime)
  - Slack maintains reliability and uptime through automated incident detection, real-time collaboration, proactive monitoring, and a resilient microservices architecture.
- [EP 43: How Amazon Personalizes Product Recommendations](https://scaleengineer.com/blog/how-amazon-personalizes-product-recommendations)
  - Amazon personalizes product recommendations using machine learning, collaborative filtering, and user interaction data to tailor suggestions based on individual preferences
- [EP 46: How Uber Manages Real-Time Analytics with Apache Flink](https://scaleengineer.com/blog/how-uber-manages-real-time-analytics-with-apache-flink)
  - Uber Eats uses real-time data processing with Apache Kafka, Flink, and Pinot to manage order updates, optimize delivery logistics, and provide quick analytics for efficient and accurate food delivery.
- [EP 39: How Twitter Manages High Availability with Kubernetes](https://scaleengineer.com/blog/how-twitter-manages-high-availability-with-kubernetes)
  - Twitter achieves high availability with Kubernetes through multi-node deployments, load balancing, and data center redundancy.
- [EP 37: What is OAuth?](https://scaleengineer.com/blog/what-is-oauth)
  - OAuth is an open standard protocol that allows users to grant apps access to their data without sharing their passwords.
- [EP 41: How Facebook Handles Billions of Messages Daily](https://scaleengineer.com/blog/how-facebook-handles-billions-of-messages-daily)
  - Facebook manages billions of daily messages using scalable servers, distributed systems and advanced algorithms for efficient processing and real-time delivery.
- [How Zoom Ensures Low Latency Video Calls](https://scaleengineer.com/blog/how-zoom-ensures-low-latency-video-calls)
  - Zoom ensures low latency by using distributed data centers, optimized video encoding, and adaptive bitrate streaming to maintain real-time communication quality.
- [EP 59: How Reddit designed their Metadata Store to serve 100k req/sec?](https://scaleengineer.com/blog/how-reddit-designed-their-metadata-store-to-serve-100k-req-sec)
  - Reddit built a high-performance metadata store using Aurora Postgres, range-based partitions, PgBouncer, and JSONB fields, handling 100k req/sec.
- [EP 42: How Pinterest Scales Their Image Search with Elasticsearch](https://scaleengineer.com/blog/how-pinterest-scales-their-image-search-with-elasticsearch)
  - Pinterest scales its image search by using Elasticsearch for fast indexing, real-time search, and advanced machine learning features.
- [EP 55: How did Magic Pocket help Dropbox save millions? ](https://scaleengineer.com/blog/how-did-magic-pocket-help-dropbox-save-millions)
  - Dropbox scaled its storage with its custom-built system- Magic Pocket, and utilized high-density SMR drives, increasing its gross revenue by 75%.
- [EP 57: How Airbnb Processes a Million User Events Every Second?](https://scaleengineer.com/blog/how-airbnb-s-user-signals-platform-process-a-million-user-events-every-second)
  - How Airbnb made 1.9 Billion in 6 months and how Airbnb’s User Signals Platform uses Apache Flink & Lambda Architecture to process millions of events per second for real-time personalization.
- [EP 58: How Facebook built its Video Delivery System?](https://scaleengineer.com/blog/how-facebook-built-its-video-delivery-system)
  - Facebook unified Reels, Watch, and Live, optimizing ranking, servers, and mobile to deliver personalized, efficient, and fresh video experiences.
- [EP 50: How Google search works?](https://scaleengineer.com/blog/how-google-search-works)
  - Google Search works by using crawlers to scan and index web pages, then processes your queries to rank and display relevant results in seconds.
- [EP 48: How Tinder Streams to 75 Million Users with HTTP Live Streaming](https://scaleengineer.com/blog/how-tinder-streams-to-75-million-users-with-http-live-streaming)
  - Tinder used HTTP Live Streaming (HLS) & AWS CloudFront to deliver Swipe Night videos efficiently, ensuring seamless, adaptive playback.
- [EP 49: How Stripe Handles Global Payments Technology](https://scaleengineer.com/blog/how-stripe-handles-global-payments-technology)
  - Stripe utilizes a tech stack of Ruby and JavaScript to enable secure, compliant global payments and currency conversion.
- [EP 52: How GitHub manages continuous integration and deployment](https://scaleengineer.com/blog/how-github-manages-continuous-integration-and-deployment)
  - GitHub manages CI/CD by automating testing, building, and deploying code changes, allowing developers to release updates faster and with confidence.
- [EP 51: How Instagram handled user growth and scale?](https://scaleengineer.com/blog/how-instagram-handled-user-growth-and-scale)
  - Instagram achieved rapid user growth by maintaining a simple and efficient tech stack, utilizing AWS, Django, and Postgres also effectively managing traffic with load balancing, caching, and data sharding to handle the increasing demand.
- [EP 54: How Dropbox scaled its storage infrastructure? ](https://scaleengineer.com/blog/how-dropbox-scaled-its-storage-infrastructure)
  - Dropbox scaled its storage infrastructure with a custom-built system called Magic Pocket, utilizing high-density SMR drives and advanced data replication for durability and scalability.
- [EP 56: How LinkedIn Scaled to 1 billion Users?](https://scaleengineer.com/blog/how-linkedin-scaled-to-1-billion-users)
  - By shifting to microservices from monoliths, using tools like Hadoop, Kafka, Rest.li, LinkedIn scaled to a billion of users globally.
- [EP 68: How Stripe uses Similarity Clustering to detect fraud](https://scaleengineer.com/blog/how-stripe-uses-similarity-clustering-to-detect-fraud)
  - Stripe uses similarity clustering with XGBoost to detect fraud, linking accounts by shared traits to block fraud rings in real-time and reduce false positives.
- [EP 70: How Wayfair built their Ad Bidding System?](https://scaleengineer.com/blog/how-wayfair-built-their-ad-bidding-system)
  - Wayfair built a smart Ad Bidding System using automation, ML, and real-time data to optimize bids, maximize ROI, and scale efficiently.
- [EP 69:  How Airbnb Rebuilt its Payment System to achieve 150x performance gains?](https://scaleengineer.com/blog/how-airbnb-rebuilt-its-payment-system-to-achieve-150x-performance-gains)
  - Airbnb rebuilt its payments system with SOA, a unified read layer & denormalization, boosting scalability, reliability & 150x faster transactions.
- [EP 63: How Quora Optimized their Databases?](https://scaleengineer.com/blog/how-quora-optimized-their-databases)
  - Quora optimized databases with caching, MyRocks for storage efficiency, and MySQL sharding to boost performance, cut costs, and handle scale.
- [EP 61: How Stripe achieved 99.999% uptime with DocDB (Document Database)](https://scaleengineer.com/blog/how-stripe-achieved-99-999-uptime-with-docdb-document-database)
  - Stripe achieved 99.999% uptime by building DocDB, a custom solution on MongoDB, enabling efficient data migration, scaling, and high availability.
- [EP 67: How BBC uses Serverless to handle Millions of visitors](https://scaleengineer.com/blog/how-bbc-uses-serverless-to-handle-millions-of-visitors)
  - BBC uses AWS Lambda to scale instantly, optimize caching, and reduce cold starts, ensuring fast, cost-efficient performance for millions of visitors.
- [EP 62: How Robinhood prevents Fraud using Graph Algorithms ](https://scaleengineer.com/blog/how-robinhood-prevents-fraud-using-graph-algorithms)
  - Robinhood prevents fraud using graph algorithms to analyze user connections, detect patterns, and enable real-time, smarter fraud detection.
- [EP 65: How Quora Improved its Search System with Qdrant?](https://scaleengineer.com/blog/how-quora-improved-its-search-system-with-qdrant)
  - Quora moved to Qdrant for faster, scalable embedding search, improving recommendations with real-time updates, bulk loads, and optimized storage.
- [EP 66: How Meta distributes Exabytes of Data across the World so fast?](https://scaleengineer.com/blog/how-meta-distributes-exabytes-of-data-across-the-world-so-fast)
  - Meta uses Owl, a hybrid system that mixes peer-to-peer caching with smart tracking, making data move faster, smoother, and at scale.
- [How Stripe Scales its APIs using Rate Limiters](https://scaleengineer.com/blog/how-stripe-scales-its-apis-using-rate-limiters)
  - Stripe uses token buckets, concurrency limits & load shedders to scale APIs, prevent abuse & keep critical traffic flowing reliably.
- [EP 75: How Netflix built a Distributed Counter for Billions of User Interactions](https://scaleengineer.com/blog/how-netflix-built-a-distributed-counter-for-billions-of-user-interactions)
  - Netflix uses a smart Distributed Counter system to track billions of user actions daily with speed, accuracy, and massive scale.
- [EP 76: How Mixpanel Fixed Their Load Balancing Problem using Power of 2 Choices](https://scaleengineer.com/blog/how-mixpanel-fixed-their-load-balancing-problem-using-power-of-2-choices)
  - Mixpanel fixed Compacter’s load imbalance using Power-of-2-Choices, boosting efficiency and cutting costs by 70% with minimal changes!
- [EP 81: How Pinterest Built “Holiday Finds” to make Gift Shopping easier?](https://scaleengineer.com/blog/how-pinterest-built-holiday-finds-to-make-gift-shopping-easier)
  - Pinterest Holiday Finds uses smart recommendations, auto wishlists and a fresh UI to make holiday gifting easy!
- [How does UPI work?](https://scaleengineer.com/blog/how-upi-works)
  - UPI enables instant bank-to-bank transfers using just a UPI ID or mobile number, no bank details needed, just your app and secure PIN.
- [EP 77: How GitHub made Push Processing faster and more Reliable](https://scaleengineer.com/blog/how-github-made-push-processing-faster-and-more-reliable)
  - GitHub sped up and stabilized push processing by splitting one big job into parallel Kafka-triggered tasks with better retries and monitoring.
- [EP 79: How Grab enabled near Real-Time analytics on their Data Lake](https://scaleengineer.com/blog/how-grab-enabled-near-real-time-analytics-on-their-data-lake)
  - Grab used Apache Hudi with Flink and Spark to enable near real-time analytics, ensuring fast ingestion and low-latency queries on their data lake.
- [EP 80: How Pinterest improved ABR Video Performance?](https://scaleengineer.com/blog/how-pinterest-improved-abr-video-performance)
  - Pinterest sped up video playback by embedding manifests in API responses and using Memcache to reduce startup latency.
- [How Discord’s "Go Live" streaming works](https://scaleengineer.com/blog/how-discord-s-go-live-streaming-works)
  - Discord’s “Go Live” streams in real-time by capturing, encoding, transmitting, and decoding adapting quality to your network and device.
- [EP 86: How Facebook Scales Live Streaming for Millions of Viewers at Once?](https://scaleengineer.com/blog/how-facebook-scales-live-streaming-for-millions-of-viewers-at-once)
  - Facebook scaled Live streaming for millions by building robust ingestion, delivery, and ISP optimizations, powering events like the UEFA Final.
- [EP 88: How Pinterest Evolved its Architecture to Serve 500 Million Users](https://scaleengineer.com/blog/how-pinterest-evolved-its-architecture-to-serve-500-million-users)
  - Pinterest began as a simple side project and scaled by simplifying tech, embracing microservices, and building strong pipelines and monitoring.
- [EP 84: How Pinterest Built Text-to-SQL to make Data analysis easier](https://scaleengineer.com/blog/how-pinterest-built-text-to-sql-to-make-data-analysis-easier)
  - Pinterest built a Text-to-SQL tool using LLMs and RAG to help analysts convert questions into SQL and find the right data faster and easier.
- [How Uber Eats Scaled Search to Handle Billions of Daily Queries](https://scaleengineer.com/blog/how-uber-eats-scaled-search-to-handle-billions-of-daily-queries)
  - Uber Eats scaled search by revamping indexing, geo-sharding & ranking, supporting billions of queries daily without compromising latency.
- [EP 87: How Uber Handles 40 Million+ Reads Per Second Using an Integrated Cache](https://scaleengineer.com/blog/how-uber-handles-40-million-reads-per-second-using-an-integrated-cache)
  - Uber serves 40M+ reads/sec by pairing Docstore with a smart Redis cache, using CDC for near-instant updates and clever sharding for scale.
- [EP 83: How Pinterest Rebuilt its $3B+ Ads System without any Downtime](https://scaleengineer.com/blog/how-pinterest-rebuilt-its-3b-ads-system-without-any-downtime)
  - Pinterest rebuilt its \$3B+ ad system with a graph-based design for better scale, safety & dev speed, launched with zero downtime and big cost wins.
- [How Spotify Powers Music Streaming for Millions](https://scaleengineer.com/blog/how-spotify-powers-music-streaming-for-millions)
  - Spotify uses Kafka, microservices, and ML to deliver real-time, personalized music to millions, powered by a fast, scalable cloud backend.
- [How Meta Powers its Cloud Gaming Infrastructure at Scale](https://scaleengineer.com/blog/how-meta-powers-its-cloud-gaming-infrastructure-at-scale)
  - Meta streams games from cloud GPUs to your device with ultra-low latency, using real-time encoding, smart networking, and fast decoding.
- [Understanding API Gateway in Microservices: Key Benefits & Use Cases](https://scaleengineer.com/blog/understanding-api-gateway-in-microservices-key-benefits-use-cases)
  - Learn how an API gateway in microservices optimizes architecture. Explore core functions, patterns, and best practices to enhance your system.
- [Authentication & Access Control](https://scaleengineer.com/blog/authentication-access-control)
  - You sign in to your bank account and can only view your balance. The bank manager logs in and can approve loans. Same system, different powers but how does the app decide?
- [Data Management in Applications](https://scaleengineer.com/blog/data-management-in-applications)
  - Whether you’re building a simple note-taking app, a social media platform, or a large-scale e-commerce system, your application’s success depends on how well...
- [Performance and Scalability in Web Applications](https://scaleengineer.com/blog/performance-and-scalability-in-web-applications)
  - Ever wondered why some apps stay smooth at 100 users but crash at 10k? That is where performance meets scalability.
- [Vertical vs Horizontal Scaling](https://scaleengineer.com/blog/vertical-vs-horizontal-scaling)
  - Is it better to make one server stronger or add more servers?
- [System Design Tutorial](https://scaleengineer.com/blog/system-design-fundamentals)
  - When applications grow beyond a handful of users, writing code alone isn’t enough. To scale, stay reliable, and support complex features, software needs strong...
- [How Salesforce Reinvented Task Execution for the Cloud Era](https://scaleengineer.com/blog/how-salesforce-reinvented-task-execution-for-the-cloud-era)
  - Salesforce built a cloud-native task execution system in Hyperforce, replacing SSH with secure, scalable, multi-cloud automation using recipes & workers.
- [How Salesforce migrated 200,000 Machines from CentOS 7 to RHEL 9](https://scaleengineer.com/blog/how-salesforce-migrated-200-000-machines-from-centos-7-to-rhel-9)
  - Using automation for zero downtime, stronger security & faster parallel upgrades, Salesforce successfully migrated 200,000 machines from CentOS 7 to RHEL 9
- [How Razorpay Capital Detects Duplicate or Fraudulent Merchants](https://scaleengineer.com/blog/how-razorpay-capital-detects-duplicate-or-fraudulent-merchants)
  - Razorpay scaled payments to billions of transactions by re-engineering its core systems, ensuring speed, security & reliability at scale.
- [What Is an Application Server? Role & Importance](https://scaleengineer.com/blog/what-is-an-application-server-88a2)
  - Ever wondered what happens behind the curtain when you log into an app, book a flight, or add something to your online shopping cart? That seamless, interactive experience is powered by an unseen engine...
- [How Salesforce Migrated 760+ Kafka Nodes Handling 1M Messages per Second with Zero Downtime](https://scaleengineer.com/blog/how-salesforce-migrated-760-kafka-nodes-handling-1m-messages-per-second-with-zero-downtime)
  - Salesforce upgraded 760+ Kafka nodes handling 1M+ msg/sec with zero downtime, scaling Marketing Cloud seamlessly for the future.
- [How Hyperforce Edge Networking Scaled to 20 Million Domains With Less Than 30GB of RAM](https://scaleengineer.com/blog/how-hyperforce-edge-networking-scaled-to-20-million-domains-with-less-than-30gb-of-ram)
  - Scaled from 3M→20M+ domains, Salesforce Hyperforce Edge cut memory <30GB with new storage design, boosting speed, reliability & security.
- [Sharding vs Partitioning: What's the Difference?](https://scaleengineer.com/blog/difference-between-data-sharding-and-partitioning)
  - Partitioning splits data within one database for faster retrieval, while sharding spreads data across multiple databases to handle scale and traffic.
- [How LinkedIn Built a Faster, Safer, and Smarter HDFS Ecosystem](https://scaleengineer.com/blog/how-linkedin-built-a-faster-safer-and-smarter-hdfs-ecosystem)
  - LinkedIn scaled HDFS with HA, Observer nodes, encryption & Wormhole, boosting speed, reliability & secure data access for massive growth.
- [Write-Through, Write-Back & Write-Around in Cache: A Practical Guide](https://scaleengineer.com/blog/write-through-write-back-and-write-around-in-cache)
  - Your app writes data every second but how it writes can change everything. Write-Through, Write-Back & Write-Around hide big trade-offs.
- [How Swiggy Improved Video Performance with Smart Caching](https://scaleengineer.com/blog/how-swiggy-improved-video-performance-with-smart-caching)
  - Swiggy boosted video cache hits & cut costs by clustering widths with K-means, reducing redundant processing while keeping playback seamless.
- [Latency vs Throughput: A Guide for System Performance](https://scaleengineer.com/blog/latency-vs-throughput)
  - When you hear engineers talk about latency vs throughput, they are discussing two sides of the same coin: speed versus capacity.
- [Edge Computing vs Fog Computing: Making the Right Choice](https://scaleengineer.com/blog/edge-computing-vs-fog-computing)
  - When comparing edge computing vs. fog computing, the main difference comes down to a simple question: where does the data processing happen?
- [How Zomato Handles 100 Million Daily Search Queries](https://scaleengineer.com/blog/how-zomato-handles-100-million-daily-search-queries)
  - Zomato fixed search scale issues by moving from Field Cache to DocValues and using nested docs, cutting costs, OOM errors & boosting speed.
- [How Razorpay prepared for Chrome’s Third-Party Cookie Deprecation](https://scaleengineer.com/blog/how-razorpay-prepared-for-chrome-s-third-party-cookie-deprecation)
  - Razorpay uses partitioned cookies (CHIPS) to tackle Chrome’s 3P cookie phaseout, cutting drop-offs while ensuring a smooth, reliable checkout.
- [How Swiggy Scaled and Maintained Postgres](https://scaleengineer.com/blog/how-swiggy-scaled-and-maintained-postgres)
  - Swiggy scaled Postgres by cleaning unused indexes, controlling auto-vacuum, and using pg_repack for online maintenance and better performance.
- [How Shopify Built Super-Fast Search at C++ Speed](https://scaleengineer.com/blog/how-shopify-built-super-fast-search-at-c-speed)
  - Shopify built RankFlow to run ML-powered search at C++ speed, letting data scientists iterate fast without sacrificing latency or scale.
- [How Razorpay Uses Terraform to Simplify and Scale Infrastructure Management](https://scaleengineer.com/blog/how-razorpay-uses-terraform-to-simplify-and-scale-infrastructure-management)
  - Razorpay leverages Terraform + Atlantis to automate, secure, and scale infrastructure with GitOps workflows and modular IaC practices.
- [What is Token Bucket Algorithm?](https://scaleengineer.com/blog/what-is-token-bucket-algorithm)
  - Discover how context switching lets operating systems multitask smoothly, switching between processes to keep your system fast and efficient.
- [How Swiggy Cut QA Regression Time by 66% Using Automated Event Testing](https://scaleengineer.com/blog/how-swiggy-cut-qa-regression-time-by-66-using-automated-event-testing)
  - Swiggy built ARD Automator to automate mobile event verification using contracts and validators, cutting QA time by 66% and boosting accuracy.
- [How Airbnb builds Products 10x Faster Using GraphQL and Apollo](https://scaleengineer.com/blog/how-airbnb-builds-products-10x-faster-using-graphql-and-apollo)
  - Airbnb ships faster by using GraphQL and Apollo to power backend-driven UI, automatic types, and tooling that lets engineers focus on building features
- [How LinkedIn Made the “My Network” Tab Faster, Smoother, and More Flexible](https://scaleengineer.com/blog/how-linkedin-made-the-my-network-tab-faster-smoother-and-more-flexible)
  - LinkedIn sped up My Network by unifying APIs, adding pagination, and using a backend-driven render model, cutting latency and improving the overall UX.
- [What Is CQRS?](https://scaleengineer.com/blog/what-is-cqrs)
  - What is CQRS? This guide explains the CQRS pattern with simple analogies and practical examples to help you build scalable and high-performance applications.
- [How Zomato Improved their Android App Startup Time by Over 20% Using Baseline Profiles](https://scaleengineer.com/blog/how-zomato-improved-their-android-app-startup-time-by-over-20-using-baseline-profiles)
  - Zomato cut Android app startup time by 20% using Baseline Profiles, pre-optimizing key code paths for faster launches and a smoother, consistent user experience.
- [What Happens During a Database Migration?](https://scaleengineer.com/blog/what-happens-during-a-database-migration)
  - Discover what happens during a database migration. This practical guide covers planning, execution, validation, and strategies for a smooth transition.
- [How Airbnb Measures the Lifetime Value of a Listing](https://scaleengineer.com/blog/how-airbnb-measures-the-lifetime-value-of-a-listing)
  - Airbnb’s LTV framework shows which listings drive value, supports hosts, and adapts to market changes for smarter, data-driven decisions.
- [How Lyft Rebuilt its Iconic Dashboard Emblem and its entire IoT Platform along with it?](https://scaleengineer.com/blog/how-lyft-rebuilt-its-iconic-dashboard-emblem-and-its-entire-iot-platform-along-with-it)
  - Lyft’s Glow is more than an emblem, it’s a unified IoT platform with secure provisioning, real-time control, device shadowing, and safe OTA updates.
- [How LinkedIn Reduced Latency and Cost by Merging Two Critical Systems](https://scaleengineer.com/blog/how-linkedin-reduced-latency-and-cost-by-merging-two-critical-systems)
  - LinkedIn merged identity midtier and data services, cutting network hops to reduce latency, memory use, and cost while keeping APIs unchanged.
- [How Instagram Improved HDR Video on iOS With Dolby Vision](https://scaleengineer.com/blog/how-instagram-improved-hdr-video-on-ios-with-dolby-vision)
  - Dolby Vision first hurt Reels due to load delays from metadata. Compression fixed it, boosting watch time and enabling rollout on Instagram iOS
- [How LinkedIn Rebuilt its Profile Highlights System](https://scaleengineer.com/blog/how-linkedin-rebuilt-its-profile-highlights-system)
  - LinkedIn rebuilt Profile Highlights into a plug-in platform, enabling faster experiments, independent teams, better performance, and ~50% lower costs.
- [What is Lazy Loading vs Eager Loading?](https://scaleengineer.com/blog/lazy-loading-vs-eager-loading-explained)
  - lazy loading vs eager loading explained with practical examples. Learn when to apply each approach for performance and resource efficiency.
- [What Is MLOps?](https://scaleengineer.com/blog/what-is-mlops-bridging-the-gap-between-devops-and-machine-learning)
  - Learn what is MLOps: bridging the gap between devops and machine learning. Explore its lifecycle, tools, and best practices for scaling AI effectively.
- [How Dropbox Dash Uses a Feature Store for Real-Time AI](https://scaleengineer.com/blog/how-dropbox-dash-uses-a-feature-store-for-real-time-ai)
  - Dropbox Dash uses a hybrid feature store to deliver fast, fresh signals at scale, keeping AI search accurate, low-latency, and reliable at scale.
- [How Dropbox Dash Uses Context Engineering to Build Smarter AI](https://scaleengineer.com/blog/how-dropbox-dash-uses-context-engineering-to-build-smarter-ai)
  - Dropbox Dash evolved into agentic AI by engineering context fewer tools, relevant data, and specialized agents making AI faster, smarter at work.
- [Why Spotify’s Shuffle Never Felt Random (and What They Did About It)](https://scaleengineer.com/blog/why-spotify-s-shuffle-never-felt-random-and-what-they-did-about-it)
  - Spotify kept Shuffle random but made it feel fair by choosing the least repetitive random order, so songs feel fresher without breaking true randomness.
- [What is Publish-Subscribe Pattern? ](https://scaleengineer.com/blog/what-is-publish-subscribe-pattern)
  - What is publish-subscribe pattern? Learn how pub/sub decouples components, with real-world examples and benefits for scalable systems.
- [How Spotify Scaled Content Annotations to Millions (Without Losing Quality)](https://scaleengineer.com/blog/how-spotify-scaled-content-annotations-to-millions-without-losing-quality)
  - Spotify built a scalable annotation platform by combining human experts, smart tools, and strong infrastructure to power high-quality ML training data
- [What is Configuration Drift?](https://scaleengineer.com/blog/what-is-configuration-drift-a-guide-for-devops)
  - What is Configuration Drift? Learn causes, risks, and best practices to detect, prevent, and fix drift with IaC and GitOps.
- [How Slack cut their E2E Build Time by 80%?](https://scaleengineer.com/blog/how-slack-cut-their-e2e-build-time-by-80)
  - Slack cut E2E time 80% by skipping redundant frontend builds and reusing cached assets, saving compute, storage, and hours.
- [How Slack Built Secure Enterprise Search?](https://scaleengineer.com/blog/how-slack-built-secure-enterprise-search)
  - Slack enables secure enterprise search using real-time fetch, RAG, ACL & OAuth, no data storage, always permission-aware & private across tools.
- [What is Backpressure?](https://scaleengineer.com/blog/what-is-backpressure)
  - Learn what is backpressure in distributed systems, why it’s vital for stability, and key strategies to prevent overloads in large-scale systems.
- [How Airbnb Migrated a Petabyte Without Users Noticing](https://scaleengineer.com/blog/how-airbnb-migrated-a-petabyte-without-users-noticing)
  - Airbnb rebuilt Mussel into a cloud-native KV store and migrated 1PB+ data using Apache Kafka with zero downtime.
- [Replication vs Redundancy. What's the Difference?](https://scaleengineer.com/blog/difference-between-replication-and-redundancy)
  - Learn the key differences between replication and redundancy to optimize your data protection strategies. Discover which method suits your needs best.
- [How Slack Automatically Stops Suspicious Activity in Real Time](https://scaleengineer.com/blog/how-slack-automatically-stops-suspicious-activity-in-real-time)
  - Slack’s AER detects suspicious activity and automatically terminates user sessions, shrinking response time from hours to minutes.
- [What are Immutable Data Structures?](https://scaleengineer.com/blog/immutable-data-structures-why-they-matter-in-modern-coding)
  - Explore why immutable data structures: why they matter in modern coding. Discover how they enhance reliability, simplify concurrency, and prevent bugs.
- [API Gateway vs Load Balancer](https://scaleengineer.com/blog/api-gateway-vs-load-balancer)
  - Discover the differences between API gateway vs load balancer and find out which is best for your system's performance and security needs.
- [How GitHub Redesigned CLI Accessibility Without a Rulebook](https://scaleengineer.com/blog/how-github-redesigned-cli-accessibility-without-a-rulebook)
  - GitHub makes CLI accessible by improving prompts, colors, and output, helping screen readers, low-vision users, and making terminals usable for all devs
- [How DoorDash transitioned from Monolith to Microservices](https://scaleengineer.com/blog/how-doordash-transitioned-from-monolith-to-microservices-be29)
  - DoorDash used the strangler fig pattern, scream tests, and multi-tenant architecture to smoothly transition from monolith to microservices.
- [How Netflix Secures Content Delivery using Open Connect CDN?](https://scaleengineer.com/blog/how-netflix-secures-content-delivery)
  - Netflix secures content delivery through its proprietary Open Connect CDN, which caches content on local servers, ensuring low-latency streaming and minimizing network congestion.
- [How Jira moved from JSON to Protobuf saved them 55% cost and 75% CPU?](https://scaleengineer.com/blog/how-jira-saved-55-cost-and-75-cpu-by-moving-from-json-to-protobuf)
  - Jira cut data size by 80%, reduced Memcached CPU by 75%, and saved 55% in costs by switching from JSON to Protobuf, improving speed and efficiency.
- [EP 38: How Spotify Optimized Their Recommendation System](https://scaleengineer.com/blog/how-spotify-optimized-their-recommendation-system)
  - Spotify optimized recommendations by combining collaborative filtering, content-based filtering, and audio analysis to deliver highly personalized music recommendations.
- [EP 82: How Pinterest uses LLMs to make your Search Results more Relevant?](https://scaleengineer.com/blog/how-pinterest-uses-llms-to-make-your-search-results-more-relevant)
  - Pinterest's AI teacher-student system improved search by 19.7%, understanding user intent beyond keywords for better relevance globally
- [How Salesforce migrated 200,000 Machines from CentOS 7 to RHEL 9](https://scaleengineer.com/blog/sending-email-how-salesforce-migrated-200-000-machines-from-centos-7-to-rhel-9)
  - Using automation for zero downtime, stronger security & faster parallel upgrades, Salesforce successfully migrated 200,000 machines from CentOS 7 to RHEL 9
- [How X (Formerly Twitter) Handles Millions of Tweets Every Second](https://scaleengineer.com/blog/how-x-formerly-twitter-handles-millions-of-tweets-every-second)
  - X scaled from Ruby to Java, microservices, real-time data, and AI to handle millions of tweets, searches, and users with speed and reliability.
- [Circuit Breaker vs Retry in Microservices](https://scaleengineer.com/blog/circuit-breaker-vs-retry)
  - When building resilient systems, the debate of circuit breaker vs retry is about choosing the right tool for the right kind of failure. A Retry pattern is...
- [How Amazon Key Unlocks 100 Million Doors a Year](https://scaleengineer.com/blog/how-amazon-key-unlocks-100-million-doors-a-year)
  - Amazon Key lets drivers unlock gates for faster deliveries. From serverless to microservices, it now powers 100M+ secure unlocks yearly.
- [EP 71: How PayPal Solved the Thundering Herd Problem Efficiently](https://scaleengineer.com/blog/how-paypal-solved-the-thundering-herd-problem-efficiently)
  - PayPal’s Braintree fixed the Thundering Herd Problem using Exponential Backoff with Jitter and simplified their architecture for better scaling.
- [How Shopify Made Commerce Data Queryable Without SQL](https://scaleengineer.com/blog/how-shopify-made-commerce-data-queryable-without-sql)
  - ShopifyQL Notebooks lets merchants explore business data without SQL, using commerce-focused models built for clarity, speed, and action.
- [What Is the N+1 Query Problem?](https://scaleengineer.com/blog/what-is-the-n-1-query-problem)
  - What is the n+1 query problem? Learn how it slows apps, why it happens, and practical fixes with code examples to speed up performance.
- [How LinkedIn Cut Build Times from 30 Minutes to 10 Seconds](https://scaleengineer.com/blog/how-linkedin-cut-build-times-from-30-minutes-to-10-seconds)
  - LinkedIn’s RDev lets engineers code in the cloud with pre-built containers, cutting setup from 30 mins to 10 secs while keeping CI consistent.
- [How Lyft Built an In-App Messaging Without Annoying Riders](https://scaleengineer.com/blog/how-lyft-built-an-in-app-messaging-without-annoying-riders)
  - Lyft built in-app messaging by starting with simple banners and scaling into a smart, context-aware system that delivers timely messages without annoying riders.
- [How GitHub Uses CodeQL to Secure Code at Scale](https://scaleengineer.com/blog/how-github-uses-codeql-to-secure-code-at-scale)
  - GitHub uses CodeQL to scan code as data, detect vulnerabilities, and secure thousands of repos automatically at scale.
- [What is Micro Frontend Architecture?](https://scaleengineer.com/blog/micro-frontend-architecture)
  - Micro frontends are extending the concepts of micro services to the frontend world.
- [How Snowflake Reduced Query Time by 20% (Without You Doing Anything)](https://scaleengineer.com/blog/how-snowflake-reduced-query-time-by-20-without-you-doing-anything)
  - Snowflake reduces query time by 20% via continuous engine optimizations, improving real workloads automatically without user changes.
- [EP 53: How TikTok Optimizes Video Streaming](https://scaleengineer.com/blog/how-tiktok-optimizes-video-streaming)
  - TikTok boosts streaming by preloading videos, optimizing buffers, and reusing media players, with on-device upscaling and task distribution for smooth playback on all networks.
- [How LinkedIn Rebuilt Service Discovery to Scale to Millions of Services](https://scaleengineer.com/blog/how-linkedin-rebuilt-service-discovery-to-scale-to-millions-of-services)
  - LinkedIn rebuilt service discovery using Kafka and Observer, enabling scalable, push-based updates with lower latency and higher availability.
- [How LinkedIn Uses Machine Learning to Moderate Content at Scale](https://scaleengineer.com/blog/how-linkedin-uses-machine-learning-to-moderate-content-at-scale)
  - LinkedIn is using ML to prioritize content smarter, not replace humans but helping reviewers act faster, scale better, and keep the platform safe without losing judgment or nuance.
- [How Snowflake Improved Performance by 27% (Without Users Noticing)](https://scaleengineer.com/blog/how-snowflake-improved-performance-by-27-without-users-noticing)
  - Snowflake boosts performance by 27% via backend optimizations in ingestion, planning, and execution thus faster queries and lower cost automatically
- [EP 25: What is DMARC Record? Why is it used?](https://scaleengineer.com/blog/what-is-dmarc-record-why-is-it-used)
  - DMARC prevents email spoofing and phishing by authenticating email senders.
- [How Slack Built Accessibility Checks into Its Testing Pipeline](https://scaleengineer.com/blog/how-slack-built-accessibility-checks-into-its-testing-pipeline)
  - Slack added Axe-based accessibility checks to Playwright tests, balancing automation with reliability, better reports, and easy developer workflows.
- [How Nomad by HashiCorp Reduced Scheduler Load by 90%](https://scaleengineer.com/blog/how-nomad-by-hashicorp-reduced-scheduler-load-by-90)
  - Nomad reduces scheduler load by canceling redundant evaluations, improving system performance and speeding up recovery during failures.
- [What are SOLID Principles?](https://scaleengineer.com/blog/solid-principles-in-software-engineering-explained-with-examples)
  - Learn solid principles in software engineering: explained with examples to write clean, maintainable, and scalable code. A practical guide for developers.
- [How Slack makes its Mobile App Feel Seamless (Even on Bad Internet)](https://scaleengineer.com/blog/how-slack-makes-its-mobile-app-feel-seamless-even-on-bad-internet)
  - Slack optimizes mobile performance using prioritized APIs, caching + versioning, offline sync, and scalable architecture for reliable user experience.
- [When Microservices Get Messy: How API Federation Brings Order](https://scaleengineer.com/blog/when-microservices-get-messy-how-api-federation-brings-order)
  - API Federation combines multiple services into one API using shared models and modular features, simplifying complex microservice architectures.
- [How Uber Standardized Mobile Analytics (Without Slowing Down Teams)](https://scaleengineer.com/blog/how-uber-standardized-mobile-analytics-without-slowing-down-teams)
  - Uber standardized mobile analytics by moving event logic to the platform, automating metadata, and ensuring consistent, reliable data across apps.
- [How LinkedIn Built Northguard and Xinfra to Move Beyond Kafka](https://scaleengineer.com/blog/how-linkedin-built-northguard-and-xinfra-to-move-beyond-kafka)
  - LinkedIn built Northguard and Xinfra to overcome Kafka's scaling limits with self-balancing storage, distributed metadata, and seamless migration.
- [How Uber Uses Pull-Based Ingestion to Keep Search Data Fresh at Massive Scale](https://scaleengineer.com/blog/how-uber-uses-pull-based-ingestion-to-keep-search-data-fresh-at-massive-scale)
  - Uber uses Kafka-based pull ingestion in OpenSearch to handle traffic spikes, simplify recovery, and maintain global search consistency.
- [How Cloudflare Built an Intelligent Maintenance Scheduler using Workers](https://scaleengineer.com/blog/how-cloudflare-built-an-intelligent-maintenance-scheduler-using-workers)
  - Cloudflare uses graphs, caching, and real-time analytics to automate maintenance scheduling and prevent infrastructure conflicts at scale.
- [How Anthropic Built Claude's Multi-Agent Research System](https://scaleengineer.com/blog/how-anthropic-built-claude-s-multi-agent-research-system)
  - Anthropic's Claude Research uses multiple AI agents that collaborate, reason, and coordinate to tackle complex research more effectively.
- [Why Netflix Replaced Its Custom Batch Scheduler with Kueue](https://scaleengineer.com/blog/why-netflix-replaced-its-custom-batch-scheduler-with-kueue)
  - Netflix replaced its custom batch scheduler with Kueue, simplifying scheduling and migrating millions of batch jobs seamlessly.
- [How Meta Migrated One of the World's Largest Data Ingestion Systems Without Downtime](https://scaleengineer.com/blog/how-meta-migrated-one-of-the-world-s-largest-data-ingestion-systems-without-downtime)
  - Meta migrated thousands of data ingestion jobs using shadow testing, automated validation, and safe rollbacks without disrupting users.
- [AI Can Write Code in Seconds. But Can It Stop Malware Too?](https://scaleengineer.com/blog/ai-can-write-code-in-seconds-but-can-it-stop-malware-too)
  - Replit integrates Socket Firewall to analyze AI-suggested packages in real time, blocking malicious dependencies before installation.
- [How NVIDIA is Using Agentic AI to Build Autonomous Telecom Networks](https://scaleengineer.com/blog/how-nvidia-is-using-agentic-ai-to-build-autonomous-telecom-networks)
  - Telcos are moving beyond predefined automation toward agentic AI that can reason, research, optimize, and safely operate networks.
- [How Google’s A2A Is Changing How AI Agents Collaborate?](https://scaleengineer.com/blog/how-google-s-a2a-is-changing-how-ai-agents-collaborate)
  - A2A lets specialized AI agents securely collaborate and delegate tasks, turning isolated agents into a connected ecosystem of autonomous capabilities.
- [How DeepSeek Runs 380,000 Agent Sandboxes at Once](https://scaleengineer.com/blog/how-deepseek-runs-380-000-agent-sandboxes-at-once)
  - How DeepSeek built DSec to run 380K+ agent sandboxes concurrently while optimizing isolation, resources, images, and security.
- [How Zomato Made Its Restaurant Partner App Over 90% Faster](https://scaleengineer.com/blog/how-zomato-made-its-restaurant-partner-app-over-90-faster)
  - How Zomato optimized Android startup, rendering, memory, and background work to make its Restaurant Partner App significantly faster.
