Staff Software Engineer - FDB Platform

Staff Engineer · Staff · Full Time

US-CA-Menlo ParkUSD 236k – 339k1mo ago
Apply for this role

Opens Snowflake's application page

Role

What you'll do.

Staff Software Engineer role at Snowflake focused on designing and implementing scalable distributed systems for the FDB (FoundationDB-based) platform that powers the Snowflake Data Cloud across AWS, Azure, and GCP. This position requires 8+ years of infrastructure experience with deep expertise in distributed systems, container orchestration, and large-scale database technologies to architect cloud-agnostic solutions for autoscaling, self-healing clusters, and cost-optimized operations.

Responsibilities

  • Design Scalable Distributed System Solutions: Architect and design cloud-agnostic distributed system solutions for the FDB platform infrastructure, considering multi-cloud deployment across AWS, Azure, and GCP with focus on elastic scalability, fault tolerance, and operational simplicity
  • Solve Fault-Tolerance and High Availability Challenges: Analyze complex fault-tolerance scenarios and high availability requirements across distributed infrastructure, implement robust solutions that prevent single points of failure and ensure continuous service availability at scale
  • Address Performance and Scale Challenges: Identify performance bottlenecks and scalability limitations in the FDB platform infrastructure, conduct thorough analysis, implement optimizations, and validate improvements through testing and production monitoring
  • Own End-to-End Project Delivery: Take complete ownership of complex infrastructure projects from problem identification and solution design through implementation, comprehensive testing, performance validation, and safe production rollout with minimal operational risk
  • Build Consistency-Aware Solutions: Understand and apply nuanced trade-offs between consistency guarantees, durability requirements, and cost implications to build solutions that meet the demands of rapidly growing Snowflake services while optimizing operational expenses
  • Develop Next-Generation Transaction and Storage Systems: Build advanced transaction processing systems, intelligent caching layers, optimized storage engines, and multi-tenant isolation capabilities that power Snowflake's expanding product portfolio
  • Evangelize Database Best Practices: Share expertise across the organization regarding distributed database usage patterns, end-to-end system architecture principles, and operational excellence to elevate engineering practices company-wide
  • Instrument and Debug Production Systems: Implement comprehensive instrumentation and observability across FDB platform components, diagnose complex production issues through systematic analysis, and develop solutions that address root causes and improve system resilience
  • Drive Infrastructure Automation: Develop and enhance self-managing infrastructure capabilities including autoscaling based on utilization and traffic patterns, automatic cluster provisioning with zero manual intervention, and self-healing mechanisms that prevent or mitigate production impact
  • Optimize Operational Cost Efficiency: Design and implement self-optimizing systems that ensure FDB clusters run at optimal resource utilization, minimize cost-of-goods-sold (COGS), and maintain efficiency as workload patterns evolve

Qualifications

What we look for.

Technical

  • Large-Scale Distributed Systems Design

    Proven expertise designing, building, and operating production distributed systems infrastructure at scale, with demonstrated understanding of trade-offs between consistency, durability, availability, and cost

  • Container Orchestration Mastery

    Solid understanding of Kubernetes, Mesos, OpenShift, or equivalent container platforms, including internal architecture, scheduling algorithms, networking models, and operational patterns

  • Systems Programming

    Fluency in Java with strong foundation in multi-threading, concurrency primitives, memory management, and performance optimization for high-throughput distributed systems

  • Key-Value Store Expertise

    Practical experience designing, deploying, or operating scalable key-value stores such as FoundationDB, RocksDB/LevelDB, DynamoDB, Redis, or Cassandra in production environments

  • Operating Systems Knowledge

    Deep understanding of kernel concepts including multi-threading models, memory management strategies, networking stack internals, storage I/O optimization, and performance profiling

  • Cloud-Native Architecture

    Experience building and managing distributed systems across multiple cloud providers (AWS, Azure, GCP) with understanding of cloud-agnostic design principles and multi-cloud orchestration

Education

  • Bachelor's Degree in Computer Science

    Required: BS in Computer Science or equivalent practical experience demonstrating advanced systems knowledge

  • Advanced Degree Preferred

    Preferred: Master's degree or PhD in Computer Science, distributed systems, or related technical field that demonstrates deepened expertise in algorithms, systems design, and theoretical foundations

Experience

  • 8+ Years Infrastructure Engineering

    Minimum eight years of industry experience designing, building, deploying, and supporting large-scale infrastructure systems in production environments with responsibility for system reliability and performance

  • Stateful Service Operations

    Significant hands-on experience designing, implementing, and operating large-scale distributed systems infrastructure that manages stateful services with high availability requirements

  • Complex Project Delivery

    Demonstrated track record of owning and delivering highly complex projects in the distributed systems space from conception through design, implementation, testing, and production rollout

  • Big Data and Storage Technologies

    Professional experience with big data storage technologies, distributed file systems (HDFS), columnar databases, or related data infrastructure platforms

Skills

Required

  • Distributed Systems Architecture

    Design and implementation of fault-tolerant, highly available distributed systems with expertise in consensus algorithms, replication strategies, and failure recovery mechanisms

  • Java Programming

    Production-level systems programming in Java with deep expertise in concurrency, multi-threading, performance optimization, and low-latency system design

  • Kubernetes or Container Orchestration

    Deep knowledge of container orchestration platforms for cluster management, resource scheduling, service discovery, and automated workload orchestration

  • Database Systems Knowledge

    Understanding of distributed database internals, transaction processing, consistency models, durability mechanisms, and multi-tenancy architecture

  • Cloud Infrastructure Operations

    Hands-on experience deploying and managing distributed systems on public cloud platforms including AWS, Azure, or GCP with understanding of cloud-native operational patterns

  • Performance Analysis and Optimization

    Ability to identify bottlenecks, instrument systems for observability, analyze performance metrics, and implement optimizations for scale and efficiency

Preferred

  • FoundationDB Experience

    Nice to have

    Prior hands-on experience with FoundationDB architecture, operational characteristics, or contribution to FoundationDB ecosystem highly valued

  • Multi-Cloud Architecture

    Nice to have

    Expertise building and operating systems across multiple cloud providers with understanding of cloud-agnostic abstraction patterns and provider-agnostic infrastructure-as-code

  • Autoscaling and Self-Healing Systems

    Nice to have

    Design and implementation experience building self-managing infrastructure with auto-scaling capabilities, failure detection, and self-healing mechanisms

  • Big Data Technologies

    Nice to have

    Experience with HDFS, Cassandra, columnar databases, or other big data storage technologies and their distributed operation at scale

  • Go or C++ Systems Programming

    Nice to have

    Additional expertise in systems languages like Go or C++ for performance-critical infrastructure components

  • Infrastructure-as-Code

    Nice to have

    Proficiency with Terraform, CloudFormation, or similar infrastructure automation tools for declarative infrastructure management

  • Observability and Monitoring

    Nice to have

    Experience building comprehensive monitoring, logging, and distributed tracing systems for production distributed systems

Tech stack

Languages

JavaPythonGoC++

Frameworks

KubernetesMesosOpenShift

Databases

FoundationDBRocksDBDynamoDBCassandraRedis

Tools

Amazon Web Services (AWS)Microsoft AzureGoogle Cloud Platform (GCP)TerraformDocker

Other

Distributed Systems DesignMulti-threading and ConcurrencyOperating SystemsCloud Infrastructure AutomationHDFS and Big Data Technologies

Compensation

Pay and benefits.

Base·USD 236,000 – 339,250

Equity·Stock options

Benefits

  • Equity Compensation

    Stock options providing ownership stake and long-term value participation in Snowflake's continued growth as a high-growth cloud computing company

  • Comprehensive Health Insurance

    Medical, dental, and vision coverage for employees and eligible dependents

  • 401(k) Retirement Plans

    Tax-advantaged retirement savings plans with company match contributions

  • Professional Development

    Learning and development budgets to support continued technical growth, conference attendance, and skill development in distributed systems and cloud technologies

  • Flexible Work Environment

    Flexibility to work remotely or in office with support for work-life balance as a high-growth technology company culture

  • Collaborative Innovation Culture

    Environment built on impact, innovation, and collaboration where technical excellence is valued and challenging problems drive career advancement

Process

Interview steps.

  1. 01

    Initial Screening

    Recruiter phone screen to assess background, experience with distributed systems, and alignment with Staff-level expectations for ownership and complexity handling

  2. 02

    Technical Phone Screen

    Senior engineer conversation exploring distributed systems design thinking, specific experience with key-value stores or database infrastructure, and approach to solving complex architectural challenges

  3. 03

    System Design Interview

    In-depth technical interview requiring you to design a large-scale distributed system component, articulate trade-offs (consistency vs. availability vs. cost), explain failure modes, and discuss operational considerations

  4. 04

    Deep Dive Technical Assessment

    Discussion of past complex infrastructure project you owned, focusing on design decisions, challenges overcome, lessons learned, and how you would approach similar problems differently with current knowledge

  5. 05

    Infrastructure Knowledge Assessment

    Technical evaluation of expertise in Kubernetes or similar container orchestration platforms, demonstrating understanding of scheduling algorithms, networking models, and operational patterns

  6. 06

    Leadership and Collaboration Discussion

    Conversation exploring your approach to architecture evangelism, how you influence engineering practices, mentoring approach, and track record of elevating team technical capabilities

  7. 07

    Hiring Manager Interview

    Discussion with engineering leadership regarding vision alignment, appetite for solving hard distributed systems problems, working style in fast-moving environments, and long-term technical growth aspirations

Full posting

Original listing.

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done.

Snowflake is a high-growth, cloud-native data platform company committed to empowering enterprises to achieve their full potential. With a culture built on impact, innovation, and collaboration, we offer an environment where you can build large-scale systems, move fast, and take your technology career to the next level.

We are seeking an outstanding Staff Software Engineer with a passion for large scale databases and distributed systems to help us take the FDB platform to the next level. A massive new market opportunity is being created at the intersection of Cloud and Data, and the Snowflake Data Cloud is leading the way, all powered by the database engine we are building from the ground up.

Key to Snowflake’s Database Engine is our large scale distributed transactional Key-Value store - called FDB - which powers all of Snowflake’s products and services and is rapidly evolving to meet Snowflake’s future needs.

FDB runs on multiple cloud providers including Amazon Web Services, Microsoft Azure and Google Cloud. The elastic infrastructure FDB runs on is being built from the ground up and is envisioned to be a cloud agnostic, fully automated manageability platform that provides:

  • Autoscaling and auto-balancing of clusters based on utilization, traffic and workloads

  • Auto-provisioning of new clusters with zero manual intervention

  • Self-healing capabilities that prevent, mitigate and resolve any production impact

  • Built-in configuration management that guarantees FDB runs correctly and on the intended topologies

  • Self-optimizing COGS efficiency, ensuring we run our clusters at optimal utilization

AS A STAFF SOFTWARE ENGINEER AT SNOWFLAKE, YOU WILL:

  • Design and implement scalable distributed system solutions for our cloud agnostic platform.

  • Analyze fault-tolerance and high availability issues, performance and scale challenges, and solve them.

  • Own the end to end delivery of your projects, from identifying a solution, to design, implementation, test and safe production rollout

  • Understand trade-offs between consistency, durability and costs to build solutions which can meet the demands of rapidly growing services.

  • Build the next generation transaction system, caching, storage engine and multi tenant capabilities

  • Evangelize best practices in database usage and end-to-end architecture.

  • Pinpoint problems, instrument relevant components as needed, and ultimately implement solutions.

OUR IDEAL STAFF SOFTWARE ENGINEER WILL HAVE:

  • 8+ years industry experience designing, building and supporting large scale infrastructure in production.

  • Experience designing, building, and operating large-scale distributed systems infrastructure supporting stateful services

  • Experience in container orchestration, cluster management, or autoscaling.

  • Excellent understanding of operating systems concepts including multi-threading, memory management, networking and storage, performance and scale.

  • Systems programming skills including multi-threading, concurrency, etc. Fluency in Java

  • Solid understanding of the internals of Kubernetes, Mesos, OpenShift, or other container platforms

  • Experience with scalable Key-Value stores such as FoundationDB, RocksDB/LevelDB, DynamoDB, Redis, etc. a plus.

  • Track record of delivering highly complex projects in the distributed systems space

  • Intense curiosity, willingness to question and passion for making systems better

  • Experience with one or more of the following highly desired:

    • Big Data storage technologies and their applications (HDFS, Cassandra, Columnar Databases, etc.)

    • Scalable Key-Value stores such as FoundationDB, RocksDB/LevelDB, DynamoDB, Redis, Cassandra, etc.

  • BS in Computer Science; Masters or PhD Preferred.

Snowflake is growing fast, and we’re scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake.

How do you want to make your impact?

For jobs located in the United States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.com

Redirects to Snowflake's application page.

Other roles

More at Snowflake.

View all 78 roles