Principal Software Engineer - Elastic Global Services
Principal Engineer · Principal · Full Time
Opens Snowflake's application page
Role
What you'll do.
As a Principal Software Engineer in Elastic Global Services at Snowflake, you will architect and lead the development of highly available, scalable, multi-tenant cloud services infrastructure that powers the Data Cloud platform across AWS, Azure, and GCP. This role combines hands-on distributed systems engineering with technical leadership, requiring 15+ years of experience designing large-scale fault-tolerant infrastructure, container orchestration, and cluster management. You'll solve critical challenges in autoscaling, workload orchestration, and operational excellence while mentoring engineers and shaping the future of AI-native enterprise computing.
Responsibilities
- Design and Architect Distributed Systems: Lead the design and implementation of scalable distributed systems for cloud services infrastructure, making critical architectural decisions around consistency models, fault tolerance, and system reliability across multi-cloud environments.
- Drive Large-Scale Infrastructure Solutions: Apply software engineering and analytical problem-solving skills to solve real business needs at enterprise scale, addressing infrastructure challenges that directly impact Snowflake's ability to serve thousands of customers globally.
- Optimize Performance and Availability: Analyze and resolve fault-tolerance, high availability, performance, and scalability challenges in production systems, ensuring services meet stringent SLAs and customer commitments for availability and performance metrics.
- Technical Leadership and Mentorship: Mentor and develop junior and mid-level engineers on the team, sharing distributed systems expertise and guiding them through complex technical challenges in building infrastructure systems.
- Manage Infrastructure Trade-offs: Evaluate and balance critical trade-offs between consistency, durability, cost, and performance to build solutions that meet the demands of rapidly evolving services and growing customer needs.
- Ensure Operational Excellence: Drive operational readiness of production services through robust monitoring, incident response, and disaster recovery practices, maintaining infrastructure reliability and meeting customer SLOs.
Qualifications
What we look for.
Technical
Large-Scale Distributed Systems Architecture
Demonstrated expertise designing, building, and supporting distributed fault-tolerant infrastructure in production environments at scale, with deep understanding of distributed system principles including consensus algorithms, replication, and eventual consistency.
Container Orchestration and Cluster Management
Proven experience with container orchestration platforms and cluster management technologies, including hands-on knowledge of Kubernetes, Mesos, OpenShift, or comparable container platforms with understanding of their internals and operational characteristics.
Autoscaling and Resource Management
Experience designing and implementing autoscaling solutions for cloud infrastructure, including VM management across cloud providers, workload orchestration, and dynamic resource allocation in multi-tenant environments.
Operating Systems and Systems Programming
Expert-level understanding of operating system concepts including multi-threading, memory management, networking protocols, storage systems, and performance optimization techniques at the systems level.
Cloud Platform Architecture
Deep experience architecting solutions across multiple cloud providers (AWS, Azure, GCP), understanding provider-specific infrastructure primitives, networking capabilities, and optimization strategies for multi-cloud deployments.
Performance and Scale Engineering
Track record of optimizing systems for performance and scale, including profiling, benchmarking, bottleneck identification, and implementation of solutions that operate reliably under extreme load conditions.
Education
Computer Science or Related Field
Bachelor's degree in Computer Science, Computer Engineering, Mathematics, or equivalent practical experience demonstrating mastery of fundamental computing concepts and theoretical foundations.
Experience
Infrastructure and Platform Engineering
15+ years of software engineering experience with substantial time spent designing, building, and operating large-scale production infrastructure systems serving millions of requests or managing petabyte-scale data.
Distributed Systems Production Experience
Extensive hands-on experience troubleshooting and resolving production issues in distributed systems, including debugging network partitions, managing state consistency, and designing recovery mechanisms.
Leadership in Technical Projects
Demonstrated ability to lead complex technical initiatives, drive architectural decisions across large systems, and influence organizational technical direction through technical excellence and strategic thinking.
Cross-Functional Collaboration
Experience working closely with customers, product teams, and partners to understand requirements, translate business needs into technical solutions, and innovate strategically within infrastructure constraints.
Skills
Required
Distributed Systems Design
Ability to architect systems that handle consistency, availability, and partition tolerance trade-offs; experience with consensus algorithms, replication strategies, and failure recovery.
Kubernetes Administration and Architecture
Deep proficiency with Kubernetes including cluster architecture, networking, storage orchestration, and operational management of large-scale deployments.
Cloud Infrastructure
Expert knowledge of cloud services from major providers (AWS, Azure, GCP), including compute, networking, and storage offerings, with ability to architect multi-cloud solutions.
Systems Programming Languages
Production experience with systems-level programming languages such as C++, Go, or Rust for infrastructure components requiring performance and reliability.
Scalable System Architecture
Proven ability to design and implement systems that scale linearly or sub-linearly with load, including horizontal scaling strategies and resource optimization.
Preferred
Container Runtime Internals
Nice to haveUnderstanding of container runtime implementations, including Docker, containerd, or other container engines, and their integration with orchestration platforms.
Kubernetes Advanced Features
Nice to haveExperience with Kubernetes advanced capabilities such as custom controllers, operators, custom resource definitions (CRDs), and Kubernetes API extensions.
Multi-Cloud Orchestration
Nice to haveExperience orchestrating workloads across multiple cloud providers, managing cloud-agnostic abstractions, and optimizing costs in heterogeneous cloud environments.
Infrastructure as Code
Nice to haveProficiency with infrastructure automation tools such as Terraform, CloudFormation, or Helm for managing complex infrastructure deployments at scale.
Observability and Monitoring
Nice to haveExperience designing and implementing comprehensive monitoring, logging, and tracing solutions for distributed systems, including tools like Prometheus, ELK, and distributed tracing platforms.
Data Platform Experience
Nice to haveBackground working on data warehousing, analytics platforms, or data cloud infrastructure, with understanding of data processing pipelines and large-scale data management.
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 264,000 – 379,500
Equity·Stock options
Benefits
Comprehensive Health Coverage
Medical, dental, and vision insurance with options for employee and family coverage, designed to meet diverse healthcare needs.
Retirement Planning
401(k) retirement savings plan with company matching contributions to support long-term financial security and retirement planning.
Equity Compensation
Stock options and RSU grants aligned with company performance, allowing you to share in Snowflake's long-term success and value creation.
Paid Time Off
Generous paid vacation, sick leave, and personal days to maintain work-life balance and recharge while working on challenging technical problems.
Professional Development
Learning and development budget for conference attendance, training programs, and educational resources to advance technical expertise and career growth.
Remote Work Flexibility
Flexible work arrangements with options for remote work and distributed team collaboration, enabling you to balance professional responsibilities with personal needs.
Wellness Programs
Access to fitness facilities, mental health resources, wellness challenges, and employee assistance programs supporting overall health and wellbeing.
Parental Leave
Paid parental leave benefits for welcoming new family members while maintaining financial stability during important life transitions.
Process
Interview steps.
- 01
Initial Recruiter Conversation
Phone or video screening with a recruiting team member to discuss your background, experience with distributed systems, and interest in the Elastic Global Services team at Snowflake.
- 02
Technical Screening
Detailed technical conversation with a senior engineer covering distributed systems concepts, past project experiences, and approaches to solving large-scale infrastructure challenges.
- 03
System Design Discussion
In-depth discussion on designing and architecting large-scale distributed systems, covering trade-offs between consistency, availability, durability, and cost considerations.
- 04
Infrastructure Architecture Presentation
Presentation or detailed discussion of a significant infrastructure project you've led, including architectural decisions, challenges faced, and lessons learned in scaling systems.
- 05
Leadership and Collaboration Assessment
Conversation with team leadership or peers focused on your experience mentoring engineers, driving technical decisions, and collaborating across organizations on complex initiatives.
- 06
Final Leadership Round
Interview with engineering leadership or the team's management to discuss your vision for infrastructure engineering, career aspirations, and alignment with Snowflake's technical direction.
Full posting
Original listing.
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done.
Elastic Global Services team is responsible for building the highly available, scalable, multi-tenant “Cloud Services” platform that underpin Snowflake services. Areas we work on include, the autoscaling of VMs from cloud providers, managing topologies of our compute clusters, cluster management, workload orchestration and many others. We are looking at expanding our team to handle the next big challenges for Snowflake customers.
Our product offering runs on multiple cloud providers including Amazon Web Services, Microsoft Azure and Google Cloud. Our infrastructure self-optimizes, provides high availability and data protection across cloud providers so our users can focus on using their data, not managing it. In our effort to enable our Data Cloud vision, we are actively hiring talented distributed systems engineers. This role is a unique opportunity to make a significant impact on our elastic, large scale, high-performance computing environment.
To learn more about the team’s tech stack, see our recent talk at ACM Symposium on Cloud Computing! https://acmsocc.org/2022/assets/slides/99.pdf
AS A PRINCIPAL SOFTWARE ENGINEER AT SNOWFLAKE YOU WILL:
Solving real business needs at large scale by applying your software engineering and analytical problem solving skills.
Design and implement scalable distributed systems for our cloud services.
Analyze fault-tolerance and high availability issues, performance and scale challenges, and solve them.
Mentor and grow junior engineers.
Understand trade-offs between consistency, durability and costs to build solutions which can meet the demands of rapidly growing services.
Ensure operational readiness of the services and meet the commitments to our customers regarding availability and performance.
OUR IDEAL DISTRIBUTED SYSTEMS ENGINEER WILL HAVE:
15+ years of industry experience designing, building and supporting large scale infrastructure in production.
Experience building large scale distributed fault tolerant infrastructure.
Experience in container orchestration, cluster management, or autoscaling.
Excellent understanding of operating systems concepts including. multi-threading, memory management, networking and storage, performance and scale.
Solid understanding of the internals of Kubernetes, Mesos, OpenShift, or other container platforms.
WHY JOIN THE ENGINEERING TEAM AT SNOWFLAKE?
Build an industry-leading Cloud Data Platform.
Solve challenging technical problems related to security, parallel and distributed systems, programming, resource management, large-scale system maintenance, and more!
Work closely with our customers & partners, understand their use cases & needs, think strategically to seek the right problem to solve at the right time, and innovate with rigor.
Join a world-class team of both industry veterans and rising stars.
Snowflake is growing fast, and we’re scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake.
How do you want to make your impact?
Snowflake is growing fast, and we’re scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake.
How do you want to make your impact?
For jobs located in the United States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.com
Redirects to Snowflake's application page.
Other roles
More at Snowflake.
Staff Software Engineer – Engineering Systems, Continuous Integration Team
Staff
Account Engineer
Mid
Director of Engineering - Traffic & Networking
Director
Principal Software Engineer 2 - OLTP
Principal
Engineering Manager, Release Validation
Manager