Senior Software Engineer, Accelerated Delivery
Backend Engineer · Senior · Full Time
Opens Snowflake's application page
Role
What you'll do.
Join Snowflake's Release Engineering team as a Senior Software Engineer to design and build large-scale continuous deployment infrastructure for multi-cloud production environments. This role combines platform engineering, distributed systems reliability, and DevOps expertise to create safe, scalable, and efficient software delivery systems augmented by AI-driven automation. You'll partner with engineering teams across Snowflake to eliminate deployment friction, implement progressive delivery patterns, and build self-service developer tooling while maintaining operational safety at global scale.
Responsibilities
- Design and Build Continuous Deployment Infrastructure: Design and architect continuous deployment and rollout infrastructure capable of safely shipping changes across Snowflake's large-scale, multi-cloud production environment. Focus on creating systems that provide reliability, observability, and auditability for every production deployment, supporting rapid iteration while minimizing operational risk.
- Implement Progressive Delivery Capabilities: Build and evolve platform capabilities for progressive delivery patterns including staged rollouts, canary deployments, automated health checks, intelligent rollback controls, and blast radius minimization guardrails. Design mechanisms that enable teams to validate changes safely in production before full rollout.
- Eliminate Release Pipeline Friction: Improve engineering velocity by identifying and removing friction from release pipelines, replacing error-prone manual workflows with durable, maintainable platform abstractions, robust automation, and self-documenting processes that reduce cognitive load on engineering teams.
- Build Internal Release Orchestration Platforms: Develop large-scale release orchestration platforms supporting application rollouts on Kubernetes, production change workflows, and cross-service coordination. Create abstractions that allow teams to adopt consistent deployment patterns without requiring deep release engineering expertise.
- Develop Observability and Health Evaluation Systems: Build systems that evaluate rollout health using metrics, logs, alerts, and operational signals from observability platforms like Prometheus, Datadog, or Grafana. Implement automated regression detection and safe mitigation paths that trigger intelligent rollback decisions with minimal blast radius.
- Partner with Product and Infrastructure Teams: Collaborate with product and infrastructure teams to design platform capabilities that make services easier to deploy, validate, observe, and operate. Drive adoption of release best practices through well-designed developer experiences and self-service tooling that reduces operational toil.
- Implement Advanced Deployment Methodologies: Research, design, and implement deployment methodologies including GitOps-inspired workflows, infrastructure-as-code practices, policy-driven automation, and progressive delivery patterns tailored to Snowflake's multi-cloud environment and operational requirements.
- Build AI-Assisted and Autonomous Workflows: Design and implement AI-assisted, agentic-driven, and increasingly autonomous release workflows that enhance rollout intelligence, improve developer productivity, and strengthen deployment safety through intelligent automation and predictive analysis.
- Create Self-Service Developer Tooling: Develop self-service developer tools and platforms that enable teams across Snowflake to adopt safe deployment patterns, manage their own rollouts, and participate in deployment decisions without requiring deep release engineering expertise or tribal knowledge.
- Build Automation and Operational Guardrails: Design and implement guardrails and automation systems that reduce operational toil, make production change workflows more consistent and resilient, enforce best practices automatically, and ensure the right operational path remains the easiest path.
Qualifications
What we look for.
Technical
Continuous Deployment Platform Experience
Proven experience building or operating continuous deployment, release engineering, or production change platforms at scale, with demonstrated expertise in managing safe deployments across large, complex distributed systems.
Kubernetes and Container Orchestration
Strong hands-on experience with Kubernetes-based systems and understanding of how to safely orchestrate and validate changes across distributed production environments, including knowledge of Helm, service meshes, and deployment strategies.
Systems Programming Languages
Strong software engineering skills in at least one systems language such as Golang, Java, or C++, with ability to build performant, reliable infrastructure components and understand low-level system behavior.
Scripting and Automation
Proficiency in scripting languages such as Python or Bash for infrastructure automation, creating runbooks, implementing CI/CD pipeline logic, and building operational tools that improve deployment efficiency.
Distributed Systems Design
Deep understanding of distributed systems principles including consistency models, failure modes, eventual consistency, replication strategies, and architectural patterns that apply to large-scale deployment infrastructure.
Infrastructure Automation and IaC
Experience with infrastructure automation tools, infrastructure-as-code practices, configuration management, and managing large-scale deployments across multiple cloud providers and on-premises environments.
CI/CD Pipeline Design
Expertise in designing, building, and maintaining sophisticated CI/CD pipelines that support automated testing, building, and deployment processes for large-scale systems with high uptime requirements.
Multi-Cloud Infrastructure
Experience working with multi-cloud environments and understanding how to abstract away cloud-specific details while leveraging platform-specific optimizations and managed services effectively.
Observability and Monitoring
Hands-on experience with observability platforms such as Prometheus, Datadog, Grafana, or similar tools for metrics collection, distributed tracing, log aggregation, and building data-driven deployment decisions.
Safe Production Rollout Practices
Deep commitment to and proven experience implementing safe production change practices including blast radius minimization, automated rollback mechanisms, canary deployments, and progressive delivery patterns.
Education
Computer Science or Related Field
Bachelor's degree in Computer Science, Computer Engineering, or related discipline, or equivalent professional experience demonstrating strong foundational knowledge of distributed systems, algorithms, and software architecture principles.
Experience
Large-Scale Systems Operation
5+ years of experience operating or building systems at significant scale, dealing with distributed system challenges, failure scenarios, and infrastructure decisions affecting high-availability production environments serving thousands of users.
Platform Engineering Leadership
Demonstrated ability to design platform abstractions that improve developer productivity across large engineering organizations, with experience translating complex operational requirements into elegant, maintainable systems.
DevOps and Release Engineering
Substantial experience in DevOps practices, release engineering methodologies, or site reliability engineering roles with focus on improving deployment safety, reducing operational overhead, and scaling delivery capabilities.
Open Source Infrastructure Projects
Familiarity with open-source infrastructure and deployment projects such as Kubernetes, GitOps tools, progressive delivery platforms, or similar systems that demonstrate engagement with modern infrastructure challenges.
AI and Automation in Operations
Experience or demonstrated interest in applying AI, machine learning, and intelligent automation to operational workflows, deployment decisions, and autonomous systems that improve efficiency and reduce human error.
Skills
Required
Golang or Java
Strong proficiency in Golang or Java for building reliable, high-performance infrastructure and platform components that handle complex deployment orchestration and system reliability challenges.
Python or Bash
Proficiency in Python or Bash scripting for infrastructure automation, operational tooling, CI/CD pipeline logic, and building the glue that connects various deployment system components.
Kubernetes
Substantial hands-on experience with Kubernetes architecture, API resources, deployment strategies, and operational best practices for managing containerized applications at scale in production.
Distributed Systems Fundamentals
Solid understanding of distributed systems concepts including consensus algorithms, eventual consistency, failure handling, state management, and architectural tradeoffs relevant to deployment infrastructure.
CI/CD Platforms
Experience with CI/CD platforms and tools such as GitHub Actions, GitLab CI, Jenkins, or similar systems for building automated delivery pipelines that support safe production deployments.
Infrastructure Observability
Ability to instrument systems with metrics, logs, and traces, and use observability platforms to understand system behavior, diagnose issues, and make data-driven deployment decisions.
Production Incident Management
Demonstrated experience handling production incidents, performing root cause analysis, designing safeguards to prevent recurrence, and improving system resilience based on operational learnings.
Preferred
Canary Deployment Patterns
Nice to haveExperience designing or implementing canary deployment systems, progressive delivery frameworks, or similar mechanisms that reduce risk during production rollouts.
GitOps Workflows
Nice to haveFamiliarity with GitOps principles and tools like Flux or ArgoCD that use git repositories as the single source of truth for infrastructure and deployment configurations.
Service Mesh Technologies
Nice to haveExperience with service mesh technologies such as Istio, Linkerd, or Envoy that provide traffic management, observability, and security capabilities for distributed microservice systems.
Infrastructure as Code Tools
Nice to haveExpertise with IaC tools such as Terraform, Pulumi, CloudFormation, or similar platforms for declaratively managing cloud infrastructure and deployment configurations at scale.
Prometheus and Grafana
Nice to haveHands-on experience with Prometheus for metrics collection and Grafana for visualization, or similar observability stacks for building comprehensive deployment health dashboards and alerting systems.
Policy as Code
Nice to haveFamiliarity with policy-as-code frameworks such as OPA/Rego, Kyverno, or similar tools that enforce deployment rules, security policies, and operational guardrails automatically.
ML and Anomaly Detection
Nice to haveInterest or experience in applying machine learning or anomaly detection techniques to operational data for early regression detection, predictive rollout health analysis, or autonomous deployment decisions.
Multi-Cloud Orchestration
Nice to haveExperience managing deployments across multiple cloud providers (AWS, GCP, Azure) or hybrid cloud environments with understanding of cloud-agnostic abstractions and cross-cloud tooling.
Agentic AI Systems
Nice to haveExperience or demonstrated understanding of agentic AI systems, autonomous workflows, or AI-driven operational decision making that could enhance release engineering automation.
Database Systems Knowledge
Nice to haveUnderstanding of data warehousing, analytics platforms, or distributed database systems that would aid in appreciating Snowflake's product and deployment requirements.
Compensation
Pay and benefits.
Base·USD 200,000 – 287,500
Full posting
Original listing.
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done.
Snowflake’s Release Engineering team builds and operates the systems that safely deliver infrastructure, platform, and product changes to production at global scale. We own the release platforms, rollout orchestration, and safety mechanisms that allow engineering teams across Snowflake to ship quickly while minimizing operational risk. Our mission is to make production deployments fast, safe, self-service, increasingly autonomous, and augmented by AI-driven intelligence and automation.
This role sits at the intersection of developer productivity, distributed systems reliability, and large-scale multi-cloud infrastructure orchestration. At Snowflake, Release Engineering is a platform engineering function focused on building the systems, abstractions, and automation that make software delivery safe, scalable, and efficient across the company.
In this role, you will
Design and build continuous deployment and rollout infrastructure that safely ships changes across Snowflake’s large-scale, multi-cloud production environment.
Build and evolve platform capabilities for progressive delivery, including staged rollouts, canarying, automated health checks, rollback controls, and guardrails that reduce blast radius during production change events.
Improve engineering velocity by removing friction from release pipelines and replacing manual workflows with durable platform abstractions and automation.
Build internal platforms that support large-scale release orchestration, application rollouts on Kubernetes, and broader production change workflows.
Partner with product and infrastructure teams to make their services easier to deploy, validate, observe, and operate through well-designed platform capabilities.
Implement and evolve deployment methodologies such as GitOps-inspired workflows, infrastructure as code, policy-driven automation, and progressive delivery patterns appropriate for Snowflake’s environment.
Build systems that evaluate rollout health using metrics, logs, alerts, and operational signals to detect regressions early and trigger safe mitigation or rollback paths.
Develop self-service developer tooling that enables teams across Snowflake to adopt safe deployment patterns without requiring deep release expertise.
Build automation and guardrails that reduce operational toil and make production change workflows more consistent, scalable, and resilient.
Design and build AI-assisted, agentic-driven, and increasingly autonomous release workflows that improve rollout intelligence, developer productivity, and deployment safety.
You may be a strong fit if you
Have experience building or operating continuous deployment, release engineering, or production change platforms at scale.
Have worked with Kubernetes-based systems and understand how to safely roll changes across distributed production environments.
Have strong software engineering skills in Golang, Java, C++, or similar systems languages, along with Python, Bash, or similar scripting languages.
Have experience with distributed systems, infrastructure automation, CI/CD pipelines, and cloud environments.
Bring a data-driven mindset and have experience using observability platforms such as Prometheus, Datadog, or Grafana to evaluate system and rollout health.
Care deeply about safe production rollouts, developer experience, minimizing blast radius, and building systems that make the right operational path the easiest one.
Enjoy building internal platforms and self-service systems that improve developer productivity across a large engineering organization.
Apply a combined software engineering and DevOps mindset to design, build, and continuously improve large-scale delivery platforms in production.
Are excited about applying AI and intelligent automation to release operations, deployment safety, and autonomous workflows.
Snowflake is growing fast, and we’re scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake.
How do you want to make your impact?
For jobs located in the United States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.com
Redirects to Snowflake's application page.
Other roles
More at Snowflake.
Software Engineer AI Team
Mid
Staff/Principal AI Software Engineer - Snowflake CoWork
Principal
Sr Manager, Applied Field Engineering - AI/ML
Manager
Principal Data Platform Architect
Principal
Senior Software Engineer - NatSec
Senior