Infrastructure Engineer, Observe by Snowflake

Infrastructure Engineer · Senior · Full Time

US-CA-Menlo ParkUSD 160k – 210k4mo ago
Apply for this role

Opens Snowflake's application page

Role

What you'll do.

Observe by Snowflake is seeking an Infrastructure Engineer to design, build, and operate scalable cloud infrastructure supporting an AI-powered observability platform. The ideal candidate will have strong cloud engineering skills, expertise in infrastructure-as-code, and experience with container orchestration and distributed systems.

Responsibilities

  • Cloud Infrastructure Management: Design, build, and operate scalable cloud infrastructure in AWS supporting a high-scale observability platform
  • System Reliability: Improve system reliability, performance, and operational visibility across development and production environments
  • CI/CD Development: Develop and maintain CI/CD pipelines and internal tooling to improve developer productivity and deployment safety
  • Security Management: Identify and mitigate security risks, and help maintain internal security standards and compliance requirements
  • Infrastructure Resilience: Build infrastructure that supports high availability, scalability, and operational resilience
  • On-Call Support: Participate in an on-call rotation, contributing to incident response and post-incident improvements
  • Cross-Team Collaboration: Partner closely with engineering teams to ensure infrastructure supports evolving product and platform needs

Qualifications

What we look for.

Technical

  • Cloud Platforms

    Experience with AWS, GCP, and Azure cloud infrastructure

  • Container Orchestration

    Hands-on experience with Kubernetes and container management

  • Infrastructure as Code

    Proficiency in tools like Terraform, Ansible, or similar IaC technologies

  • Programming Languages

    Strong programming skills in Go, Python, or similar languages focused on automation and systems development

Education

  • Computer Science

    Bachelor's degree in Computer Science, Software Engineering, or related technical field preferred

Experience

  • Infrastructure Engineering

    Minimum 2+ years of experience in Infrastructure Engineering, SRE, or DevOps roles

  • Production Systems

    Experience supporting production systems at scale with a focus on reliability and operational excellence

Skills

Required

  • AWS Infrastructure

    Comprehensive knowledge of AWS cloud services and infrastructure design

  • Kubernetes

    Hands-on container orchestration and management skills

  • Infrastructure as Code

    Proficient in defining and managing infrastructure using code

  • System Automation

    Strong skills in developing automated solutions for infrastructure management

Preferred

  • Distributed Systems

    Nice to have

    Experience operating large-scale distributed systems

  • Observability Platforms

    Nice to have

    Familiarity with observability platforms, telemetry pipelines, or monitoring infrastructure

  • Developer Tooling

    Nice to have

    Experience improving developer platform tooling or internal infrastructure platforms

Tech stack

Languages

GoPython

Frameworks

Kubernetes

Tools

TerraformAnsible

Other

AWS

Compensation

Pay and benefits.

Base·USD 160,000 – 210,000

Equity·Stock options

Benefits

  • Competitive Compensation

    Salary range of $160K - $210K with potential stock options

  • Innovative Work Environment

    Opportunity to work at a cutting-edge cloud computing and observability platform

Process

Interview steps.

  1. 01

    Initial Screening

    Phone or video call with recruiting team to assess basic qualifications and background

  2. 02

    Technical Interview

    In-depth technical discussion focusing on infrastructure engineering skills and experience

  3. 03

    System Design Challenge

    Evaluate candidate's ability to design scalable and reliable infrastructure solutions

  4. 04

    Team Interview

    Meet with potential team members to assess cultural fit and collaboration potential

  5. 05

    Final Interview

    Discussion with hiring manager about role expectations and long-term goals

Full posting

Original listing.

Snowflake is about empowering enterprises to achieve their full potential — and people too. With a culture that’s all in on impact, innovation, and collaboration, Snowflake is the sweet spot for building big, moving fast, and taking technology — and careers — to the next level.

Observe by Snowflake is an AI-powered observability platform built on the Snowflake Data Cloud and engineered for scale. We ingest and store logs, metrics, traces, and events on an open, scalable data lake using open formats like Apache Iceberg, delivering deep correlation and long-term analytics at dramatically lower cost. A dynamic Knowledge Graph and chat-based AI SRE provide rich context and guided workflows so teams can move from detection to root cause and resolution significantly faster.

The Infrastructure team at Observe by Snowflake is responsible for building, scaling, and operating the development and production environments that power our observability platform. We are a small, highly collaborative team with a broad scope, focused on delivering reliable infrastructure while continuously improving the systems that support our engineers and customers.

What You’ll Do

  • Design, build, and operate scalable cloud infrastructure in AWS supporting a high-scale observability platform.

  • Improve system reliability, performance, and operational visibility across development and production environments.

  • Develop and maintain CI/CD pipelines and internal tooling to improve developer productivity and deployment safety.

  • Identify and mitigate security risks, and help maintain internal security standards and compliance requirements.

  • Build infrastructure that supports high availability, scalability, and operational resilience

  • Participate in an on-call rotation, contributing to incident response and post-incident improvements.

  • Partner closely with engineering teams to ensure infrastructure supports evolving product and platform needs.

What We’re Looking For

  • 2+ years of experience in Infrastructure Engineering, Site Reliability Engineering (SRE), DevOps, or related roles.

  • Experience operating container orchestration platforms such as Kubernetes

  • Hands-on experience managing cloud infrastructure using Infrastructure-as-Code tools such as Terraform, Ansible, or similar.

  • Strong programming skills in Go, Python, or similar languages, with a focus on automation and systems development.

  • Experience supporting production systems at scale, with a focus on reliability and operational excellence.

  • Strong problem-solving skills and the ability to balance short-term operational needs with long-term infrastructure design.

  • Experience with AWS, GCP, and Azure

Nice to Have

  • Experience operating large-scale distributed systems.

  • Familiarity with observability platforms, telemetry pipelines, or monitoring infrastructure.

  • Experience improving developer platform tooling or internal infrastructure platforms.

  • Experience working in high-growth or rapidly evolving engineering environments.

Snowflake is growing fast, and we’re scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake.

How do you want to make your impact?

For jobs located in the United States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.com

Redirects to Snowflake's application page.

Other roles

More at Snowflake.

View all 87 roles