Senior Platform Engineer

Platform Engineer · Senior · Full Time · Remote

United States or Canada (remote) · RemoteUSD 165k – 195k1mo ago
Apply for this role

Opens Apollo GraphQL's application page

Role

What you'll do.

Senior Platform Engineer at Apollo GraphQL, a leader in GraphQL technology, seeking a distributed systems expert to drive platform engineering excellence. This role focuses on infrastructure automation, Kubernetes service delivery, SLO-driven reliability, and cross-organizational leadership within a high-performing team that values elegant solutions and operational excellence.

Responsibilities

  • Lead Cross-Organizational Platform Initiatives: Take ownership of medium to large-impact subsystems and projects, driving infrastructure modernization efforts across Apollo GraphQL. Execute on strategic platform engineering initiatives that span multiple teams, ensuring successful delivery of scalable distributed systems solutions.
  • Design Infrastructure and Service Delivery Architecture: Create forward-thinking technical designs for Kubernetes-based service delivery, Terraform infrastructure-as-code management, and CI/CD pipelines. Proactively address cost efficiency, security posture, and observability in all architectural decisions using data-driven approaches.
  • Automate Infrastructure Operations and DevOps Workflows: Build and enhance automation frameworks across container orchestration, infrastructure provisioning, and deployment pipelines. Leverage tools like ArgoCD, Atlantis, and CircleCI to eliminate manual processes and reduce operational overhead.
  • Drive Developer Velocity and Operational Excellence: Implement platform capabilities that accelerate developer productivity while maintaining high reliability standards. Utilize DORA metrics and observability frameworks to measure and continuously improve platform performance, reducing deployment friction and incident response times.
  • Deliver Technical Documentation and Design Reviews: Produce comprehensive technical artifacts including design documents, one-pagers, decision records (DRs), and operational runbooks. Participate in design review processes to ensure architectural decisions align with company standards and future scalability requirements.
  • Lead On-Call Operations and Production Support: Participate in on-call rotations and take full ownership of incident resolution, root cause analysis, and prevention. Identify and resolve systemic reliability issues, eliminating noisy monitoring alerts through proactive remediation and infrastructure hardening.
  • Conduct Technical Interviews and Mentor Engineers: Participate in recruiting efforts by conducting technical interviews for platform engineering candidates. Mentor and guide team members on distributed systems concepts, infrastructure best practices, and operational reliability patterns.
  • Collaborate Across Teams on Platform Capabilities: Work with internal stakeholders to understand requirements for logging, monitoring, deployment, and infrastructure services. Build consensus on platform direction and foster a collaborative culture where platform improvements benefit the entire organization.

Qualifications

What we look for.

Technical

  • Distributed Systems Architecture

    Deep expertise designing and operating stateless, fault-tolerant systems with understanding of eventual consistency models, event-driven architectures, and asynchronous patterns. Proven ability to reason about system behavior under failure conditions and design for high availability.

  • Kubernetes Container Orchestration

    Production-grade experience operating and optimizing Kubernetes clusters, including workload management, resource allocation, networking, and troubleshooting. Understanding of declarative infrastructure patterns and cluster operations at scale.

  • Infrastructure-as-Code and Terraform

    Proficiency with Terraform for managing cloud infrastructure across multiple environments. Experience writing modular, maintainable IaC that follows best practices for state management, modularity, and reproducibility.

  • CI/CD Pipeline Design and Automation

    Experience designing and implementing continuous integration and deployment pipelines. Familiarity with GitOps principles, deployment automation tools, and strategies for managing infrastructure and application deployments safely at scale.

  • Cloud Platform Operations

    Strong working knowledge of Google Cloud Platform (GCP), with transferable expertise applicable to AWS or Azure. Understanding of compute, networking, storage services, and cloud-native operational patterns.

  • Observability and Monitoring Architecture

    Experience implementing comprehensive observability solutions including logging, metrics collection, distributed tracing, and alerting. Proficiency with monitoring tools like DataDog and understanding of SLO-driven reliability practices.

  • AI-Assisted Development and Production Intelligence

    Demonstrated ability leveraging agentic tooling and AI systems to enhance daily engineering workflows and detect, diagnose, and mitigate production issues. Experience using automation to improve operational efficiency and system reliability.

Education

  • Bachelor's Degree in Computer Science, Engineering, or Related Field

    Strong foundation in computer science fundamentals, distributed systems theory, and software engineering principles. Equivalent professional experience demonstrating mastery of these core concepts may substitute for formal degree.

Experience

  • Senior-Level Platform or Infrastructure Engineering

    5+ years of progressive experience in platform engineering, site reliability engineering (SRE), or infrastructure operations roles. Demonstrated track record of owning critical systems, leading architectural decisions, and driving organizational platform improvements.

  • Cross-Team Leadership and Collaboration

    Proven ability to work effectively across organizational boundaries, building consensus on platform direction and aligning diverse stakeholder interests. Experience mentoring junior engineers and contributing to technical hiring decisions.

  • Operational Incident Management

    Substantial on-call experience responding to and resolving production incidents. Track record of performing effective root cause analysis, implementing permanent fixes, and eliminating systemic reliability issues through infrastructure improvements.

  • Data-Driven Decision Making

    Strong track record of using metrics, observability data, and analytics to drive technical and business decisions. Experience applying frameworks like DORA to measure and improve engineering effectiveness and operational performance.

  • Technical Design and Documentation

    Experience writing technical designs, architectural decision records, and operational documentation. Ability to communicate complex infrastructure concepts to both technical and non-technical audiences through clear written and verbal communication.

Skills

Required

  • Kubernetes

    Production-grade expertise with container orchestration, cluster operations, workload management, and troubleshooting.

  • Terraform

    Proficiency with infrastructure-as-code for provisioning and managing cloud resources across multiple environments.

  • Google Cloud Platform (GCP)

    Strong working knowledge of GCP services, cloud architecture patterns, and cloud-native operations.

  • Distributed Systems Design

    Deep understanding of distributed computing patterns, eventual consistency, fault tolerance, and event-driven architectures.

  • CI/CD Pipelines

    Experience designing and implementing continuous integration and continuous deployment workflows and automation.

  • System Observability

    Expertise in implementing monitoring, logging, metrics collection, and alerting for complex distributed systems.

  • On-Call Operations

    Production incident response, root cause analysis, and implementation of permanent reliability fixes.

  • Technical Architecture and Design

    Ability to design scalable, maintainable systems and communicate architectural decisions through technical documentation.

Preferred

  • Helm

    Nice to have

    Package management and templating for Kubernetes applications, enabling standardized deployments across environments.

  • ArgoCD

    Nice to have

    GitOps-based continuous delivery tool for Kubernetes, enabling declarative application deployment workflows.

  • Atlantis

    Nice to have

    Infrastructure-as-code workflow tool that enables collaborative infrastructure management and GitOps practices.

  • CircleCI

    Nice to have

    Continuous integration platform for automated testing and deployment pipeline orchestration.

  • DataDog

    Nice to have

    Enterprise observability platform for monitoring, logging, and analytics across distributed systems.

  • Docker

    Nice to have

    Container runtime and containerization best practices for packaging applications and services.

  • GraphQL

    Nice to have

    Experience with GraphQL APIs and understanding of Apollo GraphQL's role in the GraphQL ecosystem.

  • AI/ML-Powered Tooling

    Nice to have

    Experience with agentic systems and AI tools for automating operations, diagnostics, and incident response workflows.

Tech stack

Languages

GoPythonBash/ShellYAML

Frameworks

KubernetesHelmArgoCD

Databases

PostgreSQLCloud Datastore/Firestore

Tools

TerraformDockerArgoCDAtlantisCircleCIDataDogGit/GitHub

Other

Google Cloud Platform (GCP)SLO/SLI/SLA FrameworkDORA MetricsGitOpsObservability and Monitoring

Compensation

Pay and benefits.

Base·USD 165,000 – 195,000

Equity·Stock options

Benefits

  • Equity and Stock Options

    Competitive equity package providing ownership stake in Apollo GraphQL's future success.

  • Health Insurance

    Comprehensive medical, dental, and vision coverage for employee and family.

  • Retirement Planning

    401(k) retirement plan with employer contribution matching.

  • Professional Development

    Education budget for conferences, courses, and professional growth in platform engineering and distributed systems.

  • Flexible Work Arrangement

    Remote-first work environment enabling geographic flexibility and work-life balance.

  • Unlimited Paid Time Off

    Flexible PTO policy enabling engineers to manage personal time and wellness needs.

  • Parental Leave

    Paid leave for new parents supporting family growth and work-life integration.

Process

Interview steps.

  1. 01

    Initial Screening Call

    30-minute conversation with recruiting team to assess background, career goals, and alignment with platform engineering at Apollo GraphQL.

  2. 02

    Technical Deep Dive Interview

    60-90 minute session with platform engineering team members covering distributed systems concepts, Kubernetes architecture, infrastructure design patterns, and real-world problem-solving scenarios.

  3. 03

    System Design Discussion

    Technical interview focusing on designing scalable infrastructure solutions, addressing trade-offs between cost, reliability, and observability. May include infrastructure-as-code design or Kubernetes architecture scenarios.

  4. 04

    Cross-Functional Collaboration Interview

    Conversation with team members outside platform engineering to assess collaboration style, communication ability, and cross-team impact potential.

  5. 05

    Leadership and Culture Fit Interview

    Discussion with senior engineering leadership to explore mentorship philosophy, decision-making approach, and alignment with Apollo's engineering culture emphasizing humility and mindfulness.

  6. 06

    Final Offer Discussion

    Detailed conversation with hiring manager covering role expectations, growth opportunities, and compensation discussion.

Full posting

Original listing.

Are you a talented and driven distributed systems engineer, with a proven track record, capable of leading cross-organizational features? Does tech debt quiver in its boots at the sound of your name? Do you butter your bread with evented systems? Have the letters C, Q, R, and S been ruined for you in an eventually consistent manner? Have we got news for you: you’re not alone!

In this role, you’ll be a key contributor to our Platform Engineering team. You'll get the chance to hone your skills alongside some of the best Platform Engineers (think SRE meets Ops with a heavy focus on automation of everything). Our team’s current focus is on Service Delivery (Kubernetes), Infrastructure Management (Terraform), CI/CD (Argo, Atlantis), and in general all things SLOs and Operational Excellence. Your teammates are talented folks who value code that is 80% of the value for 20% of the work, designs that are forward-thinking enough to be easily flexible for the next features, and leadership with a healthy dose of mindfulness and humility.

What you’ll do

  • You’ll become a key contributor to the team, taking responsibility for the success of some of our subsystems.

  • You will be participating on medium to large impact team initiatives, and within a year be able to execute on such projects.

  • You’ll help with interviewing potential teammates.

  • You’ll create technical designs that proactively address cost efficiency, security, and observability.

  • You’ll deliver technical plans, one-pagers, DRs, and other artifacts.

  • You’ll work with Kubernetes, GCP, Helm, Terraform, DataDog, ArgoCD, CircleCI, Atlantis, Docker (the list goes on) to deliver your work.

  • You’ll be responsible for improving developer velocity across the company (leveraging frameworks like DORA) and hardening our reliability and observability.

  • You’ll participate in on-call rotations and help keep all of Apollo afloat

    • You’ll be fully empowered to fix the root cause of issues, and ruthlessly quell any noisy monitors

Who you are

A description of our ideal candidate.

Minimum requirements

  • You have systems expertise and experience with stateless/fault tolerant systems, as well as familiarity with eventing patterns and distributed paradigms.

  • You’ve leveraged agentic tooling to not only enhance your day to day work, but detect and mitigate issues in production

  • You think about weighing technical and business trade-offs and are working at “seeing down the road and around corners.”

  • You enjoy cross-team collaboration and believe in a “rising tides lifts all boats” mentality. You’re a joy to those around you, bringing everyone along for the ride.

  • You are data-oriented and have leveraged data to make decisions in a concrete way.

  • You have familiarity with Kubernetes, GCP (preferably, but AWS or Azure is great too!), Terraform

Nice to have

  • Bonus points for Helm, Docker, ArgoCD, CircleCI, DataDog, and of course GraphQL!

Redirects to Apollo GraphQL's application page.

Other roles

More at Apollo GraphQL.