Senior Platform Engineer
Platform Engineer · Senior · Full Time · Remote
Opens Apollo GraphQL's application page
Role
What you'll do.
Senior Platform Engineer at Apollo GraphQL, a leader in GraphQL technology, seeking a distributed systems expert to drive platform engineering excellence. This role focuses on infrastructure automation, Kubernetes service delivery, SLO-driven reliability, and cross-organizational leadership within a high-performing team that values elegant solutions and operational excellence.
Responsibilities
- Lead Cross-Organizational Platform Initiatives: Take ownership of medium to large-impact subsystems and projects, driving infrastructure modernization efforts across Apollo GraphQL. Execute on strategic platform engineering initiatives that span multiple teams, ensuring successful delivery of scalable distributed systems solutions.
- Design Infrastructure and Service Delivery Architecture: Create forward-thinking technical designs for Kubernetes-based service delivery, Terraform infrastructure-as-code management, and CI/CD pipelines. Proactively address cost efficiency, security posture, and observability in all architectural decisions using data-driven approaches.
- Automate Infrastructure Operations and DevOps Workflows: Build and enhance automation frameworks across container orchestration, infrastructure provisioning, and deployment pipelines. Leverage tools like ArgoCD, Atlantis, and CircleCI to eliminate manual processes and reduce operational overhead.
- Drive Developer Velocity and Operational Excellence: Implement platform capabilities that accelerate developer productivity while maintaining high reliability standards. Utilize DORA metrics and observability frameworks to measure and continuously improve platform performance, reducing deployment friction and incident response times.
- Deliver Technical Documentation and Design Reviews: Produce comprehensive technical artifacts including design documents, one-pagers, decision records (DRs), and operational runbooks. Participate in design review processes to ensure architectural decisions align with company standards and future scalability requirements.
- Lead On-Call Operations and Production Support: Participate in on-call rotations and take full ownership of incident resolution, root cause analysis, and prevention. Identify and resolve systemic reliability issues, eliminating noisy monitoring alerts through proactive remediation and infrastructure hardening.
- Conduct Technical Interviews and Mentor Engineers: Participate in recruiting efforts by conducting technical interviews for platform engineering candidates. Mentor and guide team members on distributed systems concepts, infrastructure best practices, and operational reliability patterns.
- Collaborate Across Teams on Platform Capabilities: Work with internal stakeholders to understand requirements for logging, monitoring, deployment, and infrastructure services. Build consensus on platform direction and foster a collaborative culture where platform improvements benefit the entire organization.
Qualifications
What we look for.
Technical
Distributed Systems Architecture
Deep expertise designing and operating stateless, fault-tolerant systems with understanding of eventual consistency models, event-driven architectures, and asynchronous patterns. Proven ability to reason about system behavior under failure conditions and design for high availability.
Kubernetes Container Orchestration
Production-grade experience operating and optimizing Kubernetes clusters, including workload management, resource allocation, networking, and troubleshooting. Understanding of declarative infrastructure patterns and cluster operations at scale.
Infrastructure-as-Code and Terraform
Proficiency with Terraform for managing cloud infrastructure across multiple environments. Experience writing modular, maintainable IaC that follows best practices for state management, modularity, and reproducibility.
CI/CD Pipeline Design and Automation
Experience designing and implementing continuous integration and deployment pipelines. Familiarity with GitOps principles, deployment automation tools, and strategies for managing infrastructure and application deployments safely at scale.
Cloud Platform Operations
Strong working knowledge of Google Cloud Platform (GCP), with transferable expertise applicable to AWS or Azure. Understanding of compute, networking, storage services, and cloud-native operational patterns.
Observability and Monitoring Architecture
Experience implementing comprehensive observability solutions including logging, metrics collection, distributed tracing, and alerting. Proficiency with monitoring tools like DataDog and understanding of SLO-driven reliability practices.
AI-Assisted Development and Production Intelligence
Demonstrated ability leveraging agentic tooling and AI systems to enhance daily engineering workflows and detect, diagnose, and mitigate production issues. Experience using automation to improve operational efficiency and system reliability.
Education
Bachelor's Degree in Computer Science, Engineering, or Related Field
Strong foundation in computer science fundamentals, distributed systems theory, and software engineering principles. Equivalent professional experience demonstrating mastery of these core concepts may substitute for formal degree.
Experience
Senior-Level Platform or Infrastructure Engineering
5+ years of progressive experience in platform engineering, site reliability engineering (SRE), or infrastructure operations roles. Demonstrated track record of owning critical systems, leading architectural decisions, and driving organizational platform improvements.
Cross-Team Leadership and Collaboration
Proven ability to work effectively across organizational boundaries, building consensus on platform direction and aligning diverse stakeholder interests. Experience mentoring junior engineers and contributing to technical hiring decisions.
Operational Incident Management
Substantial on-call experience responding to and resolving production incidents. Track record of performing effective root cause analysis, implementing permanent fixes, and eliminating systemic reliability issues through infrastructure improvements.
Data-Driven Decision Making
Strong track record of using metrics, observability data, and analytics to drive technical and business decisions. Experience applying frameworks like DORA to measure and improve engineering effectiveness and operational performance.
Technical Design and Documentation
Experience writing technical designs, architectural decision records, and operational documentation. Ability to communicate complex infrastructure concepts to both technical and non-technical audiences through clear written and verbal communication.
Skills
Required
Kubernetes
Production-grade expertise with container orchestration, cluster operations, workload management, and troubleshooting.
Terraform
Proficiency with infrastructure-as-code for provisioning and managing cloud resources across multiple environments.
Google Cloud Platform (GCP)
Strong working knowledge of GCP services, cloud architecture patterns, and cloud-native operations.
Distributed Systems Design
Deep understanding of distributed computing patterns, eventual consistency, fault tolerance, and event-driven architectures.
CI/CD Pipelines
Experience designing and implementing continuous integration and continuous deployment workflows and automation.
System Observability
Expertise in implementing monitoring, logging, metrics collection, and alerting for complex distributed systems.
On-Call Operations
Production incident response, root cause analysis, and implementation of permanent reliability fixes.
Technical Architecture and Design
Ability to design scalable, maintainable systems and communicate architectural decisions through technical documentation.
Preferred
Helm
Nice to havePackage management and templating for Kubernetes applications, enabling standardized deployments across environments.
ArgoCD
Nice to haveGitOps-based continuous delivery tool for Kubernetes, enabling declarative application deployment workflows.
Atlantis
Nice to haveInfrastructure-as-code workflow tool that enables collaborative infrastructure management and GitOps practices.
CircleCI
Nice to haveContinuous integration platform for automated testing and deployment pipeline orchestration.
DataDog
Nice to haveEnterprise observability platform for monitoring, logging, and analytics across distributed systems.
Docker
Nice to haveContainer runtime and containerization best practices for packaging applications and services.
GraphQL
Nice to haveExperience with GraphQL APIs and understanding of Apollo GraphQL's role in the GraphQL ecosystem.
AI/ML-Powered Tooling
Nice to haveExperience with agentic systems and AI tools for automating operations, diagnostics, and incident response workflows.
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 165,000 – 195,000
Equity·Stock options
Benefits
Equity and Stock Options
Competitive equity package providing ownership stake in Apollo GraphQL's future success.
Health Insurance
Comprehensive medical, dental, and vision coverage for employee and family.
Retirement Planning
401(k) retirement plan with employer contribution matching.
Professional Development
Education budget for conferences, courses, and professional growth in platform engineering and distributed systems.
Flexible Work Arrangement
Remote-first work environment enabling geographic flexibility and work-life balance.
Unlimited Paid Time Off
Flexible PTO policy enabling engineers to manage personal time and wellness needs.
Parental Leave
Paid leave for new parents supporting family growth and work-life integration.
Process
Interview steps.
- 01
Initial Screening Call
30-minute conversation with recruiting team to assess background, career goals, and alignment with platform engineering at Apollo GraphQL.
- 02
Technical Deep Dive Interview
60-90 minute session with platform engineering team members covering distributed systems concepts, Kubernetes architecture, infrastructure design patterns, and real-world problem-solving scenarios.
- 03
System Design Discussion
Technical interview focusing on designing scalable infrastructure solutions, addressing trade-offs between cost, reliability, and observability. May include infrastructure-as-code design or Kubernetes architecture scenarios.
- 04
Cross-Functional Collaboration Interview
Conversation with team members outside platform engineering to assess collaboration style, communication ability, and cross-team impact potential.
- 05
Leadership and Culture Fit Interview
Discussion with senior engineering leadership to explore mentorship philosophy, decision-making approach, and alignment with Apollo's engineering culture emphasizing humility and mindfulness.
- 06
Final Offer Discussion
Detailed conversation with hiring manager covering role expectations, growth opportunities, and compensation discussion.
Full posting
Original listing.
Are you a talented and driven distributed systems engineer, with a proven track record, capable of leading cross-organizational features? Does tech debt quiver in its boots at the sound of your name? Do you butter your bread with evented systems? Have the letters C, Q, R, and S been ruined for you in an eventually consistent manner? Have we got news for you: you’re not alone!
In this role, you’ll be a key contributor to our Platform Engineering team. You'll get the chance to hone your skills alongside some of the best Platform Engineers (think SRE meets Ops with a heavy focus on automation of everything). Our team’s current focus is on Service Delivery (Kubernetes), Infrastructure Management (Terraform), CI/CD (Argo, Atlantis), and in general all things SLOs and Operational Excellence. Your teammates are talented folks who value code that is 80% of the value for 20% of the work, designs that are forward-thinking enough to be easily flexible for the next features, and leadership with a healthy dose of mindfulness and humility.
What you’ll do
You’ll become a key contributor to the team, taking responsibility for the success of some of our subsystems.
You will be participating on medium to large impact team initiatives, and within a year be able to execute on such projects.
You’ll help with interviewing potential teammates.
You’ll create technical designs that proactively address cost efficiency, security, and observability.
You’ll deliver technical plans, one-pagers, DRs, and other artifacts.
You’ll work with Kubernetes, GCP, Helm, Terraform, DataDog, ArgoCD, CircleCI, Atlantis, Docker (the list goes on) to deliver your work.
You’ll be responsible for improving developer velocity across the company (leveraging frameworks like DORA) and hardening our reliability and observability.
You’ll participate in on-call rotations and help keep all of Apollo afloat
You’ll be fully empowered to fix the root cause of issues, and ruthlessly quell any noisy monitors
Who you are
A description of our ideal candidate.
Minimum requirements
You have systems expertise and experience with stateless/fault tolerant systems, as well as familiarity with eventing patterns and distributed paradigms.
You’ve leveraged agentic tooling to not only enhance your day to day work, but detect and mitigate issues in production
You think about weighing technical and business trade-offs and are working at “seeing down the road and around corners.”
You enjoy cross-team collaboration and believe in a “rising tides lifts all boats” mentality. You’re a joy to those around you, bringing everyone along for the ride.
You are data-oriented and have leveraged data to make decisions in a concrete way.
You have familiarity with Kubernetes, GCP (preferably, but AWS or Azure is great too!), Terraform
Nice to have
Bonus points for Helm, Docker, ArgoCD, CircleCI, DataDog, and of course GraphQL!
Redirects to Apollo GraphQL's application page.
Other roles