Systems Integration Engineer, Build Systems | Consumer Devices

DevOps Engineer · Senior · Full Time

San FranciscoUSD 293k – 325k1mo ago
Apply for this role

Opens OpenAI's application page

Role

What you'll do.

As a Systems Integration Engineer at OpenAI's Consumer Devices team, you'll architect and maintain enterprise-scale build systems and CI pipelines that power rapid, reliable software delivery for consumer-facing AI products. You'll design Bazel-based build workflows, optimize Buildkite CI infrastructure, and develop developer tooling that increases engineering velocity while maintaining safety and correctness standards. This role requires 5+ years of infrastructure engineering experience with deep expertise in build systems, CI/CD orchestration, and distributed systems optimization.

Responsibilities

  • Build System Architecture and Maintenance: Own and evolve Bazel and Yocto-based build and test workflows in a polyrepo environment, designing Starlark rules, macros, and toolchains that ensure builds are hermetic, reproducible, and easily adoptable across teams. Implement dependency graph analysis and affected-target detection to optimize build performance.
  • CI/CD Pipeline Optimization: Design and maintain Buildkite pipelines focusing on queue time reduction, build acceleration, cache hit rate optimization, flake isolation, and intelligent retry behavior. Implement systems that minimize unnecessary CI work through advanced scheduling and batching algorithms.
  • Local and Remote Development Workflow Unification: Unify local development environments with CI workflows so engineers can reproduce CI behavior, debug build failures efficiently, and iterate rapidly without deep knowledge of the entire build stack. Reduce friction points that create developer toil.
  • Build Infrastructure Operations: Operate and optimize containerized build infrastructure across Docker/OCI images, Kubernetes-based CI runners, cloud resources, and distributed remote caching and execution systems. Manage scalability, cost efficiency, and resource utilization.
  • Observability and Instrumentation: Instrument build and CI systems with comprehensive metrics, structured logging, distributed tracing, dashboards, and analytics to measure performance indicators including speed, reliability, cost impact, and developer productivity metrics.
  • Stakeholder Collaboration and Onboarding: Partner directly with engineering teams to understand pain points, onboard projects onto the build platform, debug complex build failures, and identify and eliminate systemic infrastructure bottlenecks that impede engineering velocity.
  • AI-Powered Developer Infrastructure Innovation: Pioneer applications of modern AI tools for automated CI failure analysis, flaky test debugging, pull request triage, root cause identification, and intelligent remediation to enhance developer experience and reduce manual troubleshooting.
  • Production Reliability and On-Call Support: Own the reliability and availability of build and CI systems you architect, including participating in an on-call rotation for critical developer and manufacturing infrastructure, responding to incidents, and implementing preventative improvements.

Qualifications

What we look for.

Technical

  • Build System Expertise

    Hands-on production experience with Bazel, Buck, Gradle, or equivalent modern build systems. Deep understanding of hermetic builds, dependency graph resolution, sandboxing, remote execution, and the architectural trade-offs between local and distributed builds.

  • CI/CD Infrastructure at Scale

    Proven experience building and maintaining CI systems in environments where build time, queue time, test flakiness, and developer confidence materially impact engineering velocity. Experience optimizing CI infrastructure for speed, reliability, and cost.

  • Distributed Systems Debugging

    Demonstrated ability to debug complex failures across distributed systems including source control integration, dependency management, container runtime environments, orchestration platforms, remote caching systems, test frameworks, and microservice infrastructure.

  • Infrastructure as Code and Automation

    Proficiency with Infrastructure as Code tools such as Terraform or equivalent, container technologies (Docker, OCI), orchestration platforms (Kubernetes), and scripting languages for systems automation and infrastructure provisioning.

  • Polyrepo Management

    Production experience managing source code and build infrastructure across polyrepo environments with multiple independent repositories, cross-repository dependencies, and coordinated releases from disparate code sources.

  • Production SLA Management

    Comfort owning production software systems with strong SLA requirements, including incident response, on-call rotations, reliability engineering, and maintaining high availability standards for critical developer and manufacturing infrastructure.

  • Systems Observability

    Experience implementing comprehensive observability systems including metrics collection, distributed tracing, log aggregation, dashboard design, and analytics to identify performance bottlenecks and inform optimization priorities.

Education

  • Computer Science or Software Engineering Foundation

    Bachelor's degree in Computer Science, Software Engineering, or equivalent professional engineering experience demonstrating deep technical knowledge of software systems and infrastructure.

Experience

  • Infrastructure and Platform Engineering

    5+ years of professional engineering experience with significant focus on infrastructure, developer platforms, and tooling. Demonstrated track record of building systems that measurably improve developer productivity and engineering velocity.

  • Build Systems Operations

    Substantial production experience operating large-scale build systems at companies with complex monorepo or polyrepo architectures, high developer throughput, and strict reliability requirements.

  • DevOps and Reliability Engineering

    Professional experience managing production systems, participating in on-call rotations, troubleshooting distributed system failures, and implementing reliability improvements in high-stakes environments.

  • Developer Experience Focus

    Demonstrated empathy for developer experience with a track record of identifying small sources of friction that create operational toil, and shipping solutions that measurably improve workflow efficiency and team satisfaction.

Skills

Required

  • Bazel or Equivalent Build System

    Production-level proficiency with Bazel and Starlark, or deep expertise with alternative modern build systems like Buck, Gradle, or Pants. Understanding of build caching, remote execution, sandboxing, and hermetic build principles.

  • Buildkite or CI/CD Platform

    Hands-on experience designing, implementing, and optimizing CI/CD pipelines using Buildkite or equivalent platforms. Knowledge of pipeline orchestration, queue management, and failure handling at scale.

  • Kubernetes and Container Orchestration

    Operational experience with Kubernetes clusters managing containerized CI runners, infrastructure scaling, resource management, and troubleshooting container runtime issues in production environments.

  • Python and/or Go

    Proficiency in Python or Go for infrastructure automation, tooling development, and systems scripting. Ability to write maintainable, testable infrastructure code at scale.

  • Docker and Container Technologies

    Expert-level knowledge of Docker, image building, OCI standards, container security, and optimizing container images for build environments and runtime efficiency.

  • Terraform and Infrastructure as Code

    Production experience using Terraform to define, provision, and manage cloud infrastructure. Understanding of infrastructure automation, state management, and managing complex infrastructure dependencies.

  • Systems Debugging and Troubleshooting

    Strong ability to diagnose failures across complex distributed systems using logging, metrics, tracing, and systematic debugging methodologies. Experience with debugging tools and observability platforms.

Preferred

  • Yocto or Embedded Build Systems

    Nice to have

    Experience with Yocto, BitBake, or other embedded Linux build systems relevant to consumer device software. Understanding of cross-compilation, hardware integration testing, and device-specific build optimization.

  • Remote Caching and Execution

    Nice to have

    Experience implementing or optimizing distributed build systems with remote caching backends (Bazel Remote Caching, BuildBarn) and remote execution platforms. Knowledge of cache invalidation and distributed execution challenges.

  • Test Flakiness Mitigation

    Nice to have

    Track record of identifying root causes of flaky tests, implementing flake detection systems, and reducing test unreliability in CI environments. Experience with test isolation and deterministic test execution.

  • AI/ML Application to Infrastructure

    Nice to have

    Interest in and experience applying machine learning or large language models to infrastructure problems such as CI failure analysis, automated remediation, intelligent test selection, or predictive performance optimization.

  • TypeScript, Rust, or C++

    Nice to have

    Familiarity with TypeScript, Rust, or C++ in large monorepo environments. Understanding of polyglot build challenges and cross-language dependency management.

  • Manufacturing or Hardware Testing

    Nice to have

    Experience with hardware-in-the-loop testing infrastructure, factory automation systems, or manufacturing CI/CD pipelines relevant to consumer device production.

  • Developer Productivity Metrics

    Nice to have

    Experience measuring and optimizing developer productivity through instrumentation, A/B testing infrastructure changes, and quantifying the impact of build system improvements on engineering velocity.

  • Open Source Build Tools

    Nice to have

    Contribution to or deep knowledge of open source build system projects, CI platforms, or developer infrastructure tools. Understanding of the broader build systems ecosystem and emerging best practices.

Tech stack

Languages

PythonGoStarlarkTypeScriptRustC++

Frameworks

BazelBuildkiteKubernetesYocto

Databases

Distributed Cache SystemsTime-Series Metrics DatabasesLog Aggregation Platforms

Tools

DockerTerraformGit and Version ControlDistributed Tracing ToolsMetrics and Monitoring

Other

OCI Image SpecificationRemote Execution ProtocolHardware-in-the-Loop TestingAI/ML Infrastructure Tools

Compensation

Pay and benefits.

Base·USD 293,000 – 325,000

Equity·Stock options

Full posting

Original listing.

About the Team

The Systems Integration team is responsible for building the infrastructure, tooling, and validation systems that ensure our device software our device software is reliable, testable, and ready to ship. We design and maintain build systems, CI pipelines, automated test frameworks, and hardware-in-the-loop labs to enable rapid, safe product launches. Our work spans build systems, developer tools, systems integration, and cross-team collaboration to ensure developers can build reliably and ship with confidence.

About the Role

We are looking for an engineer to help evolve OpenAI’s Consumer Products build and continuous integration systems for a fast-growing engineering organization. This role sits at the intersection of developer productivity, build systems, distributed infrastructure, software quality, and on-device software. You will work on the systems that determine how quickly and confident engineers can move: Bazel-bazed builds, Buildkite pipelines, test coverage, remote caching and execution, CI observability, and tooling that helps engineers understand and fix failures quickly.

Our mission is to enable OpenAI to ship software running on consumer devices rapidly with a high bar for correctness, reliability, and safety. The best version of this work is invisible when it succeeds: builds are fast, tests are trusted, CI failures are understandable, and engineers can focus on shipping products instead of fighting infrastructure.

This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees.

In This Role, You Will

  • Own and evolve Bazel and yocto-based build and test workflows in a polyrepo environment

  • Design and maintain Starlark rules, macros, toolchains, and integrations that make builds hermetic, reproducible, and easy for teams to adopt

  • Improve CI performance and reliability across Buildkite pipelines, including queue time, build time, cache hit rates, retry behavior, and flake isolation.

  • Build systems that reduce unnecessary CI work through affected-target detection, dependency graph analysis, test selection, caching, batching, and smarter scheduling

  • Unify local and CI development workflows so engineers can reproduce CI behavior, debug build failures, and iterate quickly without learning every detail of the build stack

  • Operate and optimize build infrastructure across Docker/OCI images, Kubernetes-based runners, cloud resources, and remote cache/execution systems

  • Instrument build and CI systems with metrics, logs, traces, dashboards, and analytics so we can measure speed, reliability, cost, and developer impact

  • Partner directly with users to understand pain points, onboard projects, debug hard build issues, and remove systemic bottlenecks

  • Use modern AI tools to pioneer novel CI failure analysis, flaky test debugging, PR triage, automated remediation, and developer-facing information

  • Own the reliability of the systems you build, including participating in an on-call rotation for critical developer and factory infrastructure

Technologies Commonly Used In This Environment Include

  • Bazel and Starlark for build and test workflows

  • Buildkite for CI orchestration

  • Docker and OCI images for build and runtime packaging

  • Kubernetes for CI runners and infrastructure orchestration

  • Python, Go, TypeScript, Rust, C++, and other languages in a large monorepo

  • Terraform for infrastructure as code

  • Remote caching, remote execution, artifact storage, and build telemetry systems

You May Be A Strong Fit If You

  • Have 5+ years of engineering experience, including significant experience building infrastructure and tooling for developers

  • Have hands-on experience with Bazel, Buck, Gradle, or similar build systems, and understand the trade offs of hermetic builds, dependency graphs, caching, sandboxing, and remote execution

  • Have built CI systems at scale, especially in environments where build time, queue time, test flakiness, and developer trust materially affect engineering velocity

  • Are comfortable with owning production software and working in environments with strong SLA requirements

  • Can debug distributed build and CI failures across source control, dependency management, containers, runners, remote caches, test frameworks, and service infrastructure.

  • Care deeply about developer experience and have empty for the small sources of friction that slow teams down or create operational toil

  • Are excited to apply AI to developer infrastructure in ways that increase team velocity without weakening quality, reliability, or safety

  • Have operated in a polyrepo environment with source code from multiple parties

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. 

We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.

For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement.

Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.

To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance.

We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link.

OpenAI Global Applicant Privacy Policy

At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.

Redirects to OpenAI's application page.

Other roles

More at OpenAI.

View all 111 roles