Systems Integration Engineer, Build Systems | Consumer Devices
DevOps Engineer · Senior · Full Time
Opens OpenAI's application page
Role
What you'll do.
Join OpenAI's Systems Integration team to design and maintain the build infrastructure and CI systems that power consumer device software deployments. This senior engineering role focuses on Bazel-based builds, Buildkite pipelines, and distributed infrastructure while enabling thousands of developers to ship products rapidly with confidence. You'll own critical developer productivity systems including build caching, test frameworks, and CI automation in a fast-growing organization.
Responsibilities
- Own and Evolve Build Workflows: Architect and maintain Bazel and Yocto-based build and test workflows in a polyrepo environment, ensuring hermetic, reproducible builds that teams can confidently adopt across the organization.
- Design and Maintain Build Rules: Create and iterate on Starlark rules, macros, toolchains, and integrations that enhance build reliability and reproducibility while simplifying team adoption and reducing configuration overhead.
- Optimize CI Performance and Reliability: Improve Buildkite pipeline performance across queue time, build time, cache hit rates, retry behavior, and flake isolation to maximize developer productivity and reduce time-to-feedback.
- Implement Intelligent Build Optimization: Build systems that reduce unnecessary CI work through affected-target detection, dependency graph analysis, test selection optimization, smart caching, batching, and intelligent scheduling.
- Unify Development Workflows: Design systems that allow engineers to reproduce CI behavior locally, debug build failures efficiently, and iterate quickly without requiring deep expertise in the entire build infrastructure stack.
- Operate Build Infrastructure: Manage and optimize build infrastructure across Docker/OCI images, Kubernetes-based runners, cloud resources, remote cache/execution systems, and factory hardware to ensure reliability and cost efficiency.
- Instrument Systems with Observability: Build comprehensive metrics, logging, tracing, dashboards, and analytics across build and CI systems to measure speed, reliability, cost, and quantify developer impact of infrastructure improvements.
- Partner with Engineering Teams: Work directly with developers to understand pain points, onboard projects onto new systems, debug complex build issues, and identify and eliminate systemic infrastructure bottlenecks.
- Pioneer AI-Driven Developer Tools: Apply modern AI tools to innovate in CI failure analysis, flaky test debugging, PR triage, automated remediation, and create developer-facing intelligence that accelerates decision-making.
- Ensure Production Reliability: Own the reliability of critical developer and factory infrastructure systems through active participation in on-call rotations and rapid incident response to maintain high SLA standards.
Qualifications
What we look for.
Technical
Advanced Bazel Expertise
5+ years hands-on experience with Bazel, Buck, Gradle, or equivalent build systems with deep understanding of hermetic builds, dependency graphs, caching strategies, sandboxing mechanisms, and remote execution architectures.
Large-Scale CI System Design
Proven track record building and operating CI systems at scale in environments where build time, queue time, test flakiness, and developer trust directly impact engineering velocity and product delivery.
Distributed Infrastructure Debugging
Demonstrated ability to debug complex distributed build and CI failures across source control systems, dependency management, containerization, runners, remote caches, test frameworks, and service infrastructure.
Starlark and Build Configuration
Proficiency writing and maintaining Starlark code including rules, macros, toolchains, and plugins to create reproducible, maintainable build configurations across polyrepo environments.
Kubernetes and Container Orchestration
Production experience managing Kubernetes clusters for CI runners and build infrastructure including resource optimization, scheduling, networking, and reliability.
Infrastructure as Code
Hands-on experience with Terraform, Bazel, or equivalent IaC tools to define, version, and manage build infrastructure in reproducible and auditable ways.
Education
Computer Science or Related Field
Bachelor's degree in Computer Science, Computer Engineering, or equivalent practical experience demonstrating foundational software engineering and systems design knowledge.
Experience
Senior Infrastructure Engineering
5+ years of infrastructure and tooling engineering experience with demonstrated impact on developer productivity at scale in organizations managing large codebases.
Polyrepo Environment Experience
Production experience operating in polyrepo environments with source code from multiple parties, managing coordination and consistency across independent build and deployment pipelines.
Production Systems Ownership
Comfortable owning production software with strong SLA requirements, understanding trade-offs between velocity and reliability, and maintaining systems under high-stakes usage patterns.
Developer Experience Focus
Track record identifying and eliminating sources of friction that slow engineering teams, reduce operational toil, and measurably improve developer satisfaction and velocity metrics.
Skills
Required
Bazel Build Systems
Expert-level proficiency with Bazel including build file authoring, rule creation, workspace configuration, and optimization for large-scale monorepo and polyrepo environments.
CI/CD Pipeline Design
Advanced experience designing, implementing, and scaling CI/CD pipelines with focus on build performance, test reliability, and developer feedback loops.
Kubernetes Administration
Production-grade Kubernetes skills including cluster management, workload orchestration, resource optimization, and troubleshooting in high-volume environments.
Python or Go Programming
Proficiency writing infrastructure automation, build tooling, and DevOps utilities in Python, Go, or equivalent systems languages.
Distributed Systems Debugging
Expert troubleshooting of complex distributed systems issues involving multiple layers of infrastructure, networking, caching, and asynchronous execution.
Linux Systems Administration
Deep Linux expertise including kernel concepts, system performance tuning, containerization, and troubleshooting at scale.
Preferred
Buildkite CI/CD Platform
Nice to haveProduction experience with Buildkite including pipeline configuration, agent management, and optimization for high-throughput CI environments.
Yocto Build System
Nice to haveExperience with Yocto project or BitBake for embedded Linux builds, particularly in polyrepo or multi-target build environments.
Remote Execution Systems
Nice to haveFamiliarity with remote build execution platforms like Bazel Remote Execution, BuildBarn, or equivalent systems for distributed build optimization.
Hardware-in-the-Loop Testing
Nice to haveExperience designing or operating hardware-in-the-loop testing infrastructure for embedded systems, device firmware, or hardware-software integration validation.
AI/ML Application to Infrastructure
Nice to haveInterest and experience applying machine learning techniques to infrastructure problems like failure prediction, flaky test detection, or intelligent resource scheduling.
Rust or C++ Development
Nice to haveExperience building systems-level tools in Rust or C++ for performance-critical infrastructure components like build orchestration or artifact caching.
Docker and OCI Standards
Nice to haveAdvanced proficiency with Docker, container image optimization, OCI standards, and container supply chain security relevant to build infrastructure.
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 293,000 – 325,000
Equity·Stock options
Benefits
Equity and Stock Options
Competitive stock option package as part of OpenAI's equity compensation structure, with vesting schedules aligned to long-term value creation.
Comprehensive Health Coverage
Medical, dental, and vision insurance with competitive premiums and coverage levels for employees and eligible dependents.
Retirement Benefits
401(k) plan with employer matching contributions to support long-term financial planning and retirement security.
Flexible Time Off
Generous paid time off policy including vacation days, sick leave, and personal days to maintain work-life balance.
Relocation Assistance
Comprehensive relocation support for new employees joining the San Francisco office, including moving allowances and transition assistance.
Hybrid Work Model
Four days per week in-office with flexible work arrangement supporting productive collaboration while maintaining schedule flexibility.
Professional Development
Learning opportunities, conference attendance, and training budgets to support career growth and technical skill advancement.
Mental Health Support
Employee assistance programs, counseling services, and mental health resources supporting employee wellness.
Process
Interview steps.
- 01
Resume and Application Review
Initial screening of your professional background focusing on build systems experience, infrastructure engineering breadth, and relevant technology expertise with Bazel and CI systems.
- 02
Technical Phone Screen
Conversation with a hiring engineer covering your hands-on experience with Bazel, CI/CD systems, distributed infrastructure design, and approach to solving build system challenges.
- 03
Deep-Dive Technical Interview
Extended technical discussion with senior team members exploring specific projects, architectural decisions you've made, debugging approaches for complex infrastructure issues, and trade-offs between build speed and reliability.
- 04
System Design Conversation
Collaborative session designing a build or CI system from first principles, discussing scalability, reliability, developer experience, and how you'd approach evolving existing infrastructure.
- 05
Infrastructure and Operations Discussion
Conversation focused on production systems ownership, on-call experience, incident response, observability, and your philosophy on maintaining reliability at scale.
- 06
Cross-Functional Stakeholder Interviews
Meetings with engineering teams who depend on build infrastructure to understand collaboration style, communication skills, and developer experience priorities.
- 07
Leadership and Culture Fit
Final conversation with hiring manager and leadership covering career goals, OpenAI's mission, team dynamics, and alignment with organizational values around safety and impact.
Full posting
Original listing.
About the Team
The Systems Integration team is responsible for building the infrastructure, tooling, and validation systems that ensure our device software our device software is reliable, testable, and ready to ship. We design and maintain build systems, CI pipelines, automated test frameworks, and hardware-in-the-loop labs to enable rapid, safe product launches. Our work spans build systems, developer tools, systems integration, and cross-team collaboration to ensure developers can build reliably and ship with confidence.
About the Role
We are looking for an engineer to help evolve OpenAI’s Consumer Products build and continuous integration systems for a fast-growing engineering organization. This role sits at the intersection of developer productivity, build systems, distributed infrastructure, software quality, and on-device software. You will work on the systems that determine how quickly and confident engineers can move: Bazel-bazed builds, Buildkite pipelines, test coverage, remote caching and execution, CI observability, and tooling that helps engineers understand and fix failures quickly.
Our mission is to enable OpenAI to ship software running on consumer devices rapidly with a high bar for correctness, reliability, and safety. The best version of this work is invisible when it succeeds: builds are fast, tests are trusted, CI failures are understandable, and engineers can focus on shipping products instead of fighting infrastructure.
This role is based in San Francisco, CA. We use a hybrid work model of four days in the office per week and offer relocation assistance to new employees.
In This Role, You Will
Own and evolve Bazel and yocto-based build and test workflows in a polyrepo environment
Design and maintain Starlark rules, macros, toolchains, and integrations that make builds hermetic, reproducible, and easy for teams to adopt
Improve CI performance and reliability across Buildkite pipelines, including queue time, build time, cache hit rates, retry behavior, and flake isolation.
Build systems that reduce unnecessary CI work through affected-target detection, dependency graph analysis, test selection, caching, batching, and smarter scheduling
Unify local and CI development workflows so engineers can reproduce CI behavior, debug build failures, and iterate quickly without learning every detail of the build stack
Operate and optimize build infrastructure across Docker/OCI images, Kubernetes-based runners, cloud resources, and remote cache/execution systems
Instrument build and CI systems with metrics, logs, traces, dashboards, and analytics so we can measure speed, reliability, cost, and developer impact
Partner directly with users to understand pain points, onboard projects, debug hard build issues, and remove systemic bottlenecks
Use modern AI tools to pioneer novel CI failure analysis, flaky test debugging, PR triage, automated remediation, and developer-facing information
Own the reliability of the systems you build, including participating in an on-call rotation for critical developer and factory infrastructure
Technologies Commonly Used In This Environment Include
Bazel and Starlark for build and test workflows
Buildkite for CI orchestration
Docker and OCI images for build and runtime packaging
Kubernetes for CI runners and infrastructure orchestration
Python, Go, TypeScript, Rust, C++, and other languages in a large monorepo
Terraform for infrastructure as code
Remote caching, remote execution, artifact storage, and build telemetry systems
Minimum Qualifications:
5+ years of engineering experience, including significant experience building infrastructure and tooling for developers
Hands-on experience with Bazel, Buck, Gradle, or similar build systems, and understand the trade offs of hermetic builds, dependency graphs, caching, sandboxing, and remote execution
Have built CI systems at scale, especially in environments where build time, queue time, test flakiness, and developer trust materially affect engineering velocity
You May Be A Strong Fit If You
Are comfortable with owning production software and working in environments with strong SLA requirements
Can debug distributed build and CI failures across source control, dependency management, containers, runners, remote caches, test frameworks, and service infrastructure.
Care deeply about developer experience and have empty for the small sources of friction that slow teams down or create operational toil
Are excited to apply AI to developer infrastructure in ways that increase team velocity without weakening quality, reliability, or safety
Have operated in a polyrepo environment with source code from multiple parties
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity.
We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.
For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement.
Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.
To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance.
We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link.
OpenAI Global Applicant Privacy Policy
At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.
Redirects to OpenAI's application page.
Other roles
More at OpenAI.
Software Security Architect, Operating Systems | Consumer Devices
Senior
Software Engineer, API Safety
Senior
Android Systems Engineer, Consumer Devices
Senior
Data Engineer, Monetization Data Platform
Senior
Software Engineer, Plugin Developer Platform
Senior