Software Engineer, Infrastructure
Infrastructure Engineer · Senior · Full Time
Opens OpenAI's application page
Role
What you'll do.
Join OpenAI's Infrastructure organization as a Software Engineer to design, build, and maintain mission-critical distributed systems powering ChatGPT and the OpenAI API. You'll collaborate with high-impact teams across Core Distributed Systems, Reliability Engineering, Observability, Developer Productivity, Cloud Infrastructure, and Databases, working on scalable, fault-tolerant systems that support cutting-edge AI research and products. This role requires 4+ years of industry experience with 2+ years leading complex projects, deep expertise in distributed systems architecture, and proficiency in languages like Python, Go, C++, or Rust.
Responsibilities
- Design and Build Reliable Distributed Systems: Architect, design, and maintain highly scalable, available, performant, and reliable distributed systems that power OpenAI's entire technology stack, including infrastructure supporting ChatGPT and the OpenAI API. Take ownership of critical system components that require careful consideration of fault-tolerance, performance optimization, and architectural patterns.
- Define Technical Strategy and Architecture: Collaborate with your infrastructure team to establish technical direction, system architecture, and long-term infrastructure goals. Contribute to strategic planning decisions that shape how OpenAI's engineering organization scales to support advanced AI research and product development.
- Cross-functional Collaboration and Stakeholder Management: Work closely with other infrastructure engineers, software engineers, product managers, and AI researchers to understand evolving infrastructure, data, and compute requirements. Translate stakeholder needs into technical solutions while maintaining clear communication about tradeoffs, timelines, and resource constraints.
- Improve Developer Experience and Internal Tooling: Enhance internal infrastructure tooling, automation frameworks, and developer workflows to enable engineers to build and deploy high-quality software faster and more safely. Focus on reducing operational friction and improving developer productivity across the organization.
- Incident Response and System Reliability: Lead incident response efforts, participate in postmortems, and develop best practices around system reliability, scalability, and security. Debug complex system issues, identify root causes, and implement preventative measures to minimize future incidents and improve overall system resilience.
- Debug System Bottlenecks and Solve Performance Problems: Work across multiple layers of the infrastructure stack to identify and resolve performance bottlenecks, evolve core infrastructure components, and tackle novel scalability challenges. Apply deep systems thinking to optimize resource utilization and improve end-to-end system performance.
Qualifications
What we look for.
Technical
Proficiency in Systems Programming Languages
Strong expertise in at least one or more of: Python, Go, C++, or Rust. Demonstrated ability to write efficient, maintainable code in production systems handling high-scale traffic and complex computational workloads.
Distributed Systems Design and Operation
Hands-on experience designing, building, operating, or scaling distributed systems architectures. Comfortable with concepts including consensus protocols, replication strategies, eventual consistency, distributed consensus, and fault-tolerance mechanisms.
Container Orchestration and Infrastructure-as-Code
Practical experience with Kubernetes for container orchestration, Terraform or similar infrastructure-as-code tools, and CI/CD pipelines. Ability to deploy, manage, and troubleshoot containerized applications in production environments.
Linux System Administration and Operations
Strong comfort working in Linux environments, including kernel-level troubleshooting, performance analysis, and system optimization. Experience with system-level debugging tools, performance profiling, and infrastructure monitoring.
Observability and Monitoring Solutions
Experience designing, implementing, or operating modern observability stacks including metrics, structured logging, distributed tracing, and alerting systems. Familiarity with tools like Prometheus, Grafana, ELK, Jaeger, or similar platforms.
Large-Scale System Debugging
Proven ability to navigate complex, distributed systems and dig deep when debugging challenging issues. Experience with root cause analysis, performance profiling, and systematic troubleshooting methodologies.
Education
Computer Science or Equivalent
Bachelor's degree in Computer Science, Computer Engineering, or equivalent professional experience demonstrating deep systems knowledge and software engineering fundamentals.
Experience
4+ Years Industry Experience
Minimum four years of relevant software engineering experience in production environments, with exposure to distributed systems, infrastructure engineering, or platform engineering at scale.
2+ Years Leading Complex Projects or Teams
At least two years of experience leading large-scale, complex technical projects or small teams as an engineer, tech lead, or infrastructure specialist. Demonstrated ability to drive projects to completion while mentoring other engineers and managing stakeholder expectations.
Distributed Systems at Scale
Hands-on experience working with distributed systems that handle significant scale, with focus on reliability engineering, scalability architecture, security hardening, and continuous improvement methodologies.
Skills
Required
Distributed Systems Architecture
Deep understanding of distributed systems principles including consistency models, failure scenarios, recovery mechanisms, and tradeoffs between availability, consistency, and partition tolerance.
Systems Programming
Expert-level proficiency in at least one systems programming language (Python, Go, C++, or Rust) with the ability to write performance-critical code.
Cloud-Native Infrastructure
Strong expertise with Kubernetes, containerization, microservices architectures, and cloud infrastructure platforms (AWS, GCP, or Azure).
Infrastructure-as-Code and Automation
Proficiency with Terraform, CloudFormation, Ansible, or similar tools for defining infrastructure declaratively and automating deployment pipelines.
Observability and Monitoring
Experience implementing comprehensive observability solutions including metrics, logging, tracing, and alerting for large-scale distributed systems.
Technical Leadership and Communication
Excellent communication and collaboration skills with ability to build consensus among diverse stakeholders, present technical strategies to leadership, and mentor junior engineers.
Complex Problem Solving
Demonstrated ability to tackle novel, ambiguous problems in large-scale systems, apply systems thinking, and develop creative solutions to performance and reliability challenges.
Preferred
Machine Learning Infrastructure Experience
Nice to haveExperience building or optimizing infrastructure specifically for machine learning workloads, including distributed training systems, data pipeline orchestration, or model serving platforms.
Database Systems Design
Nice to haveHands-on experience designing, operating, or scaling distributed database systems, including considerations for consistency, replication, and performance at scale.
Developer Tools and Productivity Engineering
Nice to haveExperience building developer-facing tools, internal platforms, or developer experience improvements in large engineering organizations.
Reliability Engineering and SRE Practices
Nice to haveFamiliarity with Site Reliability Engineering (SRE) methodologies, error budgets, blameless postmortems, and chaos engineering practices.
Open Source Contributions
Nice to haveActive contributions to open-source infrastructure projects (Kubernetes, etcd, Prometheus, etc.) demonstrating deep systems knowledge and community engagement.
Security and Compliance for Infrastructure
Nice to haveKnowledge of infrastructure security hardening, encryption strategies, secrets management, and compliance requirements for large-scale systems.
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 210,000 – 405,000
Equity·Stock options
Benefits
Competitive Equity and Stock Options
Participate in OpenAI's success with equity compensation packages that align your financial interests with company growth and long-term value creation.
Comprehensive Health and Wellness
Medical, dental, and vision coverage with company contributions toward premiums, mental health support, wellness programs, and fitness benefits.
Generous Time Off and Flexibility
Flexible vacation policy with unlimited PTO, paid parental leave, sabbatical opportunities, and flexible work arrangements supporting work-life balance.
Professional Development and Learning
Learning stipends, conference attendance budgets, internal training programs, mentorship from senior engineers, and opportunities to work on cutting-edge AI infrastructure challenges.
Competitive Base Salary
Market-competitive base salary reflecting senior-level infrastructure engineering expertise and the San Francisco Bay Area technology market.
401(k) Retirement Planning
Employer-matched 401(k) retirement savings plan with company contributions supporting long-term financial security.
Relocation Assistance
Comprehensive relocation support including moving expenses, temporary housing, and visa sponsorship for international candidates relocating to San Francisco.
Employee Discounts and Perks
Access to OpenAI API credits, employee discounts on OpenAI services, technology purchase programs, and partnership benefits with major cloud providers.
Inclusive and Supportive Culture
Diverse, inclusive team environment with commitment to equal opportunity employment, accessibility accommodations, and psychological safety for all engineers.
Full posting
Original listing.
About the Team
We’re hiring Software Engineers to join our broader Infrastructure organization, which supports multiple high-impact teams. Depending on your interests and experience, you could work on one of several focus areas—including Core Distributed Systems, Reliability Engineering, Observability, Developer Productivity or Cloud Infrastructure.
About the Role
All teams are deeply collaborative, work on mission-critical services, and are responsible for building distributed, scalable infrastructure to bring OpenAI’s technology to the world through products like ChatGPT and the OpenAI API. You’ll work closely with stakeholders to understand infrastructure, data and compute needs, setting the technical strategy that supports cutting-edge research and product development. This is a critical role for someone who is passionate about solving complex engineering problems at scale, ensuring their performance, scalability and reliability
Team Focus Areas
Distributed Systems: Owning and building important, highly scalable, available, performant, and reliable distributed systems (and their building blocks) to power the entire stack at OpenAI
Systems Engineering: Work across layers of the stack—debugging system bottlenecks, evolving core infrastructure, and solving novel problems in performance and scalability.
Reliability Engineering: Build scalable, fault-tolerant systems and lead efforts around service health, incident response, and resilience.
Observability: Design and maintain observability tooling (metrics, logs, tracing) to give teams visibility into production systems at scale.
Developer Productivity: Create tools, environments, and workflows that help engineers ship high-quality software faster and more safely.
Cloud Infrastructure: Own the cloud-native infrastructure (compute, networking, storage) that underpins all services and research workloads.
Databases: Building high performance, distributed database systems that power all of OpenAI's product stack.
In this role you will:
Design, build, and maintain reliable and performant systems used across engineering.
Work with your team to define technical strategy, architecture, and long-term goals.Collaborate with other engineers, product managers, and researchers to build infrastructure that meets evolving needs.
Improve internal tooling, automation, and developer experience.
Contribute to incident response, postmortems, and the development of best practices around system reliability and scalability.
You might thrive in this role if you:
Strong software engineering skills with experience in Python, Go, C++, Rust, or similar languages.
Experience designing, operating, or scaling distributed systems or developer infrastructure.
Comfort working in Linux environments, and with tools like Kubernetes, Terraform, CI/CD pipelines, and modern observability stacks.
Ability to navigate complex systems and a willingness to dig deep when debugging tricky issues.
Excellent communication and collaboration skills, especially in cross-functional settings.
Qualifications:
4+ years of relevant industry experience, with 2+ years leading large scale, complex projects or teams as an engineer or tech lead
A passion for distributed systems at scale with a focus on reliability, scalability, security, and continuous improvement.
Excellent communication skills, with ability to build consensus among stakeholders both internally and externally.
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity.
We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.
For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement.
Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.
To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance.
We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link.
OpenAI Global Applicant Privacy Policy
At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.
Redirects to OpenAI's application page.
Other roles
More at OpenAI.
Software Engineer, Astral
Senior
Software Security Architect, Operating Systems | Consumer Devices
Senior
Software Engineer, API Safety
Senior
Android Systems Engineer, Consumer Devices
Senior
Data Engineer, Monetization Data Platform
Senior