Software Engineer, Infrastructure

Infrastructure Engineer · Senior · Full Time

San FranciscoUSD 210k – 405k2d ago
Apply for this role

Opens OpenAI's application page

Role

What you'll do.

Join OpenAI's Infrastructure organization as a Software Engineer to design, build, and maintain mission-critical distributed systems powering ChatGPT and the OpenAI API. You'll collaborate with high-impact teams across Core Distributed Systems, Reliability Engineering, Observability, Developer Productivity, Cloud Infrastructure, and Databases, working on scalable, fault-tolerant systems that support cutting-edge AI research and products. This role requires 4+ years of industry experience with 2+ years leading complex projects, deep expertise in distributed systems architecture, and proficiency in languages like Python, Go, C++, or Rust.

Responsibilities

  • Design and Build Reliable Distributed Systems: Architect, design, and maintain highly scalable, available, performant, and reliable distributed systems that power OpenAI's entire technology stack, including infrastructure supporting ChatGPT and the OpenAI API. Take ownership of critical system components that require careful consideration of fault-tolerance, performance optimization, and architectural patterns.
  • Define Technical Strategy and Architecture: Collaborate with your infrastructure team to establish technical direction, system architecture, and long-term infrastructure goals. Contribute to strategic planning decisions that shape how OpenAI's engineering organization scales to support advanced AI research and product development.
  • Cross-functional Collaboration and Stakeholder Management: Work closely with other infrastructure engineers, software engineers, product managers, and AI researchers to understand evolving infrastructure, data, and compute requirements. Translate stakeholder needs into technical solutions while maintaining clear communication about tradeoffs, timelines, and resource constraints.
  • Improve Developer Experience and Internal Tooling: Enhance internal infrastructure tooling, automation frameworks, and developer workflows to enable engineers to build and deploy high-quality software faster and more safely. Focus on reducing operational friction and improving developer productivity across the organization.
  • Incident Response and System Reliability: Lead incident response efforts, participate in postmortems, and develop best practices around system reliability, scalability, and security. Debug complex system issues, identify root causes, and implement preventative measures to minimize future incidents and improve overall system resilience.
  • Debug System Bottlenecks and Solve Performance Problems: Work across multiple layers of the infrastructure stack to identify and resolve performance bottlenecks, evolve core infrastructure components, and tackle novel scalability challenges. Apply deep systems thinking to optimize resource utilization and improve end-to-end system performance.

Qualifications

What we look for.

Technical

  • Proficiency in Systems Programming Languages

    Strong expertise in at least one or more of: Python, Go, C++, or Rust. Demonstrated ability to write efficient, maintainable code in production systems handling high-scale traffic and complex computational workloads.

  • Distributed Systems Design and Operation

    Hands-on experience designing, building, operating, or scaling distributed systems architectures. Comfortable with concepts including consensus protocols, replication strategies, eventual consistency, distributed consensus, and fault-tolerance mechanisms.

  • Container Orchestration and Infrastructure-as-Code

    Practical experience with Kubernetes for container orchestration, Terraform or similar infrastructure-as-code tools, and CI/CD pipelines. Ability to deploy, manage, and troubleshoot containerized applications in production environments.

  • Linux System Administration and Operations

    Strong comfort working in Linux environments, including kernel-level troubleshooting, performance analysis, and system optimization. Experience with system-level debugging tools, performance profiling, and infrastructure monitoring.

  • Observability and Monitoring Solutions

    Experience designing, implementing, or operating modern observability stacks including metrics, structured logging, distributed tracing, and alerting systems. Familiarity with tools like Prometheus, Grafana, ELK, Jaeger, or similar platforms.

  • Large-Scale System Debugging

    Proven ability to navigate complex, distributed systems and dig deep when debugging challenging issues. Experience with root cause analysis, performance profiling, and systematic troubleshooting methodologies.

Education

  • Computer Science or Equivalent

    Bachelor's degree in Computer Science, Computer Engineering, or equivalent professional experience demonstrating deep systems knowledge and software engineering fundamentals.

Experience

  • 4+ Years Industry Experience

    Minimum four years of relevant software engineering experience in production environments, with exposure to distributed systems, infrastructure engineering, or platform engineering at scale.

  • 2+ Years Leading Complex Projects or Teams

    At least two years of experience leading large-scale, complex technical projects or small teams as an engineer, tech lead, or infrastructure specialist. Demonstrated ability to drive projects to completion while mentoring other engineers and managing stakeholder expectations.

  • Distributed Systems at Scale

    Hands-on experience working with distributed systems that handle significant scale, with focus on reliability engineering, scalability architecture, security hardening, and continuous improvement methodologies.

Skills

Required

  • Distributed Systems Architecture

    Deep understanding of distributed systems principles including consistency models, failure scenarios, recovery mechanisms, and tradeoffs between availability, consistency, and partition tolerance.

  • Systems Programming

    Expert-level proficiency in at least one systems programming language (Python, Go, C++, or Rust) with the ability to write performance-critical code.

  • Cloud-Native Infrastructure

    Strong expertise with Kubernetes, containerization, microservices architectures, and cloud infrastructure platforms (AWS, GCP, or Azure).

  • Infrastructure-as-Code and Automation

    Proficiency with Terraform, CloudFormation, Ansible, or similar tools for defining infrastructure declaratively and automating deployment pipelines.

  • Observability and Monitoring

    Experience implementing comprehensive observability solutions including metrics, logging, tracing, and alerting for large-scale distributed systems.

  • Technical Leadership and Communication

    Excellent communication and collaboration skills with ability to build consensus among diverse stakeholders, present technical strategies to leadership, and mentor junior engineers.

  • Complex Problem Solving

    Demonstrated ability to tackle novel, ambiguous problems in large-scale systems, apply systems thinking, and develop creative solutions to performance and reliability challenges.

Preferred

  • Machine Learning Infrastructure Experience

    Nice to have

    Experience building or optimizing infrastructure specifically for machine learning workloads, including distributed training systems, data pipeline orchestration, or model serving platforms.

  • Database Systems Design

    Nice to have

    Hands-on experience designing, operating, or scaling distributed database systems, including considerations for consistency, replication, and performance at scale.

  • Developer Tools and Productivity Engineering

    Nice to have

    Experience building developer-facing tools, internal platforms, or developer experience improvements in large engineering organizations.

  • Reliability Engineering and SRE Practices

    Nice to have

    Familiarity with Site Reliability Engineering (SRE) methodologies, error budgets, blameless postmortems, and chaos engineering practices.

  • Open Source Contributions

    Nice to have

    Active contributions to open-source infrastructure projects (Kubernetes, etcd, Prometheus, etc.) demonstrating deep systems knowledge and community engagement.

  • Security and Compliance for Infrastructure

    Nice to have

    Knowledge of infrastructure security hardening, encryption strategies, secrets management, and compliance requirements for large-scale systems.

Tech stack

Languages

PythonGoC++Rust

Frameworks

KubernetesgRPCApache Kafka

Databases

PostgreSQLDistributed DatabasesTime-Series Databases

Tools

TerraformGit and CI/CDDockerPrometheusGrafanaELK StackDatadog / New Relic

Other

Linux Kernel and System AdministrationDistributed Consensus AlgorithmsAPI DesignSecurity and EncryptionLoad Balancing and Networking

Compensation

Pay and benefits.

Base·USD 210,000 – 405,000

Equity·Stock options

Benefits

  • Competitive Equity and Stock Options

    Participate in OpenAI's success with equity compensation packages that align your financial interests with company growth and long-term value creation.

  • Comprehensive Health and Wellness

    Medical, dental, and vision coverage with company contributions toward premiums, mental health support, wellness programs, and fitness benefits.

  • Generous Time Off and Flexibility

    Flexible vacation policy with unlimited PTO, paid parental leave, sabbatical opportunities, and flexible work arrangements supporting work-life balance.

  • Professional Development and Learning

    Learning stipends, conference attendance budgets, internal training programs, mentorship from senior engineers, and opportunities to work on cutting-edge AI infrastructure challenges.

  • Competitive Base Salary

    Market-competitive base salary reflecting senior-level infrastructure engineering expertise and the San Francisco Bay Area technology market.

  • 401(k) Retirement Planning

    Employer-matched 401(k) retirement savings plan with company contributions supporting long-term financial security.

  • Relocation Assistance

    Comprehensive relocation support including moving expenses, temporary housing, and visa sponsorship for international candidates relocating to San Francisco.

  • Employee Discounts and Perks

    Access to OpenAI API credits, employee discounts on OpenAI services, technology purchase programs, and partnership benefits with major cloud providers.

  • Inclusive and Supportive Culture

    Diverse, inclusive team environment with commitment to equal opportunity employment, accessibility accommodations, and psychological safety for all engineers.

Full posting

Original listing.

About the Team

We’re hiring Software Engineers to join our broader Infrastructure organization, which supports multiple high-impact teams. Depending on your interests and experience, you could work on one of several focus areas—including Core Distributed Systems, Reliability Engineering, Observability, Developer Productivity or Cloud Infrastructure. 

About the Role

All teams are deeply collaborative, work on mission-critical services, and are responsible for building distributed, scalable infrastructure to bring OpenAI’s technology to the world through products like ChatGPT and the OpenAI API. You’ll work closely with stakeholders to understand infrastructure, data and compute needs, setting the technical strategy that supports cutting-edge research and product development. This is a critical role for someone who is passionate about solving complex engineering problems at scale, ensuring their performance, scalability and reliability 

Team Focus Areas

  • Distributed Systems: Owning and building important, highly scalable, available, performant, and reliable distributed systems (and their building blocks) to power the entire stack at OpenAI 

  • Systems Engineering: Work across layers of the stack—debugging system bottlenecks, evolving core infrastructure, and solving novel problems in performance and scalability.

  • Reliability Engineering: Build scalable, fault-tolerant systems and lead efforts around service health, incident response, and resilience.

  • Observability: Design and maintain observability tooling (metrics, logs, tracing) to give teams visibility into production systems at scale.

  • Developer Productivity: Create tools, environments, and workflows that help engineers ship high-quality software faster and more safely.

  • Cloud Infrastructure: Own the cloud-native infrastructure (compute, networking, storage) that underpins all services and research workloads.

  • Databases: Building high performance, distributed database systems that power all of OpenAI's product stack.

In this role you will: 

  • Design, build, and maintain reliable and performant systems used across engineering.
    Work with your team to define technical strategy, architecture, and long-term goals.

  • Collaborate with other engineers, product managers, and researchers to build infrastructure that meets evolving needs.

  • Improve internal tooling, automation, and developer experience.

  • Contribute to incident response, postmortems, and the development of best practices around system reliability and scalability.

You might thrive in this role if you:

  • Strong software engineering skills with experience in Python, Go, C++, Rust, or similar languages.

  • Experience designing, operating, or scaling distributed systems or developer infrastructure.

  • Comfort working in Linux environments, and with tools like Kubernetes, Terraform, CI/CD pipelines, and modern observability stacks.

  • Ability to navigate complex systems and a willingness to dig deep when debugging tricky issues.

  • Excellent communication and collaboration skills, especially in cross-functional settings.

Qualifications: 

  • 4+ years of relevant industry experience, with 2+ years leading large scale, complex projects or teams as an engineer or tech lead

  • A passion for distributed systems at scale with a focus on reliability, scalability, security, and continuous improvement. 

  • Excellent communication skills, with ability to build consensus among stakeholders both internally and externally.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. 

We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.

For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement.

Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.

To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance.

We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link.

OpenAI Global Applicant Privacy Policy

At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.

Redirects to OpenAI's application page.

Other roles

More at OpenAI.

View all 111 roles