Software Engineer, Cloud Infrastructure

Backend Engineer · Senior · Full Time

San FranciscoUSD 230k – 490k12mo ago
Apply for this role

Opens OpenAI's application page

Role

What you'll do.

OpenAI is seeking a Senior Software Engineer to join their Cloud Infrastructure team, responsible for building and maintaining the core infrastructure that powers ChatGPT and their API services. The role involves designing scalable Kubernetes-based platforms, cloud abstractions, and ensuring infrastructure can handle massive scale while maintaining reliability and security standards.

Responsibilities

  • Infrastructure Platform Design: Design and build scalable development and production platforms that power OpenAI's products including ChatGPT and API services
  • Scale Architecture Planning: Ensure infrastructure architecture can scale to the next order of magnitude to support massive AI workloads
  • Kubernetes Operations: Operate and maintain large-scale Kubernetes clusters supporting critical AI model inference and training workloads
  • Cloud Abstraction Development: Build and maintain cloud infrastructure abstractions that enable rapid product deployment across multiple cloud providers
  • System Reliability Engineering: Maintain high availability and reliability standards for production systems supporting millions of users
  • Incident Response: Participate in on-call rotation to respond to critical infrastructure incidents and ensure rapid resolution
  • Security Implementation: Implement and maintain security best practices across all infrastructure components and deployment pipelines
  • Performance Optimization: Monitor and optimize infrastructure performance to support AI model inference at scale
  • Cross-team Collaboration: Work closely with research, product, and design teams to enable rapid iteration and deployment of AI technologies
  • Infrastructure Automation: Develop automation tools and processes to reduce manual operations and improve deployment reliability

Qualifications

What we look for.

Technical

  • Infrastructure Experience

    5+ years of experience building and maintaining core infrastructure systems at scale

  • Kubernetes Expertise

    Extensive experience operating Kubernetes orchestration systems in production environments

  • Cloud Platform Proficiency

    Deep experience building abstractions and working with major cloud platforms (AWS, GCP, Azure)

  • System Architecture

    Strong background in designing scalable, reliable, and secure distributed systems

  • Infrastructure as Code

    Proficiency with infrastructure automation tools like Terraform, Ansible, or similar

  • Monitoring and Observability

    Experience with monitoring systems, logging, and observability tools for large-scale infrastructure

  • Network Architecture

    Understanding of networking concepts including load balancing, CDNs, and service mesh technologies

Education

  • Computer Science Degree

    Bachelor's degree in Computer Science, Engineering, or equivalent practical experience

  • Advanced Technical Education

    Master's degree in Computer Science or related field preferred but not required

Experience

  • Senior Infrastructure Role

    5+ years in senior infrastructure engineering roles with increasing responsibility

  • Scale Operations

    Experience operating infrastructure supporting millions of users or high-throughput applications

  • DevOps Culture

    Background working in DevOps or SRE environments with emphasis on automation and reliability

  • Startup Environment

    Experience working in fast-paced, high-growth technology companies preferred

Skills

Required

  • Kubernetes Administration

    Expert-level skills in Kubernetes cluster management, networking, and troubleshooting

  • Cloud Architecture

    Advanced knowledge of cloud-native architecture patterns and multi-cloud strategies

  • Infrastructure Automation

    Proficiency in infrastructure as code tools and CI/CD pipeline development

  • System Design

    Strong system design skills for building scalable and reliable distributed systems

  • Incident Management

    Experience with incident response, troubleshooting, and post-mortem analysis

  • Security Best Practices

    Knowledge of infrastructure security, compliance, and vulnerability management

Preferred

  • AI/ML Infrastructure

    Nice to have

    Experience with infrastructure supporting machine learning workloads and GPU clusters

  • Service Mesh

    Nice to have

    Hands-on experience with service mesh technologies like Istio or Linkerd

  • Observability Platforms

    Nice to have

    Advanced experience with observability tools like Prometheus, Grafana, and distributed tracing

  • Multi-Cloud Strategy

    Nice to have

    Experience designing and implementing multi-cloud infrastructure strategies

  • Performance Tuning

    Nice to have

    Skills in performance optimization for high-throughput, low-latency systems

  • Open Source Contributions

    Nice to have

    Active contributions to infrastructure-related open source projects

Tech stack

Languages

GoPythonRustBash/Shell

Frameworks

KubernetesTerraformHelm

Databases

PostgreSQLRedisetcd

Tools

DockerPrometheusGrafanaArgoCDIstio

Other

AWS/GCP/AzureJenkins/GitHub ActionsAnsibleVault

Compensation

Pay and benefits.

Base·USD 230,000 – 490,000

Equity·Stock options

Benefits

  • Equity Compensation

    Competitive equity package with significant upside potential in a rapidly growing AI company

  • Health Insurance

    Comprehensive medical, dental, and vision insurance coverage

  • Relocation Assistance

    Full relocation support for candidates moving to San Francisco

  • Professional Development

    Access to cutting-edge AI research and learning opportunities

  • Flexible PTO

    Generous time off policy to support work-life balance

  • Retirement Benefits

    401(k) plan with company matching

  • Parental Leave

    Comprehensive parental leave policy for new parents

  • Mental Health Support

    Access to mental health resources and counseling services

  • Learning Budget

    Annual budget for conferences, courses, and professional development

  • Gym/Wellness

    Fitness and wellness benefits including gym membership reimbursement

Process

Interview steps.

  1. 01

    Application Review

    Initial screening of resume and technical background by recruiting team

  2. 02

    Recruiter Phone Screen

    30-minute call to discuss role fit, experience, and answer initial questions

  3. 03

    Technical Phone Interview

    60-minute technical discussion focusing on infrastructure design and Kubernetes experience

  4. 04

    System Design Interview

    90-minute session designing scalable infrastructure solutions for AI workloads

  5. 05

    Technical Deep Dive

    Detailed technical interview covering cloud platforms, monitoring, and incident response

  6. 06

    Cultural Fit Interview

    Discussion with team members about OpenAI's mission, values, and collaborative approach

  7. 07

    Final Interview Round

    Meetings with senior leadership and potential team members

  8. 08

    Reference Checks

    Professional reference verification and background check process

Full posting

Original listing.

About the Team

The Applied Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses.

You’ll join the team responsible for running the core infrastructure that supports products like ChatGPT and the API. The systems we support include our kubernetes clusters, infrastructure deployment, our networking stack, cloud abstractions, and more.

We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth.

About the Role

The cloud infrastructure team builds and maintains infrastructure abstractions allowing OpenAI to ship products quickly and scalably.

This role is based in San Francisco, CA.

In this role, you will:

  • Design and build the development and production platforms that power our products, enabling reliability and security at scale

  • Ensure our infrastructure can scale to the next order of magnitude

  • Help create a diverse, equitable, and inclusive culture that makes all feel welcome while enabling radical candor and the challenging of group think

  • Like all other teams, we are responsible for the reliability of the systems we build. This includes an on-call rotation to respond to critical incidents as needed.

You might thrive in this role if you:

  • Have 5+ years building core infrastructure

  • Have experience operating orchestration systems such as Kubernetes at scale

  • Have experience building abstractions over cloud platforms

  • Take pride in building and operating scalable, reliable, secure systems

  • Are comfortable with ambiguity and rapid change

This role is exclusively based in our San Francisco HQ. We offer relocation assistance to new employees.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. 

We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.

For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement.

Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.

To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance.

We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link.

OpenAI Global Applicant Privacy Policy

At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.

Redirects to OpenAI's application page.

Other roles

More at OpenAI.

View all 125 roles