Software Engineer, GPU Infrastructure- ChatGPT Engineering

Infrastructure Engineer · Senior · Full Time

London, UKUSD 180k – 300k1mo ago
Apply for this role

Opens OpenAI's application page

Role

What you'll do.

Join OpenAI's ChatGPT Engineering team to design and operate GPU infrastructure powering one of the world's largest AI products. This role offers the opportunity to build scalable systems managing thousands of GPUs, develop intelligent automation tooling, and directly impact AGI development through infrastructure improvements. You'll work across production engineering, distributed systems, and capacity management with 5+ years of production infrastructure experience and expertise in systems languages, Kubernetes, and distributed systems design.

Responsibilities

  • Design and operate GPU infrastructure management systems: Design, build, and operate software systems that manage large-scale GPU clusters supporting ChatGPT inference workloads. This includes developing core infrastructure services for fleet orchestration, resource allocation, and compute platform operations supporting millions of concurrent inference requests.
  • Build internal platforms and AI-powered automation: Develop internal platforms, tooling, and intelligent agents that automate fleet operations, reduce manual intervention, and enable self-healing infrastructure. Create systems that leverage AI to predict and prevent infrastructure issues before they impact production workloads.
  • Improve observability, reliability, and efficiency: Enhance observability and monitoring across thousands of GPUs to achieve high reliability and operational efficiency. Implement comprehensive distributed tracing, metrics collection, and alerting systems that provide real-time insights into fleet health and performance.
  • Develop capacity planning and fleet health systems: Build and maintain systems for capacity planning, intelligent scheduling, fleet health monitoring, and automated incident response. Create data-driven tools that optimize GPU utilization while ensuring service level agreements are consistently met.
  • Identify and eliminate infrastructure bottlenecks: Analyze production infrastructure to identify performance bottlenecks and scalability constraints. Implement solutions that improve resource utilization, reduce operational latency, and enable the infrastructure to scale with demand.
  • Cross-team collaboration and platform improvement: Partner closely with research, platform, networking, and systems teams to continuously improve the compute platform. Establish engineering best practices around operational excellence, infrastructure automation, and reliability patterns.

Qualifications

What we look for.

Technical

  • Systems programming languages

    Strong programming proficiency in Go, Python, C++, Rust, or similar systems-level languages for building high-performance infrastructure software.

  • Distributed systems architecture

    Deep understanding of distributed systems concepts including consensus protocols, eventual consistency, fault tolerance, and designing systems for high availability and scalability.

  • Linux systems expertise

    Comprehensive knowledge of Linux kernel internals, system performance tuning, networking concepts, and low-level system debugging techniques.

  • Observability tooling

    Experience with observability and monitoring stacks including metrics collection, distributed tracing, log aggregation, and alerting systems.

  • Systems design and debugging

    Excellent systems design skills with strong debugging capabilities for complex production issues, performance analysis, and infrastructure troubleshooting.

  • High-performance computing knowledge

    Understanding of GPU architecture, compute optimization, performance profiling, and resource management for compute-intensive workloads.

Education

  • Computer Science or related field

    Bachelor's degree in Computer Science, Engineering, or related field, or equivalent professional experience demonstrating systems-level expertise.

Experience

  • Production infrastructure operations

    5+ years of software engineering experience building and operating production infrastructure systems, with proven track record managing large-scale distributed systems in production environments.

  • GPU or compute infrastructure

    Demonstrated experience operating GPU infrastructure, high-performance computing clusters, ML infrastructure platforms, or other large-scale compute platforms at production scale.

  • Kubernetes and container orchestration

    Hands-on experience with Kubernetes, cloud infrastructure platforms, Linux systems administration, and container orchestration technologies in production environments.

  • Distributed systems design and operations

    Experience designing and operating highly available distributed systems with focus on reliability, fault tolerance, and high availability patterns in production settings.

  • Observability and monitoring

    Strong background in infrastructure observability, monitoring systems, capacity planning, incident management, and root cause analysis of production issues.

  • Infrastructure automation

    Proven ability to build software that automates operational workflows and reduces manual operational overhead through infrastructure automation and tooling.

Skills

Required

  • Go programming

    Production experience with Go for building scalable infrastructure services and tooling

  • Python

    Proficiency in Python for infrastructure automation, scripting, and systems tools

  • Kubernetes

    Hands-on experience deploying, scaling, and debugging Kubernetes clusters in production

  • Linux administration

    Deep knowledge of Linux system administration, kernel concepts, and system performance tuning

  • Distributed systems

    Strong understanding of distributed systems design patterns, consensus algorithms, and fault tolerance

  • Cloud infrastructure

    Experience with cloud platforms and infrastructure-as-code principles

  • Monitoring and observability

    Experience building or operating monitoring systems, metrics collection, and alerting infrastructure

  • Production debugging

    Proven ability to diagnose and resolve complex production infrastructure issues

Preferred

  • C++ or Rust

    Nice to have

    Experience with C++ or Rust for high-performance systems programming

  • GPU infrastructure experience

    Nice to have

    Previous experience managing GPU clusters or working with NVIDIA GPU infrastructure

  • ML infrastructure platforms

    Nice to have

    Familiarity with ML training infrastructure or large-scale ML platform operations

  • Network programming

    Nice to have

    Knowledge of network protocols, socket programming, and network optimization

  • Production engineering background

    Nice to have

    Background in SRE, Production Engineering, or Platform Engineering roles

  • AI-powered systems

    Nice to have

    Experience building systems that leverage AI for operational automation or prediction

  • Capacity planning tools

    Nice to have

    Experience developing capacity planning or resource optimization systems

Tech stack

Languages

GoPythonC++Rust

Frameworks

KubernetesDocker

Databases

Time-series databasesDistributed databases

Tools

PrometheusGrafanaGitLinuxCloud platforms

Other

Infrastructure-as-CodeGPU architecture and optimizationDistributed tracingIncident management

Compensation

Pay and benefits.

Base·USD 180,000 – 300,000

Equity·Stock options

Benefits

  • Competitive equity package

    Significant stock options as part of compensation package, providing upside participation in OpenAI's growth

  • Comprehensive health and wellness

    Medical, dental, and vision insurance coverage with focus on employee wellbeing

  • Retirement planning

    401(k) plan with company matching to support long-term financial security

  • Flexible work arrangements

    Collaborative work environment designed for distributed and flexible working patterns

  • Professional development

    Opportunity to work at the frontier of AI infrastructure with exposure to cutting-edge technology and industry-leading engineering practices

  • Impact-driven mission

    Opportunity to directly influence AGI development through infrastructure improvements benefiting millions of users globally

  • Inclusive workplace

    Commitment to diverse perspectives and inclusive culture where all backgrounds can do their best work

  • Learning and growth

    Access to world-class engineering talent and exposure to large-scale infrastructure challenges solving real-world problems

Process

Interview steps.

  1. 01

    Initial screening call

    Recruiter conversation to discuss your background, experience with infrastructure systems, and alignment with the role's technical requirements

  2. 02

    Technical assessment

    In-depth technical discussion covering systems design, infrastructure architecture, production debugging, and problem-solving approaches

  3. 03

    Infrastructure systems design interview

    Whiteboard-style systems design interview focused on designing large-scale GPU infrastructure systems, capacity planning, and operational automation

  4. 04

    Operational problem-solving

    Interview discussing real infrastructure challenges, incident response approaches, and how you'd approach operational problem-solving at scale

  5. 05

    Cross-team collaboration discussion

    Conversation with team members about collaboration style, communication skills, and experience working across research, platform, and product teams

  6. 06

    Leadership conversation

    Meeting with engineering leadership to discuss vision for infrastructure, long-term goals, and philosophy on infrastructure reliability and automation

Full posting

Original listing.

About the Team

ChatGPT Engineering builds and operates the compute platform powering one of the world's largest AI products. Every ChatGPT conversation relies on massive GPU clusters serving inference workloads with high reliability, efficiency, and performance.

As our GPU fleet continues to grow, we're investing in the infrastructure that operates it. Our team builds the tooling, automation, and intelligent systems that make GPU infrastructure scalable, observable, and increasingly autonomous. We work across production engineering, distributed systems, capacity management, and AI-powered operational tooling to help researchers and product teams move faster while maximizing the efficiency of every GPU.

This is a unique opportunity to work on infrastructure at the frontier of AI, where small improvements in fleet efficiency, reliability, and automation have an outsized impact on the development and deployment of AGI.

About the Role

We're looking for a Software Engineer with deep experience operating large-scale GPU or compute infrastructure.

You'll design and build the systems that manage GPU clusters at scale—from fleet health and capacity planning to operational automation and intelligent agents that reduce manual intervention. You'll partner closely with infrastructure, research, and product engineering teams to improve reliability, developer productivity, and overall compute utilization.

This role is ideal for engineers who enjoy solving complex operational challenges, building internal platforms, and working on infrastructure that directly powers frontier AI.

In This Role, You Will

  • Design, build, and operate software that manages large-scale GPU infrastructure supporting ChatGPT inference.

  • Build internal platforms, tooling, and AI-powered agents that automate fleet operations and reduce operational overhead.

  • Improve observability, reliability, and operational efficiency across thousands of GPUs.

  • Develop systems for capacity planning, scheduling, fleet health monitoring, and incident response.

  • Identify infrastructure bottlenecks and implement solutions that improve utilization, scalability, and performance.

  • Partner closely with research, platform, networking, and systems teams to continuously improve our compute platform.

  • Help establish engineering best practices around operational excellence, automation, and infrastructure reliability.

You Might Thrive in This Role If You

  • Have experience operating large-scale production infrastructure, preferably GPU clusters or other compute-intensive distributed systems.

  • Have a background in Production Engineering, Site Reliability Engineering (SRE), Infrastructure Engineering, or Platform Engineering.

  • Have built software that automates operational workflows rather than relying on manual processes.

  • Have experience with Kubernetes, Linux systems, container orchestration, or distributed infrastructure.

  • Understand infrastructure observability, monitoring, capacity planning, and incident management.

  • Enjoy identifying cross-team pain points and building reusable platforms that improve developer productivity.

  • Are comfortable working across software engineering and systems operations, owning problems end-to-end.

  • Thrive in fast-moving environments with significant technical ambiguity.

Qualifications

  • 5+ years of software engineering experience building production infrastructure.

  • Strong programming skills in Go, Python, C++, Rust, or similar systems languages.

  • Experience designing and operating highly available distributed systems.

  • Experience with GPU infrastructure, high-performance computing, ML infrastructure, or large-scale compute platforms.

  • Experience with Kubernetes, cloud infrastructure, Linux, networking, and observability tooling.

  • Excellent debugging, systems design, and operational problem-solving skills.

  • Strong communication skills and experience collaborating across engineering organizations.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We build AI systems that are capable, aligned, and broadly beneficial. Our infrastructure teams power the research and products that bring these systems to millions of users worldwide.

We believe diverse perspectives make stronger teams and better technology. We're committed to creating an inclusive workplace where people from all backgrounds can do their best work.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. 

We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.

For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement.

Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.

To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance.

We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link.

OpenAI Global Applicant Privacy Policy

At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.

Redirects to OpenAI's application page.

Other roles

More at OpenAI.

View all 107 roles