Software Engineer, GPU Infrastructure- ChatGPT Engineering
Infrastructure Engineer · Senior · Full Time
Opens OpenAI's application page
Role
What you'll do.
Join OpenAI's ChatGPT Engineering team to design and operate GPU infrastructure powering one of the world's largest AI products. This role offers the opportunity to build scalable systems managing thousands of GPUs, develop intelligent automation tooling, and directly impact AGI development through infrastructure improvements. You'll work across production engineering, distributed systems, and capacity management with 5+ years of production infrastructure experience and expertise in systems languages, Kubernetes, and distributed systems design.
Responsibilities
- Design and operate GPU infrastructure management systems: Design, build, and operate software systems that manage large-scale GPU clusters supporting ChatGPT inference workloads. This includes developing core infrastructure services for fleet orchestration, resource allocation, and compute platform operations supporting millions of concurrent inference requests.
- Build internal platforms and AI-powered automation: Develop internal platforms, tooling, and intelligent agents that automate fleet operations, reduce manual intervention, and enable self-healing infrastructure. Create systems that leverage AI to predict and prevent infrastructure issues before they impact production workloads.
- Improve observability, reliability, and efficiency: Enhance observability and monitoring across thousands of GPUs to achieve high reliability and operational efficiency. Implement comprehensive distributed tracing, metrics collection, and alerting systems that provide real-time insights into fleet health and performance.
- Develop capacity planning and fleet health systems: Build and maintain systems for capacity planning, intelligent scheduling, fleet health monitoring, and automated incident response. Create data-driven tools that optimize GPU utilization while ensuring service level agreements are consistently met.
- Identify and eliminate infrastructure bottlenecks: Analyze production infrastructure to identify performance bottlenecks and scalability constraints. Implement solutions that improve resource utilization, reduce operational latency, and enable the infrastructure to scale with demand.
- Cross-team collaboration and platform improvement: Partner closely with research, platform, networking, and systems teams to continuously improve the compute platform. Establish engineering best practices around operational excellence, infrastructure automation, and reliability patterns.
Qualifications
What we look for.
Technical
Systems programming languages
Strong programming proficiency in Go, Python, C++, Rust, or similar systems-level languages for building high-performance infrastructure software.
Distributed systems architecture
Deep understanding of distributed systems concepts including consensus protocols, eventual consistency, fault tolerance, and designing systems for high availability and scalability.
Linux systems expertise
Comprehensive knowledge of Linux kernel internals, system performance tuning, networking concepts, and low-level system debugging techniques.
Observability tooling
Experience with observability and monitoring stacks including metrics collection, distributed tracing, log aggregation, and alerting systems.
Systems design and debugging
Excellent systems design skills with strong debugging capabilities for complex production issues, performance analysis, and infrastructure troubleshooting.
High-performance computing knowledge
Understanding of GPU architecture, compute optimization, performance profiling, and resource management for compute-intensive workloads.
Education
Computer Science or related field
Bachelor's degree in Computer Science, Engineering, or related field, or equivalent professional experience demonstrating systems-level expertise.
Experience
Production infrastructure operations
5+ years of software engineering experience building and operating production infrastructure systems, with proven track record managing large-scale distributed systems in production environments.
GPU or compute infrastructure
Demonstrated experience operating GPU infrastructure, high-performance computing clusters, ML infrastructure platforms, or other large-scale compute platforms at production scale.
Kubernetes and container orchestration
Hands-on experience with Kubernetes, cloud infrastructure platforms, Linux systems administration, and container orchestration technologies in production environments.
Distributed systems design and operations
Experience designing and operating highly available distributed systems with focus on reliability, fault tolerance, and high availability patterns in production settings.
Observability and monitoring
Strong background in infrastructure observability, monitoring systems, capacity planning, incident management, and root cause analysis of production issues.
Infrastructure automation
Proven ability to build software that automates operational workflows and reduces manual operational overhead through infrastructure automation and tooling.
Skills
Required
Go programming
Production experience with Go for building scalable infrastructure services and tooling
Python
Proficiency in Python for infrastructure automation, scripting, and systems tools
Kubernetes
Hands-on experience deploying, scaling, and debugging Kubernetes clusters in production
Linux administration
Deep knowledge of Linux system administration, kernel concepts, and system performance tuning
Distributed systems
Strong understanding of distributed systems design patterns, consensus algorithms, and fault tolerance
Cloud infrastructure
Experience with cloud platforms and infrastructure-as-code principles
Monitoring and observability
Experience building or operating monitoring systems, metrics collection, and alerting infrastructure
Production debugging
Proven ability to diagnose and resolve complex production infrastructure issues
Preferred
C++ or Rust
Nice to haveExperience with C++ or Rust for high-performance systems programming
GPU infrastructure experience
Nice to havePrevious experience managing GPU clusters or working with NVIDIA GPU infrastructure
ML infrastructure platforms
Nice to haveFamiliarity with ML training infrastructure or large-scale ML platform operations
Network programming
Nice to haveKnowledge of network protocols, socket programming, and network optimization
Production engineering background
Nice to haveBackground in SRE, Production Engineering, or Platform Engineering roles
AI-powered systems
Nice to haveExperience building systems that leverage AI for operational automation or prediction
Capacity planning tools
Nice to haveExperience developing capacity planning or resource optimization systems
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 180,000 – 300,000
Equity·Stock options
Benefits
Competitive equity package
Significant stock options as part of compensation package, providing upside participation in OpenAI's growth
Comprehensive health and wellness
Medical, dental, and vision insurance coverage with focus on employee wellbeing
Retirement planning
401(k) plan with company matching to support long-term financial security
Flexible work arrangements
Collaborative work environment designed for distributed and flexible working patterns
Professional development
Opportunity to work at the frontier of AI infrastructure with exposure to cutting-edge technology and industry-leading engineering practices
Impact-driven mission
Opportunity to directly influence AGI development through infrastructure improvements benefiting millions of users globally
Inclusive workplace
Commitment to diverse perspectives and inclusive culture where all backgrounds can do their best work
Learning and growth
Access to world-class engineering talent and exposure to large-scale infrastructure challenges solving real-world problems
Process
Interview steps.
- 01
Initial screening call
Recruiter conversation to discuss your background, experience with infrastructure systems, and alignment with the role's technical requirements
- 02
Technical assessment
In-depth technical discussion covering systems design, infrastructure architecture, production debugging, and problem-solving approaches
- 03
Infrastructure systems design interview
Whiteboard-style systems design interview focused on designing large-scale GPU infrastructure systems, capacity planning, and operational automation
- 04
Operational problem-solving
Interview discussing real infrastructure challenges, incident response approaches, and how you'd approach operational problem-solving at scale
- 05
Cross-team collaboration discussion
Conversation with team members about collaboration style, communication skills, and experience working across research, platform, and product teams
- 06
Leadership conversation
Meeting with engineering leadership to discuss vision for infrastructure, long-term goals, and philosophy on infrastructure reliability and automation
Full posting
Original listing.
About the Team
ChatGPT Engineering builds and operates the compute platform powering one of the world's largest AI products. Every ChatGPT conversation relies on massive GPU clusters serving inference workloads with high reliability, efficiency, and performance.
As our GPU fleet continues to grow, we're investing in the infrastructure that operates it. Our team builds the tooling, automation, and intelligent systems that make GPU infrastructure scalable, observable, and increasingly autonomous. We work across production engineering, distributed systems, capacity management, and AI-powered operational tooling to help researchers and product teams move faster while maximizing the efficiency of every GPU.
This is a unique opportunity to work on infrastructure at the frontier of AI, where small improvements in fleet efficiency, reliability, and automation have an outsized impact on the development and deployment of AGI.
About the Role
We're looking for a Software Engineer with deep experience operating large-scale GPU or compute infrastructure.
You'll design and build the systems that manage GPU clusters at scale—from fleet health and capacity planning to operational automation and intelligent agents that reduce manual intervention. You'll partner closely with infrastructure, research, and product engineering teams to improve reliability, developer productivity, and overall compute utilization.
This role is ideal for engineers who enjoy solving complex operational challenges, building internal platforms, and working on infrastructure that directly powers frontier AI.
In This Role, You Will
Design, build, and operate software that manages large-scale GPU infrastructure supporting ChatGPT inference.
Build internal platforms, tooling, and AI-powered agents that automate fleet operations and reduce operational overhead.
Improve observability, reliability, and operational efficiency across thousands of GPUs.
Develop systems for capacity planning, scheduling, fleet health monitoring, and incident response.
Identify infrastructure bottlenecks and implement solutions that improve utilization, scalability, and performance.
Partner closely with research, platform, networking, and systems teams to continuously improve our compute platform.
Help establish engineering best practices around operational excellence, automation, and infrastructure reliability.
You Might Thrive in This Role If You
Have experience operating large-scale production infrastructure, preferably GPU clusters or other compute-intensive distributed systems.
Have a background in Production Engineering, Site Reliability Engineering (SRE), Infrastructure Engineering, or Platform Engineering.
Have built software that automates operational workflows rather than relying on manual processes.
Have experience with Kubernetes, Linux systems, container orchestration, or distributed infrastructure.
Understand infrastructure observability, monitoring, capacity planning, and incident management.
Enjoy identifying cross-team pain points and building reusable platforms that improve developer productivity.
Are comfortable working across software engineering and systems operations, owning problems end-to-end.
Thrive in fast-moving environments with significant technical ambiguity.
Qualifications
5+ years of software engineering experience building production infrastructure.
Strong programming skills in Go, Python, C++, Rust, or similar systems languages.
Experience designing and operating highly available distributed systems.
Experience with GPU infrastructure, high-performance computing, ML infrastructure, or large-scale compute platforms.
Experience with Kubernetes, cloud infrastructure, Linux, networking, and observability tooling.
Excellent debugging, systems design, and operational problem-solving skills.
Strong communication skills and experience collaborating across engineering organizations.
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We build AI systems that are capable, aligned, and broadly beneficial. Our infrastructure teams power the research and products that bring these systems to millions of users worldwide.
We believe diverse perspectives make stronger teams and better technology. We're committed to creating an inclusive workplace where people from all backgrounds can do their best work.
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity.
We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.
For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement.
Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.
To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance.
We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link.
OpenAI Global Applicant Privacy Policy
At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.
Redirects to OpenAI's application page.
Other roles
More at OpenAI.
Software Engineer, API Safety
Senior
Data Engineer, Monetization Data Platform
Senior
Software Engineer, Plugin Developer Platform
Senior
Product Engineer, Full Stack - Agents
Senior
Engineering Manager, Artifacts
Manager