Software Engineer, Cloud Infrastructure
Backend Engineer · Senior · Full Time
Opens OpenAI's application page
Role
What you'll do.
OpenAI is seeking a Senior Software Engineer to join their Cloud Infrastructure team, responsible for building and maintaining the core infrastructure that powers ChatGPT and their API services. The role involves designing scalable Kubernetes-based platforms, cloud abstractions, and ensuring infrastructure can handle massive scale while maintaining reliability and security standards.
Responsibilities
- Infrastructure Platform Design: Design and build scalable development and production platforms that power OpenAI's products including ChatGPT and API services
- Scale Architecture Planning: Ensure infrastructure architecture can scale to the next order of magnitude to support massive AI workloads
- Kubernetes Operations: Operate and maintain large-scale Kubernetes clusters supporting critical AI model inference and training workloads
- Cloud Abstraction Development: Build and maintain cloud infrastructure abstractions that enable rapid product deployment across multiple cloud providers
- System Reliability Engineering: Maintain high availability and reliability standards for production systems supporting millions of users
- Incident Response: Participate in on-call rotation to respond to critical infrastructure incidents and ensure rapid resolution
- Security Implementation: Implement and maintain security best practices across all infrastructure components and deployment pipelines
- Performance Optimization: Monitor and optimize infrastructure performance to support AI model inference at scale
- Cross-team Collaboration: Work closely with research, product, and design teams to enable rapid iteration and deployment of AI technologies
- Infrastructure Automation: Develop automation tools and processes to reduce manual operations and improve deployment reliability
Qualifications
What we look for.
Technical
Infrastructure Experience
5+ years of experience building and maintaining core infrastructure systems at scale
Kubernetes Expertise
Extensive experience operating Kubernetes orchestration systems in production environments
Cloud Platform Proficiency
Deep experience building abstractions and working with major cloud platforms (AWS, GCP, Azure)
System Architecture
Strong background in designing scalable, reliable, and secure distributed systems
Infrastructure as Code
Proficiency with infrastructure automation tools like Terraform, Ansible, or similar
Monitoring and Observability
Experience with monitoring systems, logging, and observability tools for large-scale infrastructure
Network Architecture
Understanding of networking concepts including load balancing, CDNs, and service mesh technologies
Education
Computer Science Degree
Bachelor's degree in Computer Science, Engineering, or equivalent practical experience
Advanced Technical Education
Master's degree in Computer Science or related field preferred but not required
Experience
Senior Infrastructure Role
5+ years in senior infrastructure engineering roles with increasing responsibility
Scale Operations
Experience operating infrastructure supporting millions of users or high-throughput applications
DevOps Culture
Background working in DevOps or SRE environments with emphasis on automation and reliability
Startup Environment
Experience working in fast-paced, high-growth technology companies preferred
Skills
Required
Kubernetes Administration
Expert-level skills in Kubernetes cluster management, networking, and troubleshooting
Cloud Architecture
Advanced knowledge of cloud-native architecture patterns and multi-cloud strategies
Infrastructure Automation
Proficiency in infrastructure as code tools and CI/CD pipeline development
System Design
Strong system design skills for building scalable and reliable distributed systems
Incident Management
Experience with incident response, troubleshooting, and post-mortem analysis
Security Best Practices
Knowledge of infrastructure security, compliance, and vulnerability management
Preferred
AI/ML Infrastructure
Nice to haveExperience with infrastructure supporting machine learning workloads and GPU clusters
Service Mesh
Nice to haveHands-on experience with service mesh technologies like Istio or Linkerd
Observability Platforms
Nice to haveAdvanced experience with observability tools like Prometheus, Grafana, and distributed tracing
Multi-Cloud Strategy
Nice to haveExperience designing and implementing multi-cloud infrastructure strategies
Performance Tuning
Nice to haveSkills in performance optimization for high-throughput, low-latency systems
Open Source Contributions
Nice to haveActive contributions to infrastructure-related open source projects
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 230,000 – 490,000
Equity·Stock options
Benefits
Equity Compensation
Competitive equity package with significant upside potential in a rapidly growing AI company
Health Insurance
Comprehensive medical, dental, and vision insurance coverage
Relocation Assistance
Full relocation support for candidates moving to San Francisco
Professional Development
Access to cutting-edge AI research and learning opportunities
Flexible PTO
Generous time off policy to support work-life balance
Retirement Benefits
401(k) plan with company matching
Parental Leave
Comprehensive parental leave policy for new parents
Mental Health Support
Access to mental health resources and counseling services
Learning Budget
Annual budget for conferences, courses, and professional development
Gym/Wellness
Fitness and wellness benefits including gym membership reimbursement
Process
Interview steps.
- 01
Application Review
Initial screening of resume and technical background by recruiting team
- 02
Recruiter Phone Screen
30-minute call to discuss role fit, experience, and answer initial questions
- 03
Technical Phone Interview
60-minute technical discussion focusing on infrastructure design and Kubernetes experience
- 04
System Design Interview
90-minute session designing scalable infrastructure solutions for AI workloads
- 05
Technical Deep Dive
Detailed technical interview covering cloud platforms, monitoring, and incident response
- 06
Cultural Fit Interview
Discussion with team members about OpenAI's mission, values, and collaborative approach
- 07
Final Interview Round
Meetings with senior leadership and potential team members
- 08
Reference Checks
Professional reference verification and background check process
Full posting
Original listing.
About the Team
The Applied Engineering team works across research, engineering, product, and design to bring OpenAI’s technology to consumers and businesses.
You’ll join the team responsible for running the core infrastructure that supports products like ChatGPT and the API. The systems we support include our kubernetes clusters, infrastructure deployment, our networking stack, cloud abstractions, and more.
We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth.
About the Role
The cloud infrastructure team builds and maintains infrastructure abstractions allowing OpenAI to ship products quickly and scalably.
This role is based in San Francisco, CA.
In this role, you will:
Design and build the development and production platforms that power our products, enabling reliability and security at scale
Ensure our infrastructure can scale to the next order of magnitude
Help create a diverse, equitable, and inclusive culture that makes all feel welcome while enabling radical candor and the challenging of group think
Like all other teams, we are responsible for the reliability of the systems we build. This includes an on-call rotation to respond to critical incidents as needed.
You might thrive in this role if you:
Have 5+ years building core infrastructure
Have experience operating orchestration systems such as Kubernetes at scale
Have experience building abstractions over cloud platforms
Take pride in building and operating scalable, reliable, secure systems
Are comfortable with ambiguity and rapid change
This role is exclusively based in our San Francisco HQ. We offer relocation assistance to new employees.
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity.
We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.
For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement.
Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.
To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance.
We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link.
OpenAI Global Applicant Privacy Policy
At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.
Redirects to OpenAI's application page.
Other roles
More at OpenAI.
Manager, Forward Deployed Engineer (FDE), Life Sciences
Manager
Manager, Forward Deployed Engineer - Tokyo
Manager
Software Engineer, Ads Integrity
Senior
Forward Deployed Engineer - Zurich
Senior
Software Engineer, Conversion Measurement
Senior