Software Engineer, Developer Productivity

Backend Engineer · Senior · Full Time

San FranciscoUSD 210k – 490k16mo ago
Apply for this role

Opens OpenAI's application page

Role

What you'll do.

OpenAI's Engineering Acceleration team is seeking a Software Engineer for Developer Productivity to design and build foundational systems that accelerate engineering velocity across ChatGPT and API development. This role involves leveraging cutting-edge AI tools to revolutionize developer productivity while working with large-scale GPU infrastructure and modern cloud technologies.

Responsibilities

  • Developer Tooling Architecture: Drive the design, development, and implementation of tools, systems, and processes that accelerate engineering velocity and reduce manual effort
  • AI-Powered Productivity: Use OpenAI's latest AI tools to re-think and revolutionize team productivity methodologies
  • Cross-Team Collaboration: Work closely with various teams within OpenAI to understand workflows, challenges, and needs for tool development
  • Technical Foundation Building: Partner with product engineers to lay necessary technical foundations for new features and research capabilities
  • Best Practices Guidance: Guide and advise product engineering teams on best practices for ensuring observable, scalable systems
  • System Reliability: Maintain responsibility for system reliability including on-call rotation to respond to critical incidents
  • Infrastructure Optimization: Optimize large-scale GPU node deployments across dozens of Kubernetes clusters in multiple regions
  • CI/CD Pipeline Management: Design and maintain continuous integration and deployment pipelines using Buildkite and related tools

Qualifications

What we look for.

Technical

  • Infrastructure Development

    3+ years of experience building tooling and infrastructure for developer teams

  • Programming Proficiency

    Strong programming skills in Python and experience with FastAPI framework

  • Cloud Infrastructure

    Hands-on experience with Kubernetes, Terraform, and large-scale distributed systems

  • Database Management

    Proficiency with PostgreSQL, Cosmos DB, and Kafka for data processing and storage

  • DevOps Practices

    Experience with CI/CD pipelines, containerization, and automated deployment strategies

  • System Design

    Understanding of scalable system architecture and microservices patterns

Education

  • Technical Degree

    Bachelor's degree in Computer Science, Engineering, or equivalent practical experience

  • Continuous Learning

    Demonstrated ability to rapidly learn new technologies and share knowledge effectively

Experience

  • Software Engineering

    5+ years of overall engineering experience with focus on infrastructure and tooling

  • Developer Productivity

    3+ years specifically in infrastructure building tooling for developer teams

  • Large-Scale Systems

    Experience working with high-traffic, distributed systems and GPU infrastructure

  • On-Call Operations

    Experience with production system maintenance and incident response

Skills

Required

  • Python Programming

    Expert-level proficiency in Python for backend development and automation

  • Kubernetes

    Hands-on experience with container orchestration and cluster management

  • Infrastructure as Code

    Proficiency with Terraform for cloud resource management

  • Database Technologies

    Working knowledge of PostgreSQL, Cosmos DB, and Kafka

  • CI/CD Systems

    Experience with Buildkite or similar continuous integration platforms

  • System Architecture

    Understanding of distributed systems and scalable architecture patterns

Preferred

  • AI/ML Infrastructure

    Nice to have

    Experience with GPU computing and machine learning infrastructure

  • FastAPI Framework

    Nice to have

    Specific experience with FastAPI for high-performance API development

  • Multi-Cloud Deployment

    Nice to have

    Experience with multi-region and multi-cloud infrastructure management

  • Observability Tools

    Nice to have

    Familiarity with monitoring, logging, and alerting systems

  • Performance Optimization

    Nice to have

    Experience optimizing large-scale systems for performance and efficiency

Tech stack

Languages

PythonSQL

Frameworks

FastAPIKubernetes

Databases

PostgreSQLCosmos DBKafka

Tools

TerraformBuildkiteDockerGit

Other

GPU ComputingMulti-region DeploymentObservability Tools

Compensation

Pay and benefits.

Base·USD 210,000 – 490,000

Equity·Stock options

Benefits

  • Equity Compensation

    Competitive equity package as part of total compensation

  • Relocation Assistance

    Full relocation assistance provided for new employees moving to San Francisco

  • Cutting-Edge Technology

    Access to OpenAI's latest AI tools and technologies for professional development

  • Professional Growth

    Opportunity to shape the future of AI technology and developer productivity

  • Equal Opportunity

    Inclusive workplace committed to diversity and equal employment opportunities

  • Reasonable Accommodations

    Support for applicants and employees with disabilities

Process

Interview steps.

  1. 01

    Application Review

    Initial screening of resume and application materials focusing on relevant experience

  2. 02

    Phone/Video Screen

    30-45 minute conversation with recruiter covering background and role interest

  3. 03

    Technical Phone Interview

    60-minute technical discussion covering system design and infrastructure experience

  4. 04

    Technical Deep Dive

    90-minute interview focusing on specific technical challenges and problem-solving approach

  5. 05

    System Design Interview

    60-minute session designing large-scale developer productivity systems

  6. 06

    Team Fit Interview

    45-minute cultural fit assessment with potential teammates

  7. 07

    Final Interview

    60-minute interview with senior leadership covering vision and strategic thinking

  8. 08

    Reference Check

    Verification of background and professional references

Full posting

Original listing.

About the Team

The Applied AI team safely brings OpenAI's technology to the world. We released ChatGPT, Plugins, DALL·E, and the APIs for GPT-4, GPT-3, embeddings, and fine-tuning. We also operate inference infrastructure at scale. There's a lot more on the immediate horizon.

We seek to learn from deployment and distribute the benefits of AI, while ensuring that this powerful tool is used responsibly and safely. Safety is more important to us than unfettered growth. We serve end-users directly through ChatGPT, and serve developers through our APIs, which power product features that were never before possible. 

About the Role

The Engineering Acceleration team designs, builds and maintains the foundational systems that engineers use to build ChatGPT and the API. This is a fast-growing team and you will get a chance to own and define the strategy, vision, and plan for how to increase developer productivity.

In this role, you will:

  • Drive the design, development, and implementation of tools, systems, and processes that accelerate engineering velocity, reduce manual effort, and increase the quality of output.

  • Use our latest AI tools to re-think how we can be the most productive team in the industry.

  • Work closely with various teams within OpenAI to understand their workflows, challenges, and needs, and ensure the tools and systems built by the Engineering Acceleration team address these requirements.

  • Bring new features and research capabilities to the world by partnering with product engineers to lay the necessary technical foundations.

  • Guide and advise product engineering teams on best practices for ensuring observable, scalable systems.

  • Like all other teams, we are responsible for the reliability of the systems we build. This includes an on-call rotation to respond to critical incidents as needed.

You might thrive in this role if you:

  • Have 5+ years of experience in engineering, including 3+ years of experience in infrastructure building tooling for developers.

  • Have experience-driven empathy for the tools, frustrations, and processes that slow engineering teams down and lead to toil or burnout.

  • Have a voracious and intrinsic desire to learn and fill in missing skills—and an equally strong talent for sharing learnings clearly and concisely with others.

  • Are comfortable with ambiguity and rapidly changing conditions. You view changes as an opportunity to add structure and order when necessary.

As technical context: at the heart of our infrastructure is a large-scale deployment of GPU nodes running in dozens of Kubernetes clusters across regions. Some core technologies we build with include Terraform, Buildkite, Postgres, Cosmos DB, Kafka, Python, and FastAPI.

This role is exclusively based in our San Francisco HQ. We offer relocation assistance to new employee.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. 

We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.

For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement.

Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.

To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance.

We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link.

OpenAI Global Applicant Privacy Policy

At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.

Redirects to OpenAI's application page.

Other roles

More at OpenAI.

View all 125 roles