Engineering Manager, Online Data Systems

Engineering Manager · Manager · Full Time

San FranciscoUSD 325k – 405k4mo ago
Apply for this role

Opens OpenAI's application page

Role

What you'll do.

OpenAI is seeking an Engineering Manager to lead their Online Data Systems team, responsible for building and operating hyperscale database and indexing services that power ChatGPT and other AI applications. This role involves managing world-class engineers working on distributed query execution, multi-region federation, and self-healing systems at exabyte scale.

Responsibilities

  • Team Leadership: Build, lead, and grow high-performing infrastructure engineering teams specializing in data systems
  • Technology Strategy: Drive the evolution of OpenAI's in-house online data technologies, including hyperscale database systems and indexing technologies
  • Vector Search Development: Oversee development of advanced vector search capabilities for AI applications
  • Reliability Engineering: Anchor delivery around measurable reliability goals including SLOs to ensure system performance and resiliency
  • Operational Excellence: Reduce operational toil and incident frequency through better abstractions, guardrails, and self-healing systems
  • Agent Technology Integration: Champion pragmatic use of agent technology to amplify execution velocity across the team
  • Distributed Systems Architecture: Oversee delivery of distributed query execution and multi-region federation capabilities
  • Performance Optimization: Lead initiatives for low-level performance optimization at exabyte scale
  • Cross-functional Collaboration: Work with product and research teams to enable rapid development without infrastructure bottlenecks

Qualifications

What we look for.

Technical

  • Database Systems Expertise

    Deep hands-on understanding of database, indexing, storage, and distributed systems technologies

  • Infrastructure Operations

    Hands-on experience building and operating infrastructure with strict reliability, latency, and security requirements

  • Distributed Systems

    Expertise in distributed query execution, multi-region federation, and self-healing systems

  • Performance Engineering

    Experience with low-level performance optimization and hyperscale system design

  • Vector Search

    Understanding of vector databases and similarity search technologies for AI applications

Education

  • Computer Science Degree

    Bachelor's or Master's degree in Computer Science, Engineering, or related technical field preferred

  • Distributed Systems Knowledge

    Strong academic or professional background in distributed systems and database technologies

Experience

  • Engineering Management

    Proven experience managing data-intensive software engineering teams in demanding environments

  • Senior Engineer Development

    Demonstrable track record of hiring, developing, and retaining senior engineers

  • Operational Excellence

    Spirit for operational excellence with hands-on experience in high-reliability systems

  • Technical Leadership

    Ability to modulate technical engagement to create space for highly technical teams to lead and grow

Skills

Required

  • Engineering Leadership

    5+ years managing software engineering teams, particularly in infrastructure or data systems

  • Database Technologies

    Expert-level knowledge of database internals, query optimization, and storage systems

  • Distributed Systems

    Deep understanding of distributed computing, consensus algorithms, and system reliability

  • Performance Engineering

    Experience with low-level optimization, profiling, and scaling systems to extreme loads

  • Operational Excellence

    Track record of building highly reliable systems with strong SLOs and monitoring

Preferred

  • Vector Databases

    Nice to have

    Experience with vector search, embeddings, and AI-specific data storage patterns

  • Multi-cloud Architecture

    Nice to have

    Knowledge of deploying across AWS, GCP, and Azure environments

  • AI/ML Infrastructure

    Nice to have

    Understanding of infrastructure requirements for training and serving AI models

  • Open Source Contributions

    Nice to have

    Active contributions to database or distributed systems open source projects

  • Agent Technology

    Nice to have

    Experience with AI agents and autonomous system orchestration

Tech stack

Languages

PythonGoRustJava/Scala

Frameworks

Apache SparkApache KafkaKubernetes

Databases

Custom In-house DatabasePostgreSQLRedisElasticsearchVector Databases

Tools

DockerTerraformPrometheus/GrafanaApache Airflow

Other

Multi-cloud ArchitectureService MeshCI/CD Pipelines

Compensation

Pay and benefits.

Base·USD 325,000 – 405,000

Equity·Stock options

Benefits

  • Equity Compensation

    Significant equity package in one of the world's leading AI companies

  • Health Insurance

    Comprehensive medical, dental, and vision coverage for employees and families

  • Unlimited PTO

    Flexible time off policy to promote work-life balance

  • Learning Budget

    Professional development funds for conferences, courses, and skill enhancement

  • Remote Work Support

    Stipend for home office setup and equipment

  • Commuter Benefits

    Transportation assistance for San Francisco office commute

  • Wellness Programs

    Mental health support and wellness initiatives

  • Parental Leave

    Generous maternity and paternity leave policies

  • 401k Retirement Plan

    Company matching retirement savings program

Process

Interview steps.

  1. 01

    Initial Screen

    30-minute phone/video call with recruiting team covering background and role alignment

  2. 02

    Hiring Manager Interview

    45-minute discussion with the hiring manager about leadership philosophy and technical vision

  3. 03

    Technical Leadership Interview

    60-minute session covering system design, architecture decisions, and technical problem-solving

  4. 04

    Management Philosophy Interview

    45-minute conversation about team building, performance management, and engineering culture

  5. 05

    Cross-functional Interview

    30-minute interview with stakeholders from product or research teams about collaboration

  6. 06

    Executive Interview

    30-minute final interview with senior leadership about strategic alignment and vision

  7. 07

    Reference Checks

    Verification of past performance and leadership effectiveness with previous colleagues

Full posting

Original listing.

About the Team

The Online Data team builds and operates the core online database and indexing services for OpenAI’s production AI applications, including supporting the explosive growth of ChatGPT, the #1 AI app in the world, and Codex, the fastest growing agentic development toolset in the world.

Our mission is to ensure the reliability, correctness, and scalability of our online data stack and to curate a comprehensive portfolio of services that matches the relentless ambition of OpenAI, enabling our product and research teams to build 0-100 without getting bogged down in the minutiae of multi-region, multi-cloud, exabyte-scale data infrastructure.

About the Role

We are seeking an Engineering Manager to lead our Online Data Systems team, responsible for our in-house database and indexing technology. This role is about shepherding a team of world-class engineers tasked with building and operating hyperscale data storage and retrieval technology.

You’ll be overseeing the delivery of extremely challenging engineering work in areas like distributed query execution, multi-region federation, self-orchestrating and self-healing services, low-level performance optimization, and more.

There are few companies in the world building this kind of technology in-house at this scale where you’ll still be getting in on the ground floor. Instead of being a cog in the machine spending months chasing small optimizations, you’ll play a major part of shaping our future.

In this role, you will:

  • Build, lead, and grow high-performing infrastructure engineering teams.

  • Drive the evolution of OpenAI’s in-house online data technologies, our core, hyper-scale database systems, indexing technologies, and vector search.

  • Anchor delivery around measurable reliability goals (SLOs, etc) to ensure system performance and resiliency is above reproach.

  • Champion pragmatic use of agent technology to amplify execution velocity.

  • Reduce operational toil and incident frequency through better abstractions, guardrails, and self-healing systems.

You might thrive in this role if you:

  • Have a truly insatiable spirit for operational excellence, demonstrated by having your hands-on in building and operating infrastructure with strict reliability, latency, and security requirements.

  • Have experience managing data-intensive software engineering teams in intense, demanding environments.

  • Bring deep hands-on understanding of database, indexing, storage, and distributed systems technologies.

  • Effectively modulate your technical engagement to create space for a highly technical team to lead and grow.

  • Have a demonstrable track record of hiring, developing, and retaining senior engineers.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. 

We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.

For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement.

Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.

To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance.

We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link.

OpenAI Global Applicant Privacy Policy

At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.

Redirects to OpenAI's application page.

Other roles

More at OpenAI.

View all 125 roles