Software Engineer, Online Storage

Backend Engineer · Senior · Full Time

SeattleUSD 230k – 385k15mo ago
Apply for this role

Opens OpenAI's application page

Role

What you'll do.

OpenAI is seeking a Software Engineer for their Online Storage team to build high-performance database systems serving hundreds of millions of users across ChatGPT, Sora, and OpenAI APIs. The role requires 4+ years of experience with distributed systems, systems programming expertise in C++ or Python, and hands-on experience with multi-threading and concurrency for large-scale infrastructure development.

Responsibilities

  • Database System Architecture: Design and build highly scalable, reliable, and performant database systems serving hundreds of millions of users globally
  • API Development: Design and build simple, intuitive APIs for underlying database systems with focus on developer experience
  • Performance Optimization: Analyze and resolve performance and scalability bottlenecks to improve overall system efficiency and response times
  • System Debugging & Maintenance: Debug, instrument, and fix system issues from root cause analysis to delivering long-term solutions
  • Technical Strategy Leadership: Define technical strategy and guide development of robust infrastructure supporting high-scale production systems
  • Cross-functional Collaboration: Work closely with product teams to understand requirements and deliver impactful database solutions
  • Developer Tooling: Build intuitive internal tools and systems that boost engineering productivity across teams
  • System Reliability & On-call: Own reliability of built systems including participation in on-call rotation for critical incident response
  • Infrastructure Scaling: Support evolving business needs through scalable infrastructure design and implementation
  • SLA & KPI Definition: Define and maintain service level agreements and key performance indicators that meet stakeholder expectations

Qualifications

What we look for.

Technical

  • Systems Programming Expertise

    Strong proficiency in systems programming with hands-on multi-threading and concurrency experience

  • C++ or Python Proficiency

    Expert-level programming skills in C++ and/or Python for high-performance system development

  • Distributed Systems Experience

    Hands-on experience with distributed systems including data storage, caching, search, or backend infrastructure

  • Large-scale System Design

    Proven ability to design and build systems that handle massive scale with focus on reliability and performance

  • Database Domain Knowledge

    Preferably domain experience in databases, large-scale data systems, storage, or distributed infrastructure components

  • Performance Engineering

    Experience optimizing system performance and resolving scalability bottlenecks in production environments

Education

  • Computer Science Degree

    Bachelor's degree in Computer Science, Engineering, or equivalent practical experience in systems development

Experience

  • Industry Experience

    4+ years of software engineering experience in production environments

  • Technical Leadership

    2+ years leading large-scale, complex projects or technical initiatives as engineer or tech lead

  • Production System Building

    Experience building and rebuilding production systems to support new product capabilities and growing scale

  • Distributed Systems Implementation

    Hands-on experience implementing distributed systems with focus on reliability, scalability, and security

Skills

Required

  • Systems Programming

    Expert-level systems programming with multi-threading and concurrency

  • C++ or Python

    Proficient in C++ and/or Python for high-performance system development

  • Distributed Systems

    Hands-on experience with distributed data storage, caching, and backend infrastructure

  • Database Systems

    Experience with large-scale database design, optimization, and maintenance

  • Performance Engineering

    Ability to identify and resolve performance bottlenecks in high-scale systems

  • Technical Leadership

    Experience leading complex technical projects and mentoring team members

  • Problem Solving

    Strong analytical and debugging skills for complex system issues

  • Communication

    Excellent communication skills for building consensus across technical and non-technical stakeholders

Preferred

  • Search Infrastructure

    Nice to have

    Experience with search systems like Elasticsearch or similar technologies

  • Caching Systems

    Nice to have

    Knowledge of Redis, Memcached, or other distributed caching solutions

  • Container Orchestration

    Nice to have

    Experience with Kubernetes, Docker, or similar container technologies

  • Monitoring & Observability

    Nice to have

    Familiarity with Prometheus, Grafana, or similar monitoring tools

  • API Design

    Nice to have

    Experience designing RESTful APIs and GraphQL endpoints

  • Cloud Platforms

    Nice to have

    Experience with AWS, GCP, or Azure for large-scale deployments

  • Machine Learning Infrastructure

    Nice to have

    Understanding of ML system requirements and data pipeline architecture

  • Security Engineering

    Nice to have

    Knowledge of security best practices for distributed systems

Tech stack

Languages

C++PythonSQL

Frameworks

Distributed Systems FrameworksDatabase Engine FrameworksAPI Development Frameworks

Databases

Large-scale Database SystemsDistributed Storage SystemsCaching Systems

Tools

Monitoring & Observability ToolsContainer OrchestrationCI/CD PipelinesPerformance Profiling Tools

Other

Multi-threading & ConcurrencySearch InfrastructureLoad BalancingOn-call Systems

Compensation

Pay and benefits.

Base·USD 230,000 – 385,000

Equity·Stock options

Benefits

  • Equity Compensation

    Stock options and equity participation in OpenAI's growth and success

  • Competitive Base Salary

    Industry-leading compensation range of $230K-$385K annually

  • Equal Opportunity Employment

    Inclusive workplace with equal opportunity policies and diversity commitment

  • Reasonable Accommodations

    Support for applicants and employees with disabilities through accommodation requests

  • Mission-Driven Work

    Opportunity to work on AI systems that benefit humanity and shape the future of technology

  • Cutting-edge Technology

    Access to state-of-the-art AI research and development resources

  • Global Impact

    Build systems serving hundreds of millions of users worldwide through ChatGPT and OpenAI APIs

Process

Interview steps.

  1. 01

    Application Review

    Initial screening of resume, GitHub profile, and technical background

  2. 02

    Technical Phone Screen

    45-minute technical discussion covering systems design and programming concepts

  3. 03

    Technical Deep Dive

    90-minute technical interview focusing on distributed systems, database design, and coding problems

  4. 04

    System Design Interview

    60-minute system design session covering large-scale database architecture and scalability

  5. 05

    Behavioral Interview

    45-minute discussion about past experiences, leadership, and cultural fit with OpenAI values

  6. 06

    Final Interview

    Meeting with senior team members and hiring manager for final assessment and questions

  7. 07

    Reference Check

    Verification of past work experience and technical capabilities with previous employers

Full posting

Original listing.

About the Team

We are the Online Storage team powering ChatGPT, Sora, and the OpenAI APIs. We’re a growing team set up to own the databases and online‑storage infrastructure that serve all our products.

About the Role

As OpenAI scales, we’re seeking experienced, problem‑solving engineers to build robust, high‑performance, and scalable database systems. Our ability to rapidly iterate on products while ensuring reliability and speed is key to our success.

You’ll work in a fast‑paced, collaborative environment, building systems that serve hundreds of millions of users globally, with a strong emphasis on safety, reliability, and performance.

We’re hiring skilled software engineers to join the Online Storage team. You’ll help design and build a large‑scale database, collaborate with various product teams to scale it to meet their needs, and own operational excellence by defining SLAs and KPIs that directly satisfy stakeholder expectations. This is a critical role for engineers who thrive on solving complex, large‑scale challenges and are passionate about building resilient systems that perform under load.

 In this role, you will:

  • Design and build highly scalable, reliable, and performant database

  • Design and build highly simple and intuitive APIs for the underlying database

  • Analyze and resolve performance and scalability bottlenecks to improve overall system efficiency

  • Debug, instrument, and fix system issues — from pinpointing root causes to delivering long-term solutions

  • Define technical strategy and guide the development of robust infrastructure that supports high-scale production systems and evolving business needs

  • Collaborate closely with product teams to deeply understand requirements and deliver impactful solutions

  • Boost engineering productivity by building intuitive tools and systems that empower fellow developers

  • Own the reliability of the systems you build, including participating in an on-call rotation to address critical incidents

You might thrive in this role if you:

  • Have experience building (and rebuilding) production systems to support new product capabilities and growing scale

  • Care deeply about the end-user experience and take pride in solving real customer needs

  • Embrace a humble, collaborative mindset and go the extra mile to support your teammates and the broader mission

  • Own problems end-to-end — you're comfortable learning on the fly to fill gaps and get things done

  • Build internal tools that improve workflows when off-the-shelf solutions fall short

  • Have hands-on experience with distributed systems such as data storage, caching, search, or other backend infrastructure components

  • Prioritize the reliability, scalability, and performance of large-scale systems

  • Thrive in ambiguous, fast-paced environments and enjoy iterating rapidly on product and research initiatives

Qualifications:

  • 4+ years of industry experience, including 2+ years leading large-scale, complex projects or technical initiatives as an engineer or tech lead

  • Strong passion for building distributed systems at scale, with a focus on reliability, scalability, security, and continuous improvement

  • Expertise in systems programming, with hands-on experience in multi-threading and concurrency; proficiency in C++ and/or Python is highly preferred

  • Preferably, domain experience in areas such as databases, large-scale data systems, storage, caching, search, or other core components of distributed infrastructure

  • Excellent communication skills, with the ability to build consensus across diverse technical and non-technical stakeholders

.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. 

We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.

For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement.

Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.

To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance.

We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link.

OpenAI Global Applicant Privacy Policy

At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.

Redirects to OpenAI's application page.

Other roles

More at OpenAI.

View all 125 roles