# Engineering Manager, Model Flywheel
**Company:** [OpenAI](https://scaleengineer.com/companies/openai)
Lead engineering excellence for OpenAI's ChatGPT Model Capabilities and Deployment team as an Engineering Manager, overseeing model experimentation, safe deployment automation, and comprehensive measurement systems. This role demands proven experience building and scaling production systems at enterprise scale, with deep expertise in large language model deployment lifecycle, distributed infrastructure, and cross-functional stakeholder management. You'll drive critical technical initiatives that directly impact millions of users while shaping the future of AI safety, reliability, and user experience.
**Role:** Engineering Manager
**Seniority:** Manager
**Locations:** San Francisco
**Salary:** 293000–385000 USD
[Apply](https://jobs.ashbyhq.com/openai/37ee9010-079f-4932-9fc9-45cd9372b6dd)
Canonical: https://scaleengineer.com/jobs/openai/engineering-manager-model-flywheel
---
## Responsibilities

- Lead Model Experimentation Infrastructure: Design, develop, and optimize rapid experimentation frameworks for ChatGPT and Codex product validation. Establish and maintain automated lifecycle management systems that enable safe model testing and evaluation across multiple deployment tiers. Drive innovation in experiment design patterns and automate complex testing workflows to accelerate model iteration cycles.
- Oversee Safe Model Deployment and Rollout Strategy: Architect and lead robust model deployment systems ensuring safe, scalable rollout of new capabilities. Develop sophisticated rollout automation including canary deployments, staged rollouts, and automated rollback mechanisms. Implement operational tooling that provides comprehensive visibility into model performance and system health during and after deployment.
- Manage Capacity Planning and Infrastructure Operations: Build automated capacity management systems that optimize resource utilization across ChatGPT's serving infrastructure. Integrate platform-wide health monitoring and predictive scaling mechanisms. Establish operational runbooks and incident response protocols that maintain system reliability at scale while supporting continuous model innovation.
- Build Comprehensive Model Measurement and Evaluation Systems: Design end-to-end measurement frameworks encompassing model quality evaluation, user signal integration, grader systems, and launch scorecards. Establish feedback loops that connect user metrics, research insights, and product goals. Create visibility dashboards and reporting systems that enable data-driven decision-making across research, product, and engineering teams.
- Elevate ChatGPT's Core Systems and Frameworks: Drive modernization of ChatGPT's harness infrastructure, context management systems, and system prompt frameworks. Consolidate legacy systems and establish scalable abstractions that enable rapid feature development. Lead initiatives to improve multi-tier model experience capabilities, supporting diverse use cases and performance requirements.
- Scale Self-Serve Engineering Capabilities: Expand self-serve experiment platforms enabling cross-functional teams to validate model improvements independently. Implement automated guardrails and safety checks that reduce deployment friction while maintaining rigorous quality standards. Create comprehensive documentation and training that empowers researchers and product teams.
- Lead Cross-Functional Stakeholder Collaboration: Foster strong partnerships with Model Measurement Data Science, Research, Codex, Fleet, Inference, and API teams. Establish clear communication channels, dependency management, and alignment mechanisms. Navigate complex technical tradeoffs and drive consensus on architectural decisions affecting the entire ChatGPT ecosystem.
- Build and Scale High-Performance Engineering Teams: Recruit, mentor, and develop world-class engineers with expertise in distributed systems, machine learning infrastructure, and large-scale systems. Establish strong engineering culture focused on quality, collaboration, and continuous learning. Drive technical career development and create pathways for team members to grow into leadership roles.

## Requirements

### education

- {"name":"Computer Science or Engineering Degree","description":"Bachelor's degree in Computer Science, Computer Engineering, or equivalent field. Advanced degree (Master's or PhD) in related field is a plus but not required with sufficient industry experience."}
- {"name":"Continuous Learning in AI/ML","description":"Commitment to staying current with developments in machine learning, large language models, and AI infrastructure. Active engagement with research papers, open-source projects, and industry knowledge in the AI ecosystem."}

### technical

- {"name":"Large-Scale Distributed Systems Architecture","description":"Demonstrated expertise designing and deploying large-scale distributed systems handling millions of concurrent requests. Deep understanding of service-oriented architecture, load balancing, fault tolerance, and multi-region deployment patterns. Experience optimizing latency, throughput, and reliability in complex infrastructure environments."}
- {"name":"Machine Learning Infrastructure and Deployment","description":"Proven experience with model serving infrastructure, ML pipeline orchestration, and the complete deployment lifecycle from experimentation to production. Understanding of model versioning, A/B testing frameworks for ML, canary deployments, and shadow mode testing for AI systems."}
- {"name":"Large Language Model Systems Knowledge","description":"Working knowledge of transformer-based language models, attention mechanisms, and scaling laws. Experience with LLM-specific challenges including context window management, tokenization, prompt engineering infrastructure, and multi-model systems. Understanding of safety considerations and guardrail implementations for LLM products."}
- {"name":"Production Systems Design at Scale","description":"Track record shipping and maintaining production systems that serve millions of users. Deep understanding of reliability engineering, observability, monitoring, and incident response. Experience managing technical debt and scaling systems through multiple orders of magnitude."}
- {"name":"Experimentation and Measurement Framework Design","description":"Expertise designing and building experimentation platforms supporting A/B testing, multivariate testing, and causal inference. Experience building evaluation frameworks, metric systems, and telemetry pipelines. Understanding of statistical significance, experimental design, and avoiding pitfalls in online experimentation."}
- {"name":"Software Engineering Leadership","description":"Demonstrated success leading engineering teams through complex product cycles and technical challenges. Experience with technical roadmap planning, architecture decision-making, and technical debt management. Proven ability to balance engineering excellence with business velocity and user impact."}

### experience

- {"name":"Engineering Team Leadership","description":"Minimum 5+ years leading engineering teams in complex, cross-functional environments. Success building, mentoring, and scaling teams from 3-4 engineers to 15+ members. Experience working in matrix organizations navigating multiple stakeholders and competing priorities."}
- {"name":"Production Systems at Scale","description":"Demonstrated success shipping and scaling production systems to hundreds of millions of users or billions of transactions. Hands-on experience with reliability engineering, capacity planning, and operational excellence in high-scale environments."}
- {"name":"AI/ML Infrastructure or Backend Services","description":"Strong background in backend systems, distributed infrastructure, or machine learning platforms. Prior experience at companies with significant infrastructure challenges such as hyperscalers, infrastructure companies, or AI research organizations."}
- {"name":"Cross-Functional Collaboration","description":"Proven ability to work effectively with research teams, data scientists, product managers, and infrastructure engineers. Experience translating between technical and non-technical stakeholders and driving alignment on complex technical decisions."}

## Skills

### required

- {"name":"Engineering Leadership","description":"Strategic team leadership with focus on technical direction, team development, and execution excellence. Ability to set clear technical vision while empowering engineers to make autonomous decisions."}
- {"name":"Systems Design and Architecture","description":"Expert-level ability to design scalable, reliable systems handling massive scale. Experience making architectural tradeoffs and designing for operational simplicity."}
- {"name":"Machine Learning Infrastructure","description":"Deep understanding of ML pipelines, model serving, experimentation infrastructure, and the unique challenges of deploying AI systems to production."}
- {"name":"Distributed Systems","description":"Strong foundation in distributed systems concepts, including consistency models, fault tolerance, replication, and coordination."}
- {"name":"Cross-Functional Collaboration","description":"Exceptional ability to work across research, product, and infrastructure teams while maintaining clear communication and alignment."}
- {"name":"Technical Communication","description":"Ability to communicate complex technical concepts clearly to diverse audiences, from engineers to executives to researchers."}

### preferred

- {"name":"Large Language Model Experience","description":"Hands-on experience with large language models, transformer architectures, or similar foundation models in production environments."}
- {"name":"Experimentation Platforms","description":"Experience building or scaling experimentation frameworks and A/B testing platforms for online products or AI systems."}
- {"name":"ML Operations and DevOps","description":"Familiarity with MLOps practices, model deployment pipelines, feature stores, monitoring, and the operational side of machine learning."}
- {"name":"Safety and Reliability Engineering","description":"Experience implementing safety systems, guardrails, and reliability practices for mission-critical systems or AI applications."}
- {"name":"Open Source Contributions","description":"Active participation in open-source projects related to ML infrastructure, distributed systems, or data engineering."}
- {"name":"Research Collaboration","description":"Experience working directly with research teams and understanding how to translate research advances into production systems."}

## Tech stack

### tools

- {"name":"Kubernetes and Container Orchestration","description":"Standard for managing deployed services, orchestrating model serving, and managing infrastructure at scale."}
- {"name":"Prometheus and Grafana","description":"Industry-standard monitoring and observability stack for understanding system behavior and operational health."}
- {"name":"Datadog or New Relic","description":"APM and monitoring platforms providing deep visibility into distributed systems and performance characteristics."}
- {"name":"Git and GitHub","description":"Version control and collaboration platform essential for managing infrastructure code, configuration, and team coordination."}
- {"name":"Terraform or Infrastructure-as-Code","description":"Tools for managing infrastructure declaratively and reproducibly at scale."}
- {"name":"CI/CD Systems (Jenkins, GitLab CI, GitHub Actions)","description":"Automation infrastructure for testing, building, and deploying systems reliably and repeatedly."}

### others

- {"name":"Model Serving Infrastructure","description":"Understanding of vLLM, TensorRT, or similar model serving frameworks optimized for large language models and transformer architectures."}
- {"name":"Experiment Management (Weights & Biases, MLflow)","description":"Tools for tracking experiments, managing model versions, and capturing experiment metadata across large-scale ML workflows."}
- {"name":"Feature Engineering and Feature Stores","description":"Understanding of feature engineering, feature stores (Feast, Tecton), and the operational challenges of managing features at scale."}
- {"name":"Observability and Logging (ELK Stack, Loki)","description":"Distributed logging and tracing systems essential for understanding behavior of complex systems at scale."}
- {"name":"Safety and Guardrail Systems","description":"Experience with content moderation systems, safety filters, and architectural patterns for implementing guardrails in production AI systems."}

### databases

- {"name":"PostgreSQL","description":"Reliable relational database for operational data, experiment metadata, and measurement systems. Standard for reliable data storage at scale."}
- {"name":"Redis","description":"In-memory data store for caching, feature serving, and high-throughput operational needs in model serving infrastructure."}
- {"name":"BigQuery or Snowflake","description":"Data warehousing solutions for analytics, experimentation telemetry, and building comprehensive measurement systems."}
- {"name":"Vector Databases (Pinecone, Weaviate)","description":"Emerging infrastructure for embedding storage and similarity search, increasingly important for modern LLM-based applications."}

### languages

- {"name":"Python","description":"Primary language for ML infrastructure, data processing, and system scripting. Essential for working with ML pipelines and model serving systems."}
- {"name":"TypeScript/JavaScript","description":"Increasingly used for ML infrastructure tooling and backend services. Valuable for experimentation platforms and operational dashboards."}
- {"name":"Go","description":"Common choice for high-performance distributed systems and infrastructure tooling. Used in production infrastructure and deployment automation."}
- {"name":"SQL","description":"Critical for querying telemetry data, analyzing experiments, and building measurement systems. Essential for data-driven decision-making."}

### frameworks

- {"name":"PyTorch","description":"Leading framework for large language models and deep learning. Understanding of PyTorch model serving and deployment considerations important for this role."}
- {"name":"Flask/FastAPI","description":"Common frameworks for building model serving APIs and experimentation platform backends. Critical for understanding inference serving patterns."}
- {"name":"Ray","description":"Distributed computing framework used for ML workloads, model serving, and large-scale data processing."}
- {"name":"Kubernetes","description":"Industry-standard orchestration platform for containerized services. Essential knowledge for operating large-scale ML infrastructure."}

## Benefits

### benefits

- {"name":"Equity Compensation","description":"Meaningful stock options in OpenAI, enabling you to participate in the company's growth and success as a leading AI research and deployment organization."}
- {"name":"Comprehensive Health Insurance","description":"Full medical, dental, and vision coverage with multiple plan options. Includes coverage for preventive care, specialist visits, and prescription medications."}
- {"name":"Mental Health and Wellness","description":"Access to mental health services, therapy, and wellness programs. Subsidized gym memberships, meditation apps, and wellness workshops."}
- {"name":"Flexible Paid Time Off","description":"Unlimited vacation policy enabling you to maintain work-life balance and recharge. Additionally, company holidays and sick leave are fully covered."}
- {"name":"Parental Leave","description":"Generous parental leave policies supporting both mothers and fathers, including adoption and surrogacy coverage."}
- {"name":"Professional Development","description":"Annual education budget for conferences, courses, and certifications. Access to online learning platforms and internal training programs."}
- {"name":"Commuter and Transportation Benefits","description":"Pre-tax commuter benefits for public transportation or parking. Support for remote work setups and necessary equipment."}
- {"name":"Collaborative Office Environment","description":"Access to modern office facilities in San Francisco with collaborative spaces designed for innovation and cross-team interaction."}
- {"name":"Company Events and Culture","description":"Regular team gatherings, company-wide events, and social activities fostering strong community and connection among employees."}
- {"name":"401(k) Retirement Plan","description":"Employer-matched retirement savings plan supporting your long-term financial security."}

## Compensation

- **max:** 420000
- **min:** 280000
- **currency:** USD
- **stockOptions:** true

## Interview process

### steps

## Full description
**About the Team**

The ChatGPT Model Capabilities and Deployment team unified goal is to transform model advancements into great ChatGPT user experiences through reliable serving, rapid experimentation, safe deployment, and continuous improvement. 

**Team Focus Areas**

* **Model Experimentation:**

  * Enable rapid, safe model validation for ChatGPT and Codex products through experiment automation and lifecycle management.
* **Model Deployment:**

  * Ensure safe, scalable deployment of model capabilities with robust rollout and operational tooling.
  * Automate capacity management and incorporate platform-wide health monitors.
* **Model Measurement:**

  * Build comprehensive evaluation and measurement systems for model quality, from user signals to launch scorecards.
  * Improve end-to-end feedback loops for continual model improvement.

**Key Partnerships**

Collaborate cross-functionally with teams including Model Measurement DS, Research, Codex, Fleet, Inference, and API.

**In this role, you will:**

* Elevate and consolidate ChatGPT’s harness, context management, and system prompt frameworks.
* Drive expansion and improvement of multi-tier model experiences.
* Support and scale self-serve experiment capabilities and automated guardrails.
* Lead model rollout automation, capacity management, and health monitoring.
* Shape end-to-end measurement systems (evals, grader signals, user feedback, etc.).

**You might thrive in this role if you have:**

* Proven experience leading engineering teams in complex, cross-functional environments.
* Demonstrated success shipping production systems at scale (ideally for AI or large backend services).
* Deep understanding of model-driven product development, deployment lifecycle, and measurement tooling.
* Excellent communication and collaboration skills—experience interfacing directly with engineering, research, and product stakeholders.
* Prior involvement with large language models, distributed infrastructure, or experimentation platforms is a plus.

**Why Work With Us**

* Tackle highly impactful technical challenges at the cutting edge of AI.
* Collaborate with world-class researchers, engineers, and product leaders.
* Build infrastructure and experiences used by millions.
* Shape the future of how people interact with AI.

If you’re passionate about advancing AI reliability, safety, and user impact at a global scale, we encourage you to apply!

**About OpenAI**

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. 

We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. 

For additional information, please see [OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement](https://cdn.openai.com/policies/eeo-policy-statement.pdf).

Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.

To notify OpenAI that you believe this job posting is non-compliant, please submit a report through [this form](https://form.asana.com/?d=57018692298241&k=5MqR40fZd7jlxVUh5J-UeA). No response will be provided to inquiries unrelated to job posting compliance.

We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this [link](https://form.asana.com/?k=bQ7w9h3iexRlicUdWRiwvg&d=57018692298241).

[OpenAI Global Applicant Privacy Policy](https://cdn.openai.com/policies/global-employee-and-contractor-privacy-policy.pdf)

At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.
