Engineering Manager, MLE

Engineering Manager, ML · Manager · Full Time

San FranciscoUSD 293k – 385k1mo ago
Apply for this role

Opens OpenAI's application page

Role

What you'll do.

Lead OpenAI's Integrity team as an Engineering Manager specializing in machine learning, overseeing the design and deployment of advanced ML models that detect content abuse and prevent scaled attacks on our platforms. This role combines hands-on machine learning expertise with people leadership, requiring deep knowledge of LLM training, transformer architectures, and large-scale ML system deployment to safeguard user experience and platform stability. You'll mentor engineering talent, collaborate with researchers and product teams, and drive technical strategy for mission-critical safety and trust systems.

Responsibilities

  • Lead Machine Learning Architecture & Strategy: Define and drive the technical strategy for the Integrity team's machine learning systems, overseeing the design of advanced classifiers and content understanding models that detect abuse, scaled attacks, and misuse patterns. Make architectural decisions that balance innovation with production reliability, ensuring systems scale safely and effectively as OpenAI's platforms grow.
  • Deploy Advanced ML Models to Production: Lead the design and deployment of state-of-the-art machine learning models from research concept to production implementation. Ensure models meet production requirements for latency, throughput, and accuracy while maintaining robustness against adversarial threats and edge cases that could undermine platform integrity.
  • Mentor & Develop Engineering Talent: Build and mentor a high-performing machine learning engineering team. Conduct technical interviews, provide career development guidance, lead code reviews to maintain engineering standards, and foster knowledge-sharing practices. Develop team members' expertise in LLM training, model optimization, and production ML systems.
  • Optimize & Scale Data Pipelines: Oversee implementation of scalable data infrastructure and pipelines that support model training and inference at scale. Optimize models for both performance (latency, throughput) and accuracy, ensuring efficient resource utilization and cost-effectiveness. Drive continuous improvement in data quality and pipeline reliability.
  • Drive Research-to-Production Excellence: Bridge the gap between cutting-edge machine learning research and production systems. Collaborate closely with research teams to implement research breakthroughs, evaluate novel architectures and approaches, and ensure research innovations translate into tangible improvements in platform trust and safety.
  • Establish & Maintain Production Excellence: Monitor and maintain deployed models to ensure continuous performance, reliability, and safety. Establish monitoring practices, alerting systems, and incident response procedures. Drive a culture of operational excellence that balances innovation velocity with system stability and safety.
  • Cross-Functional Collaboration & Communication: Work closely with product managers, software engineers, security teams, and researchers to understand complex business challenges related to content moderation, user safety, and platform integrity. Translate business requirements into technical specifications and ensure alignment on priorities amid potentially competing demands.
  • Set Technical Standards & Best Practices: Establish coding standards, architectural patterns, and engineering practices that promote high-quality code, reliability, and maintainability across the team. Lead by example in demonstrating technical excellence and commitment to sustainable, scalable engineering practices.

Qualifications

What we look for.

Technical

  • Machine Learning & Deep Learning

    Advanced proficiency in machine learning methodologies, neural network design, and deep learning frameworks. Strong understanding of transformer architectures, attention mechanisms, and state-of-the-art model training techniques applicable to large-scale language models.

  • Python Programming

    Expert-level Python programming skills for ML development, data processing, and system implementation. Ability to write production-quality code with attention to performance optimization and maintainability.

  • LLM Training & Fine-Tuning

    Hands-on expertise with training methodologies including supervised fine-tuning, reinforcement learning from human feedback, policy optimization algorithms, and knowledge distillation. Experience optimizing large language models for specific downstream tasks and safety objectives.

  • Distributed Systems & Scalability

    Understanding of distributed computing, data pipelines, and scalable architecture patterns. Experience optimizing machine learning systems to handle production-scale data volumes and inference latency requirements.

  • Team Leadership & Technical Strategy

    Ability to define technical direction, make architectural decisions, conduct effective code reviews, and establish engineering practices that maintain code quality and system reliability at scale.

Education

  • Master's Degree in Computer Science, Machine Learning, or Related Field

    Advanced degree providing strong theoretical foundation in computer science fundamentals, machine learning theory, algorithms, and statistical methods. Equivalent demonstrated expertise through professional experience may be considered.

  • PhD in Computer Science, Machine Learning, Data Science, or Related Discipline (Preferred)

    Doctoral degree providing deep theoretical knowledge and research experience in machine learning, artificial intelligence, or related fields. Demonstrates expertise in pushing forward the state-of-the-art in AI research and development.

Experience

  • Production Machine Learning Systems

    Significant experience designing, developing, and maintaining machine learning models in production environments. Track record of shipping models that deliver measurable business impact and scale reliably under production demands.

  • Team Leadership & Mentorship

    Demonstrated experience leading technical teams, mentoring engineers, establishing best practices, and driving technical excellence. Proven ability to grow team capabilities and foster a culture of continuous learning.

  • Large Language Model Work

    Hands-on professional experience working with large language models, including training, fine-tuning, evaluation, and deployment. Understanding of LLM capabilities, limitations, and optimization strategies for production use.

  • Cross-Functional Collaboration

    Experience working effectively with product teams, researchers, and stakeholders to translate business objectives into technical solutions. Demonstrated ability to communicate complex technical concepts to non-technical audiences.

Skills

Required

  • Machine Learning Model Development

    Expertise in designing, training, and deploying production-grade machine learning models, including deep learning architectures and large language models (LLMs). Demonstrated ability to translate research concepts into scalable, production-ready solutions with measurable impact on platform safety and user experience.

  • Large Language Model Expertise

    Advanced proficiency with transformer-based models and techniques for training and fine-tuning LLMs, including supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), policy optimization, and knowledge distillation. Experience optimizing models for performance, safety, and inference efficiency at scale.

  • Deep Learning Frameworks

    Expert-level proficiency with PyTorch or TensorFlow for implementing and optimizing neural network architectures. Strong understanding of GPU optimization, distributed training, and model deployment pipelines essential for large-scale ML systems.

  • Software Engineering Fundamentals

    Strong foundation in computer science principles including data structures, algorithms, software architecture patterns, and best practices in code organization. Ability to write clean, maintainable, and efficient code that meets production standards.

  • Engineering Leadership & Team Management

    Proven experience leading and mentoring engineering teams, conducting code reviews, fostering knowledge sharing, establishing engineering best practices, and maintaining high technical standards. Demonstrated ability to scale team capabilities and develop talent.

  • Cross-Functional Collaboration

    Experience working effectively with researchers, software engineers, product managers, and stakeholders to translate business requirements into technical solutions. Strong communication skills for complex technical concepts to diverse audiences.

  • Problem-Solving Under Ambiguity

    Proactive approach to problem-solving with excellent analytical skills. Ability to thrive in fast-paced, loosely-defined environments with competing priorities, taking ownership of end-to-end problems and acquiring knowledge as needed to drive solutions forward.

Preferred

  • Content Understanding & Abuse Prevention

    Nice to have

    Direct experience building machine learning systems for content moderation, abuse detection, adversarial threat detection, or trust and safety applications. Familiarity with classification approaches for detecting harmful content, policy violations, or scaled attack vectors.

  • MLOps & Production ML Systems

    Nice to have

    Experience deploying, monitoring, and maintaining machine learning models in production environments. Knowledge of data pipelines, model versioning, A/B testing frameworks, and continuous monitoring for model performance degradation and safety issues.

  • Research-to-Production Excellence

    Nice to have

    Track record of converting cutting-edge machine learning research into reliable, scalable production systems. Understanding of how to bridge the gap between academic research and enterprise-grade deployment with high reliability requirements.

  • Adversarial Machine Learning

    Nice to have

    Knowledge of adversarial robustness, threat modeling for ML systems, and techniques for building classifiers resilient to malicious inputs or adversarial attacks. Understanding of security considerations specific to AI systems.

Tech stack

Languages

PythonSQL

Frameworks

PyTorchTensorFlowTransformers (Hugging Face)

Databases

Data Warehousing SolutionsVector Databases

Tools

Git & Version ControlMLOps & Monitoring ToolsJupyter & Development EnvironmentsCloud Platforms

Other

Large Language Models (LLMs)Adversarial Robustness & SecurityContent Moderation & Abuse PreventionDistributed Training & GPU Optimization

Compensation

Pay and benefits.

Base·USD 293,000 – 385,000

Equity·Stock options

Benefits

  • Competitive Equity & Stock Options

    Participate in OpenAI's growth through meaningful equity grants, aligning your success with company success as we advance artificial intelligence technology.

  • Comprehensive Health Coverage

    Premium medical, dental, and vision insurance plans covering you and your family, ensuring comprehensive healthcare protection.

  • Retirement & Financial Planning

    401(k) retirement plan with company matching to support long-term financial security and retirement planning.

  • Generous Time Off

    Flexible paid time off (PTO) policy and generous vacation allowance, enabling work-life balance and recovery time.

  • Professional Development & Learning

    Investment in continuous learning through conference attendance, training programs, and access to cutting-edge resources to stay current with AI/ML advancements.

  • Mentorship & Career Growth

    Work with world-class AI researchers and engineers, providing exceptional opportunities for skill development and career advancement in artificial intelligence.

  • Mission-Driven Work

    Contribute directly to OpenAI's mission of ensuring artificial intelligence benefits all of humanity, working on problems with significant societal impact.

  • Collaborative Culture

    Join a dynamic, innovative team where ideas flow freely, creativity is encouraged, and collaboration across disciplines thrives.

Full posting

Original listing.

About the Team

The Integrity team at OpenAI is dedicated to ensuring that our cutting-edge technology is not only revolutionary, but also secure from a myriad of adversarial threats. We strive to maintain the integrity of our platforms as they scale.

The Integrity team is at the front lines of defending against misuse in all its forms: content abuse, scaled attacks, and other actions that could undermine the user experience or harm our operational stability.

About the Role

As a Machine Learning Engineer in OpenAI's Integrity team, you will have the opportunity to work with some of the brightest minds in AI. You’ll work on state-of-the-art models and classifiers, experiment with new architecture and approaches, and push forward our abilities in content and user understanding. You’ll help turn research breakthroughs into tangible solutions that improve the trust and safety of our platform. If you're excited about training LLMs and building ML models, this role is your chance to make a significant mark.

In this role, you will:

  • Innovate and Deploy: Design and deploy advanced machine learning models that solve real-world problems. Bring OpenAI's research from concept to implementation, creating AI-driven applications with a direct impact.

  • Collaborate with the Best: Work closely with researchers, software engineers, and product managers to understand complex business challenges and deliver AI-powered solutions. Be part of a dynamic team where ideas flow freely and creativity thrives.

  • Optimize and Scale: Implement scalable data pipelines, optimize models for performance and accuracy, and ensure they are production-ready. Contribute to projects that require cutting-edge technology and innovative approaches.

  • Learn and Lead: Stay ahead of the curve by engaging with the latest developments in machine learning and AI. Take part in code reviews, share knowledge, and lead by example to maintain high-quality engineering practices.

  • Make a Difference: Monitor and maintain deployed models to ensure they continue delivering value. Your work will directly influence how AI benefits individuals, businesses, and society at large.

You might thrive in this role if you:

  • Master's/ PhD degree in Computer Science, Machine Learning, Data Science, or a related field.

  • Demonstrated experience in deep learning and transformers models

  • Experience with content understanding or abuse prevention with LLMs is a plus

  • Proficiency in frameworks like PyTorch or Tensorflow

  • Strong foundation in data structures, algorithms, and software engineering principles.

  • Are familiar with methods of training and fine-tuning large language models, such as distillation, supervised fine-tuning, and policy optimization

  • Excellent problem-solving and analytical skills, with a proactive approach to challenges.

  • Ability to work collaboratively with cross-functional teams.

  • Ability to move fast in an environment where things are sometimes loosely defined and may have competing priorities or deadlines

  • Enjoy owning the problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. 

We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.

For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement.

Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.

To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance.

We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link.

OpenAI Global Applicant Privacy Policy

At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.

Redirects to OpenAI's application page.

Other roles

More at OpenAI.

View all 102 roles