Engineering Manager

Manager · Full Time

San FranciscoUSD 200k – 250k1mo ago
Apply for this role

Opens LiteLLM's application page

Role

What you'll do.

LiteLLM is seeking an experienced Engineering Manager to lead stability, engineering productivity, and open-source community growth for their AI gateway platform. The ideal candidate will drive technical excellence, manage team performance, and own critical infrastructure challenges across 100+ provider integrations and emerging AI technologies.

Responsibilities

  • Stability Management: Implement and maintain a comprehensive system for tracking and reducing software regressions, conducting root-cause analysis for each incident
  • Team Performance: Hold engineers accountable to self-defined metrics, ensure team goal achievement, and drive consistent code patterns across the engineering team
  • Incident Response: Manage weekly enterprise war rooms, provide customer updates, handle escalations, and develop resolution strategies for critical issues
  • Team Scaling: Analyze team capacity, define new roles, conduct strategic hiring, and maintain the team's growth trajectory
  • Open-Source Community: Ensure continued open-source adoption and community engagement while scaling the product and team

Qualifications

What we look for.

Technical

  • Technical Leadership

    Proven experience in managing engineering teams and driving technical excellence

  • Incident Management

    Demonstrated ability to reduce incident rates and implement effective root-cause analysis processes

Education

  • Computer Science Degree

    CS/engineering degree or equivalent hands-on engineering background

Experience

  • Engineering Management

    Previous experience as an Engineering Manager with a track record of team management and successful product delivery

  • Startup Experience

    Background in startup environments, familiarity with complex codebases and technical rewrites

Skills

Required

  • Metrics-Driven Management

    Ability to establish and track engineering productivity and reliability metrics

  • Incident Response

    Experience in managing enterprise-level technical escalations and war rooms

Preferred

  • Open-Source Experience

    Nice to have

    Background in managing or contributing to open-source projects

  • Startup Scaling

    Nice to have

    Experience in scaling engineering teams at fast-growing technology companies

Tech stack

Languages

PythonRust

Frameworks

AI Integrations

Tools

Incident Management

Other

Open-Source Infrastructure

Compensation

Pay and benefits.

Base·USD 200,000 – 250,000

Equity·Stock options

Benefits

  • Health Insurance

    Comprehensive health, dental, and vision coverage

  • High-Impact Role

    Direct involvement with leading tech companies and open-source community

  • Career Growth

    Director-level scope with direct line to Head of Engineering

Process

Interview steps.

  1. 01

    Initial Screening

    Review of candidate's engineering management experience and technical background

  2. 02

    Technical Interview

    Deep dive into incident management, team leadership, and technical problem-solving skills

  3. 03

    Leadership Assessment

    Evaluation of management approach, team scaling capabilities, and strategic thinking

  4. 04

    Final Interview

    Meeting with founder/CTO to align on vision and role expectations

Full posting

Original listing.

LiteLLM is the world's most popular AI Gateway, trusted by top companies like Adobe, Netflix, and NASA. Our platform empowers developers by providing secure, reliable access to LLMs and adjacent services, and we're looking for an Engineering Manager to own stability, engineering productivity, hiring, and the open-source community for the leading OSS AI gateway.

About The Role

You own the outcomes, the process, and the team. You won't write all the code, but you own stability, eng productivity, hiring, and the open-source community. You'll drive regressions down release over release and hold them there, hold engineers accountable to the metrics they set for themselves, and own enterprise war rooms. This is director-level scope with a direct line to Head of Engineering, reporting to the founder/CTO. Owning stability of one of the most widely adopted gateways is a genuinely hard problem because the surface is very large: 100+ provider integrations, a Rust migration, and day-zero model launches.

Responsibilities

  • Stand up a system for tracking regressions and run root-cause analysis on every one

  • Delegate fixes to the eng team and own a weekly stability sync

  • Audit stability plans so changes stay gradual, not an overnight rewrite

  • Hold engineers accountable to the metrics they set for themselves, and own the team's goals so they ship what they commit to

  • Take over weekly war rooms: give unhappy customers a plan, daily updates, and own escalations

  • Drive well-architected, consistent code patterns across the team

  • Analyze capacity needs, define the roles, and hire the engineers when more capacity is needed

  • Look out for the open-source community and don't hurt open-source adoption as the team scales

What We're Looking For

  • Been an EM before, with a track record of managing and shipping through a team

  • Has owned product stability with evidence of driving an incident or regression rate down (e.g. took P0s from X to Y)

  • Root-cause and incident management: can point to incidents they ran postmortems on and the recurrence rate dropping after

  • Metrics-driven: has stood up reliability or productivity metrics (regressions per release, MTTR, on-time ship rate) and used them to find root causes

  • Process design: built regression tracking, review cadences, or release processes from scratch, with before/after numbers to show impact

  • Customer-facing incident response: has owned enterprise war rooms and can show trust or retention recovered after

  • Hiring and scaling: has filled a defined headcount plan and can speak to time-to-hire and quality bar

  • CS/engineering degree or equivalent hands-on engineering background

  • Worked at a startup, familiar with messy codebases and doing rewrites

  • Bonus: did this at an open-source company

  • Bonus: scaled the eng team at a fast-growing startup

Why Join LiteLLM?

  • High ownership: you own stability and eng productivity for the leading OSS AI gateway, the metric, the hiring, and the team's execution

  • Director-level scope, clear path: four areas under you and a direct line to Head of Engineering. Real scope now, not after years of waiting

  • Work directly with the CTO: you report to the founder/CTO and help set direction, not execute someone else's plan three layers down

  • A genuinely hard technical problem: 100+ provider integrations, a Rust migration, and day-zero model launches. The reliability challenge is real and interesting, not process theater

  • High impact, in public: we work with NASA, Stripe, Adobe, and 84.51, and the work is open source. You'd be the person who made LiteLLM stable and well-architected, and the community sees it

  • Competitive salary, health, dental, and vision benefits

About LiteLLM (https://github.com/BerriAI/litellm) is a Python SDK and Proxy Server enabling seamless calls to 100+ LLM APIs in the OpenAI format, trusted by industry leaders worldwide.

Ready to make the most widely adopted AI gateway rock-solid?

Redirects to LiteLLM's application page.

Other roles

More at LiteLLM.

View all 7 roles