Staff Engineer

Staff · Full Time

San FranciscoUSD 240k – 270k3d ago
Apply for this role

Opens LiteLLM's application page

Role

What you'll do.

Staff Software Engineer responsible for owning product quality across LiteLLM's AI Gateway platform, which integrates 100+ provider endpoints. This role focuses on building conformance testing infrastructure, CI/CD systems, and release tooling to ensure reliability at scale while enabling rapid iteration. The ideal candidate brings startup experience, deep debugging expertise, and proven success building quality systems for platforms with extensive external integrations.

Responsibilities

  • Own Product Quality End-to-End: Establish and maintain comprehensive quality standards across LiteLLM's entire platform architecture, driving regression metrics toward zero and establishing the engineering organization's quality culture through technical leadership and strategic priority-setting.
  • Design and Build Provider Conformance Test Suite: Architect and implement a sophisticated, multi-layered test suite covering critical LLM provider integration scenarios including streaming protocols, tool calling semantics, structured output validation, error mapping consistency, cost tracking accuracy, and authentication mode handling across all 100+ supported providers.
  • Build and Own CI/CD Infrastructure: Design, develop, and maintain fast, deterministic continuous integration systems and release tooling with sufficient trust levels that build failures confidently block production deployments, including pipeline optimization, artifact management, and deployment orchestration.
  • Implement Canary and Synthetic Monitoring: Engineer canary deployment systems and synthetic traffic patterns to proactively detect upstream provider API drift, behavior changes, and degradation before they impact production customers, enabling rapid response to provider-side changes.
  • Manage Release Channels Strategy: Architect and implement multi-channel release systems enabling frequent edge deployments for rapid iteration while maintaining stability guarantees for enterprise customers, including feature flags, progressive rollouts, and rollback mechanisms.
  • Execute Root Cause Analysis: Lead thorough investigation of escaped defects in production, identifying systemic quality gaps and converting every incident into permanent test coverage improvements and infrastructure enhancements to prevent recurrence.
  • Raise Engineering Standards Through Technical Excellence: Elevate team quality through hands-on code contributions, thoughtful code reviews, and technical mentorship that establishes higher standards organically, focusing on sustainable quality practices rather than process overhead.

Qualifications

What we look for.

Technical

  • Test Infrastructure Architecture

    Demonstrated expertise building comprehensive test frameworks and conformance testing suites for systems with multiple external dependencies, including experience with property-based testing, fuzzing, and integration testing patterns at scale.

  • CI/CD Pipeline Development

    Strong background designing and implementing continuous integration systems, build orchestration, artifact management, and deployment automation using modern tooling like GitHub Actions, Jenkins, or cloud-native CI platforms.

  • API Integration and Provider Management

    Deep experience working with REST and gRPC APIs, handling distributed provider integrations, managing API versioning, backward compatibility, and error handling across heterogeneous external services.

  • Debugging and Root Cause Analysis

    Advanced debugging proficiency including distributed systems debugging, log analysis, profiling, tracing, and the ability to efficiently navigate unfamiliar codebases to identify true root causes rather than symptoms.

  • Monitoring and Observability

    Experience designing and implementing comprehensive monitoring systems including synthetic monitoring, canary deployments, metrics collection, alerting strategies, and observability tooling for production systems.

Education

  • Computer Science or Related Field

    Bachelor's degree in Computer Science, Software Engineering, or equivalent hands-on professional software development experience demonstrating strong computer science fundamentals.

  • Systems Thinking

    Deep understanding of software systems architecture, distributed systems principles, API design patterns, and the ability to reason about system behavior at scale.

Experience

  • Startup Infrastructure Experience

    3-8 years at a startup environment (Seed through Series D) where you've worn multiple hats, operated with limited resources, and maintained high technical standards while moving quickly with strong startup execution mindset.

  • Large Integration Surface or OSS Contribution Management

    Proven track record building test or release infrastructure for systems with extensive external integration surfaces (100+ integrations), or maintaining CI systems for open source projects with heavy community contributions and complex dependency management.

  • Quality Systems at Scale

    Documented success implementing quality systems that meaningfully improved product reliability without becoming a team bottleneck, including experience mentoring team members on testing best practices and quality mindset.

  • Open Source Contributions

    Active participation in open source communities with public contributions on GitHub or equivalent platforms, demonstrating collaborative development practices and commitment to code quality in transparent environments.

Skills

Required

  • Python

    Expert-level Python proficiency for building test frameworks, CI tooling, and infrastructure automation, with deep understanding of async patterns, testing libraries, and performance optimization.

  • Test Framework Design

    Sophisticated understanding of testing methodologies including unit, integration, contract, and end-to-end testing; expertise in test organization, maintainability, and flakiness elimination.

  • CI/CD Tooling

    Hands-on proficiency with continuous integration platforms (GitHub Actions, GitLab CI, Jenkins, or equivalent), release orchestration, and build system design.

  • Systems Debugging

    Expert debugging skills for complex systems including distributed tracing, log analysis, profiling, and the ability to reason about system behavior under various failure modes.

  • API Integration Patterns

    Deep knowledge of REST API design, error handling, authentication modes, rate limiting, and the challenges of building robust systems on top of unreliable external dependencies.

Preferred

  • TypeScript/Node.js

    Nice to have

    Experience with TypeScript and Node.js ecosystems, particularly for building infrastructure tooling, monitoring systems, and developer-facing tools.

  • Kubernetes and Container Orchestration

    Nice to have

    Familiarity with Kubernetes, Docker, and container orchestration patterns for managing complex deployment scenarios and multi-environment testing.

  • Observability Tools

    Nice to have

    Hands-on experience with monitoring, logging, and tracing tools like Datadog, New Relic, Prometheus, ELK stack, or OpenTelemetry for building production observability systems.

  • LLM/AI Platform Knowledge

    Nice to have

    Understanding of large language model APIs, their behavioral patterns, common integration challenges, provider differences, and the unique testing requirements of AI systems.

  • Infrastructure as Code

    Nice to have

    Experience with Infrastructure as Code tools like Terraform or CloudFormation for managing reproducible, version-controlled infrastructure and environments.

  • Open Source Project Leadership

    Nice to have

    Experience maintaining open source projects, managing community contributions, triaging issues, and building contributor-friendly infrastructure and documentation.

Tech stack

Languages

PythonJavaScript/TypeScript

Frameworks

pytestpytest-asyncioFastAPI

Databases

PostgreSQLRedis

Tools

GitHub ActionsGitDockerDatadog or similar APM

Other

LLM Provider APIsDistributed TracingSynthetic Monitoring PatternsCanary Deployment Systems

Compensation

Pay and benefits.

Base·USD 240,000 – 270,000

Equity·Stock options

Benefits

  • Competitive Equity Package

    Meaningful stock options with early-stage growth potential as part of a Series-funded AI infrastructure company experiencing rapid adoption and market expansion.

  • Comprehensive Health Coverage

    Medical, dental, and vision insurance with company contributions and flexible plan options for you and your family.

  • Unlimited PTO

    Flexible time-off policy reflecting trust in self-management and recognition that sustainable work-life balance drives better engineering outcomes.

  • Professional Development Budget

    Annual budget for conferences, courses, and learning materials to stay current with rapidly evolving AI infrastructure and testing technologies.

  • Technical Mentorship Opportunities

    Work alongside seasoned infrastructure engineers and directly influence architecture decisions at a scale-stage startup with real production demands.

  • Open Source Contribution Time

    Encouraged participation in open source development and LiteLLM's public repository, giving you GitHub portfolio building opportunities.

Process

Interview steps.

  1. 01

    Initial Screening

    Conversation with a technical recruiter or hiring manager to understand your background, motivation for joining LiteLLM, and alignment with the role's focus on quality systems and infrastructure.

  2. 02

    Technical Assessment

    In-depth technical interview covering test architecture design, CI/CD systems thinking, debugging methodology, and your approach to ensuring quality at scale. Likely includes specific scenarios around the 100+ provider integration challenge.

  3. 03

    Systems Design Discussion

    Collaborative conversation about how you would design the conformance test suite, canary deployment system, and release infrastructure for a platform like LiteLLM with multiple external integrations.

  4. 04

    Infrastructure Deep Dive

    Technical discussion with the engineering team about your experience with test flakiness, CI optimization, debugging approaches, and how you've handled quality at scale in previous roles.

  5. 05

    Leadership and Impact Conversation

    Discussion with senior team members about how you raise engineering standards, mentor others, and avoid becoming a bottleneck while owning critical quality systems.

  6. 06

    Executive Alignment

    Optional conversation with company leadership to ensure strategic alignment on LiteLLM's vision, your career goals, and the impact this role plays in the company's growth trajectory.

Full posting

Original listing.

Staff Software Engineer, Product Quality

LiteLLM is the world's most popular AI Gateway, trusted by top companies like Adobe, Netflix, and NASA. Our platform empowers developers with secure, reliable access to LLMs and adjacent services. We're searching for a Staff Software Engineer to own product quality across our testing, CI, and release systems.

About The Role

You will own LiteLLM's product quality across a surface area that spans 100+ provider integrations and a large volume of outside contributions. Your mandate is to make quality automatic at our scale, without slowing down how fast we ship. You'll build the conformance testing, CI, and release infrastructure that catches problems before they reach production, and you'll set the bar on the code paths that matter most.

Responsibilities

  • Own product quality end to end and drive regressions toward zero

  • Build a provider conformance test suite covering streaming, tool calling, structured output, error mapping, cost tracking, and auth modes

  • Own CI and release tooling that is fast, deterministic, and trusted enough that a red gate blocks a ship

  • Build canary and synthetic traffic detection to catch upstream provider drift before customers do

  • Own release channels so we ship daily at the edge while enterprise stays stable

  • Do real root cause analysis and turn every escaped defect into a permanent test

  • Raise the engineering bar through code and review, not through process documents

What We're Looking For

  • Experience at a startup (Seed to Series D) with a willingness to wear multiple hats

  • Built test or release infrastructure for a system with a large external integration surface, or maintained CI for an OSS project with heavy outside contribution

  • Strong debugging depth: can enter an unfamiliar codebase and find real root causes, not just symptoms

  • Opinions about what makes a test suite worth trusting

  • Track record of raising quality on a team without becoming the bottleneck

  • Open source contributions on GitHub or similar

Why Join LiteLLM?

  • Own the quality mandate at a high-impact technical company

  • Deep technical ownership, in the codebase, not in meetings

  • Fast-paced environment with real scale behind it

  • Competitive salary and benefits

About LiteLLM

LiteLLM is a Python SDK and Proxy Server enabling seamless calls to 100+ LLM APIs in the OpenAI format, trusted by industry leaders worldwide.

Ready to own quality at LiteLLM? Apply now!

Redirects to LiteLLM's application page.

Other roles

More at LiteLLM.

View all 7 roles