Staff Engineer
Staff · Full Time
Opens LiteLLM's application page
Role
What you'll do.
Staff Software Engineer responsible for owning product quality across LiteLLM's AI Gateway platform, which integrates 100+ provider endpoints. This role focuses on building conformance testing infrastructure, CI/CD systems, and release tooling to ensure reliability at scale while enabling rapid iteration. The ideal candidate brings startup experience, deep debugging expertise, and proven success building quality systems for platforms with extensive external integrations.
Responsibilities
- Own Product Quality End-to-End: Establish and maintain comprehensive quality standards across LiteLLM's entire platform architecture, driving regression metrics toward zero and establishing the engineering organization's quality culture through technical leadership and strategic priority-setting.
- Design and Build Provider Conformance Test Suite: Architect and implement a sophisticated, multi-layered test suite covering critical LLM provider integration scenarios including streaming protocols, tool calling semantics, structured output validation, error mapping consistency, cost tracking accuracy, and authentication mode handling across all 100+ supported providers.
- Build and Own CI/CD Infrastructure: Design, develop, and maintain fast, deterministic continuous integration systems and release tooling with sufficient trust levels that build failures confidently block production deployments, including pipeline optimization, artifact management, and deployment orchestration.
- Implement Canary and Synthetic Monitoring: Engineer canary deployment systems and synthetic traffic patterns to proactively detect upstream provider API drift, behavior changes, and degradation before they impact production customers, enabling rapid response to provider-side changes.
- Manage Release Channels Strategy: Architect and implement multi-channel release systems enabling frequent edge deployments for rapid iteration while maintaining stability guarantees for enterprise customers, including feature flags, progressive rollouts, and rollback mechanisms.
- Execute Root Cause Analysis: Lead thorough investigation of escaped defects in production, identifying systemic quality gaps and converting every incident into permanent test coverage improvements and infrastructure enhancements to prevent recurrence.
- Raise Engineering Standards Through Technical Excellence: Elevate team quality through hands-on code contributions, thoughtful code reviews, and technical mentorship that establishes higher standards organically, focusing on sustainable quality practices rather than process overhead.
Qualifications
What we look for.
Technical
Test Infrastructure Architecture
Demonstrated expertise building comprehensive test frameworks and conformance testing suites for systems with multiple external dependencies, including experience with property-based testing, fuzzing, and integration testing patterns at scale.
CI/CD Pipeline Development
Strong background designing and implementing continuous integration systems, build orchestration, artifact management, and deployment automation using modern tooling like GitHub Actions, Jenkins, or cloud-native CI platforms.
API Integration and Provider Management
Deep experience working with REST and gRPC APIs, handling distributed provider integrations, managing API versioning, backward compatibility, and error handling across heterogeneous external services.
Debugging and Root Cause Analysis
Advanced debugging proficiency including distributed systems debugging, log analysis, profiling, tracing, and the ability to efficiently navigate unfamiliar codebases to identify true root causes rather than symptoms.
Monitoring and Observability
Experience designing and implementing comprehensive monitoring systems including synthetic monitoring, canary deployments, metrics collection, alerting strategies, and observability tooling for production systems.
Education
Computer Science or Related Field
Bachelor's degree in Computer Science, Software Engineering, or equivalent hands-on professional software development experience demonstrating strong computer science fundamentals.
Systems Thinking
Deep understanding of software systems architecture, distributed systems principles, API design patterns, and the ability to reason about system behavior at scale.
Experience
Startup Infrastructure Experience
3-8 years at a startup environment (Seed through Series D) where you've worn multiple hats, operated with limited resources, and maintained high technical standards while moving quickly with strong startup execution mindset.
Large Integration Surface or OSS Contribution Management
Proven track record building test or release infrastructure for systems with extensive external integration surfaces (100+ integrations), or maintaining CI systems for open source projects with heavy community contributions and complex dependency management.
Quality Systems at Scale
Documented success implementing quality systems that meaningfully improved product reliability without becoming a team bottleneck, including experience mentoring team members on testing best practices and quality mindset.
Open Source Contributions
Active participation in open source communities with public contributions on GitHub or equivalent platforms, demonstrating collaborative development practices and commitment to code quality in transparent environments.
Skills
Required
Python
Expert-level Python proficiency for building test frameworks, CI tooling, and infrastructure automation, with deep understanding of async patterns, testing libraries, and performance optimization.
Test Framework Design
Sophisticated understanding of testing methodologies including unit, integration, contract, and end-to-end testing; expertise in test organization, maintainability, and flakiness elimination.
CI/CD Tooling
Hands-on proficiency with continuous integration platforms (GitHub Actions, GitLab CI, Jenkins, or equivalent), release orchestration, and build system design.
Systems Debugging
Expert debugging skills for complex systems including distributed tracing, log analysis, profiling, and the ability to reason about system behavior under various failure modes.
API Integration Patterns
Deep knowledge of REST API design, error handling, authentication modes, rate limiting, and the challenges of building robust systems on top of unreliable external dependencies.
Preferred
TypeScript/Node.js
Nice to haveExperience with TypeScript and Node.js ecosystems, particularly for building infrastructure tooling, monitoring systems, and developer-facing tools.
Kubernetes and Container Orchestration
Nice to haveFamiliarity with Kubernetes, Docker, and container orchestration patterns for managing complex deployment scenarios and multi-environment testing.
Observability Tools
Nice to haveHands-on experience with monitoring, logging, and tracing tools like Datadog, New Relic, Prometheus, ELK stack, or OpenTelemetry for building production observability systems.
LLM/AI Platform Knowledge
Nice to haveUnderstanding of large language model APIs, their behavioral patterns, common integration challenges, provider differences, and the unique testing requirements of AI systems.
Infrastructure as Code
Nice to haveExperience with Infrastructure as Code tools like Terraform or CloudFormation for managing reproducible, version-controlled infrastructure and environments.
Open Source Project Leadership
Nice to haveExperience maintaining open source projects, managing community contributions, triaging issues, and building contributor-friendly infrastructure and documentation.
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 240,000 – 270,000
Equity·Stock options
Benefits
Competitive Equity Package
Meaningful stock options with early-stage growth potential as part of a Series-funded AI infrastructure company experiencing rapid adoption and market expansion.
Comprehensive Health Coverage
Medical, dental, and vision insurance with company contributions and flexible plan options for you and your family.
Unlimited PTO
Flexible time-off policy reflecting trust in self-management and recognition that sustainable work-life balance drives better engineering outcomes.
Professional Development Budget
Annual budget for conferences, courses, and learning materials to stay current with rapidly evolving AI infrastructure and testing technologies.
Technical Mentorship Opportunities
Work alongside seasoned infrastructure engineers and directly influence architecture decisions at a scale-stage startup with real production demands.
Open Source Contribution Time
Encouraged participation in open source development and LiteLLM's public repository, giving you GitHub portfolio building opportunities.
Process
Interview steps.
- 01
Initial Screening
Conversation with a technical recruiter or hiring manager to understand your background, motivation for joining LiteLLM, and alignment with the role's focus on quality systems and infrastructure.
- 02
Technical Assessment
In-depth technical interview covering test architecture design, CI/CD systems thinking, debugging methodology, and your approach to ensuring quality at scale. Likely includes specific scenarios around the 100+ provider integration challenge.
- 03
Systems Design Discussion
Collaborative conversation about how you would design the conformance test suite, canary deployment system, and release infrastructure for a platform like LiteLLM with multiple external integrations.
- 04
Infrastructure Deep Dive
Technical discussion with the engineering team about your experience with test flakiness, CI optimization, debugging approaches, and how you've handled quality at scale in previous roles.
- 05
Leadership and Impact Conversation
Discussion with senior team members about how you raise engineering standards, mentor others, and avoid becoming a bottleneck while owning critical quality systems.
- 06
Executive Alignment
Optional conversation with company leadership to ensure strategic alignment on LiteLLM's vision, your career goals, and the impact this role plays in the company's growth trajectory.
Full posting
Original listing.
Staff Software Engineer, Product Quality
LiteLLM is the world's most popular AI Gateway, trusted by top companies like Adobe, Netflix, and NASA. Our platform empowers developers with secure, reliable access to LLMs and adjacent services. We're searching for a Staff Software Engineer to own product quality across our testing, CI, and release systems.
About The Role
You will own LiteLLM's product quality across a surface area that spans 100+ provider integrations and a large volume of outside contributions. Your mandate is to make quality automatic at our scale, without slowing down how fast we ship. You'll build the conformance testing, CI, and release infrastructure that catches problems before they reach production, and you'll set the bar on the code paths that matter most.
Responsibilities
Own product quality end to end and drive regressions toward zero
Build a provider conformance test suite covering streaming, tool calling, structured output, error mapping, cost tracking, and auth modes
Own CI and release tooling that is fast, deterministic, and trusted enough that a red gate blocks a ship
Build canary and synthetic traffic detection to catch upstream provider drift before customers do
Own release channels so we ship daily at the edge while enterprise stays stable
Do real root cause analysis and turn every escaped defect into a permanent test
Raise the engineering bar through code and review, not through process documents
What We're Looking For
Experience at a startup (Seed to Series D) with a willingness to wear multiple hats
Built test or release infrastructure for a system with a large external integration surface, or maintained CI for an OSS project with heavy outside contribution
Strong debugging depth: can enter an unfamiliar codebase and find real root causes, not just symptoms
Opinions about what makes a test suite worth trusting
Track record of raising quality on a team without becoming the bottleneck
Open source contributions on GitHub or similar
Why Join LiteLLM?
Own the quality mandate at a high-impact technical company
Deep technical ownership, in the codebase, not in meetings
Fast-paced environment with real scale behind it
Competitive salary and benefits
About LiteLLM
LiteLLM is a Python SDK and Proxy Server enabling seamless calls to 100+ LLM APIs in the OpenAI format, trusted by industry leaders worldwide.
Ready to own quality at LiteLLM? Apply now!
Redirects to LiteLLM's application page.
Other roles
More at LiteLLM.
Security Engineer
Mid
Senior Support Engineer - India
Senior
Senior Support Engineer - Brazil
Senior
Support Engineer
Mid
Rust Engineer
Senior