Senior Backend Engineer

Backend Engineer · Senior · Full Time

San FranciscoUSD 170k – 230k4mo ago
Apply for this role

Opens LiteLLM's application page

Role

What you'll do.

LiteLLM is seeking a Senior Backend Engineer to enhance their AI Gateway platform's guardrails, observability, and logging infrastructure. The ideal candidate will develop robust backend systems that ensure secure, traceable, and high-performance interactions with large language models (LLMs) for enterprise customers.

Responsibilities

  • Backend System Development: Build and scale product infrastructure ensuring high performance and reliability
  • Logging and Traceability: Implement comprehensive logging for guardrail and policy enforcement calls
  • Security Implementation: Design CPU-level guardrails to protect against common LLM API attacks
  • Error Handling: Develop robust error handling mechanisms with transparent user feedback
  • Observability Integration: Configure and maintain monitoring solutions for backend systems handling 1B+ monthly requests

Qualifications

What we look for.

Technical

  • Backend Development

    Expertise in building scalable backend systems with Python

  • Observability

    Advanced knowledge of monitoring and logging platforms

Education

  • Computer Science Degree

    Bachelor's or Master's in Computer Science or related technical field

Experience

  • Backend Engineering

    4+ years of experience with Python and backend web frameworks

  • Systems Observability

    Proven track record in implementing comprehensive logging and monitoring solutions

Skills

Required

  • Python

    Primary backend programming language with extensive experience

  • Backend Frameworks

    Proficiency in FastAPI or Flask for building scalable web services

  • Database Technologies

    Experience with PostgreSQL, Redis, and database integration

  • Observability Tools

    Expertise in Datadog, Splunk, Prometheus, and OpenTelemetry

Preferred

  • AI Infrastructure

    Nice to have

    Understanding of LLM API integration and security

  • Error Handling

    Nice to have

    Advanced techniques for robust logging and error tracing

  • System Performance

    Nice to have

    Experience optimizing high-volume request processing

Tech stack

Languages

Python

Frameworks

FastAPIFlask

Databases

PostgreSQLRedis

Tools

DatadogSplunkPrometheusOpenTelemetry

Other

LLM Integration

Compensation

Pay and benefits.

Base·USD 170,000 – 230,000

Equity·Stock options

Benefits

  • Health Insurance

    Comprehensive medical, dental, and vision coverage

  • Equity

    Stock options in a high-growth AI infrastructure startup

  • Technical Growth

    Opportunity to work on cutting-edge AI and observability technologies

  • Impact

    Contribute to infrastructure used by enterprise customers like Adobe, Netflix, and NASA

Process

Interview steps.

  1. 01

    Initial Screening

    Phone or video call with recruiting team to discuss background and experience

  2. 02

    Technical Assessment

    Coding challenge focused on backend development, logging, and system design

  3. 03

    Technical Interview

    In-depth discussion of system architecture, observability, and AI infrastructure

  4. 04

    System Design Interview

    Evaluate candidate's approach to building scalable, secure backend systems

  5. 05

    Final Interview

    Meeting with engineering leadership to assess team fit and long-term potential

Full posting

Original listing.

Senior Backend Engineer

LiteLLM is the world's most popular AI Gateway, trusted by top companies like Adobe, Netflix, and NASA. Our platform empowers developers by providing secure, reliable access to LLMs and adjacent services, and we're looking for a Backend Engineer (New Grad) to help us build rock-solid guardrails and observability tooling at scale.

About The Role
You’ll focus on owning our guardrails and logging world-class. You will be in charge of the backend code that ensures all guardrail calls are consistently logged, errors are surfaced to users (not silently swallowed), and our observability instrumentation works for real-world, high-volume traffic. Your attention to detail in areas like latency metrics, logging traceability, and backend guardrail registration will directly impact user trust in our security and compliance features.

Responsibilities

  • Build and scale our product, ensuring performance, reliability, and continuous improvement.

  • Ensure all guardrail and policy enforcement calls (e.g., applyguardrail) are properly logged and traceable through our SpendLogs and relevant database tables

  • Build and design CPU-level guardrails to cover common attacks on LLM API's / MCP servers / Agents

  • Identify and fix areas where silent failures occur in guardrail creation, registration, and policy application—ensuring robust error handling and transparency to end users

  • Work with observability integrations, including Datadog, Splunk, Prometheus, and OpenTelemetry, to maintain accurate, configurable, and usable monitoring and logging for backend systems

  • Enhance observability integrations to work for 1B+ requests/mo., with minimal latency overhead and no memory leaks (e.g. due to cardinality of Prometheus metrics)

  • Collaborate cross-functionally on backend engineering priorities (performance, reliability, security)

What We’re Looking For

  • Bachelor’s or Master’s in Computer Science or related field

  • 4+ years of experience with Python and backend frameworks (e.g. FastAPI, Flask)

  • Understanding of logging best practices, error handling, and secure backend development

  • Exposure to monitoring, logging, or metrics platforms (Datadog, Splunk, Prometheus, OpenTelemetry)

  • Familiarity with database integration and troubleshooting (PostgreSQL, Redis, etc.)

  • Driven to deliver high-quality backend code with strong guardrails, auditing, and debugging capabilities

  • Eagerness to tackle hard bugs and ensure system transparency for end users

Why Join LiteLLM?

  • High-impact, mission-critical work on the core of compliance and reliability

  • Contribute directly to features used by enterprise customers at global scale

  • Fast-paced growth environment with room for technical ownership

  • Competitive salary, health, dental, and vision benefits

About LiteLLM
LiteLLM (https://github.com/BerriAI/litellm) is a Python SDK and Proxy Server enabling seamless calls to 100+ LLM APIs in the OpenAI format, trusted by industry leaders worldwide.

Ready to shape the future of secure, observable AI infrastructure? Apply now!

Redirects to LiteLLM's application page.

Other roles

More at LiteLLM.

View all 7 roles