Software Engineer, Model Routing & Inference

Software Engineer · Senior · Full Time

New YorkUSD 180k – 250k4mo ago
Apply for this role

Opens Cursor's application page

Role

What you'll do.

Cursor is seeking a highly skilled Software Engineer to join their Model Routing & Inference team, responsible for building a cutting-edge inference platform that powers AI interactions across their product. The ideal candidate will design and implement high-performance, reliable distributed systems for routing and serving AI model inference at massive scale.

Responsibilities

  • Inference Platform Development: Build and evolve the inference gateway that provides a unified abstraction over multiple AI model provider APIs, enabling seamless model onboarding and configuration.
  • System Reliability Engineering: Design intelligent cross-provider failover mechanisms to ensure high availability and minimize user-facing service disruptions.
  • Traffic Management: Implement advanced routing backpressure and admission control strategies to manage traffic spikes and optimize system performance.
  • Scalability Optimization: Analyze and optimize GPU utilization, provider economics, and capacity planning to improve system cost-effectiveness and performance.

Qualifications

What we look for.

Technical

  • Distributed Systems

    Extensive experience building high-throughput, low-latency distributed systems with a focus on inference serving or real-time data pipelines

  • Performance Optimization

    Proven ability to reason about complex performance and cost tradeoffs at large-scale infrastructure

Education

  • Computer Science Degree

    Bachelor's or Master's degree in Computer Science, Software Engineering, or related technical field

Experience

  • Production Systems

    Demonstrated track record of shipping robust production systems handling millions of requests

  • AI Infrastructure

    Strong background in AI model inference, routing, and infrastructure technologies

Skills

Required

  • Distributed Systems

    Deep expertise in designing and implementing high-performance distributed systems

  • System Design

    Advanced system design skills with ability to make nuanced architectural decisions

Preferred

  • Machine Learning Infrastructure

    Nice to have

    Experience with AI model serving, routing, and infrastructure technologies

  • Cloud Platforms

    Nice to have

    Familiarity with cloud infrastructure and scalable computing environments

Tech stack

Languages

PythonGo

Frameworks

gRPCKubernetes

Databases

RedisPrometheus

Tools

DockerGrafana

Other

AI Model Providers

Compensation

Pay and benefits.

Base·USD 180,000 – 250,000

Equity·Stock options

Benefits

  • Competitive Compensation

    Highly competitive salary with potential equity compensation

  • Innovative Work Environment

    Opportunity to work on cutting-edge AI infrastructure with a talented, passionate team

  • Professional Growth

    Flat organizational structure with opportunities for rapid skill development and impact

Process

Interview steps.

  1. 01

    Initial Screening

    Preliminary review of candidate's background and qualifications

  2. 02

    Technical Interviews

    2-3 focused technical interviews assessing system design and problem-solving skills

  3. 03

    Onsite Interview

    In-office project work, technical discussions, and team meetings

Full posting

Original listing.

Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.

About the Role

As a Software Engineer on the Model Routing & Inference team at Cursor, you'll build the inference platform that powers every AI interaction in the product.

This team owns the full inference path: making Cursor's AI faster, more reliable, and more cost-effective at a scale few teams in the world get to operate at. Every agent session, every tab completion, and every chat message flows through your stack.

Example projects include...

  • Building and evolving our inference gateway, a single abstraction over every provider's API semantics, so model onboarding becomes a config change.

  • Designing intelligent cross-provider failover so no single provider outage causes user-visible degradation.

  • Designing routing backpressure and admission control so traffic spikes don't cascade into providers.

You may be a fit if

  • You have deep experience building high-throughput, low-latency distributed systems, especially in inference serving, traffic routing, or real-time data pipelines.

  • You're comfortable reasoning about cost/performance tradeoffs at scale (GPU utilization, provider economics, capacity planning).

  • You have strong software engineering fundamentals and enjoy shipping production systems that handle millions of requests.

  • You make good calls in the gray area: weighing reliability, cost, latency, and user experience when there isn't a single "right" answer.

Applying

If there appears to be a fit, we'll reach to schedule 2-3 short technicals. After, we'll schedule an onsite in our office, where you'll work on a small project, discuss ideas, and meet the team.

#LI-DNI

Redirects to Cursor's application page.

Other roles

More at Cursor.

View all 37 roles