Software Engineer, ML Research

ML Engineer · Senior · Full Time

SF / NYUSD 180k – 300k6mo ago
Apply for this role

Opens Cursor's application page

Role

What you'll do.

Cursor is seeking a Research Engineer to build training, inference, and data systems for frontier coding AI models. The role involves working directly with researchers to develop distributed ML infrastructure and scale reinforcement learning on real user data to automate coding.

Responsibilities

  • Infrastructure Development: Build distributed training, inference, and reinforcement learning infrastructure to support frontier coding models
  • Research Support: Work directly with researchers to make progress repeatable and enable fast iteration cycles
  • Library Development: Write and maintain libraries that simplify large-scale data processing jobs for research teams
  • Data Pipeline Architecture: Architect systems that transform Cursor user data into effective training datasets for ML models
  • System Optimization: Optimize distributed systems performance for large-scale model training and inference workloads
  • End-to-End Ownership: Take full ownership of projects from architecture design through deployment and maintenance
  • Scalability Engineering: Design and implement systems that can handle massive scale data processing and model serving
  • Research Collaboration: Collaborate closely with ML researchers to translate research ideas into production-ready systems

Qualifications

What we look for.

Technical

  • Distributed Systems

    Strong background in building and maintaining large-scale distributed systems

  • Machine Learning Infrastructure

    Experience with ML training pipelines, model serving, and MLOps practices

  • Programming Proficiency

    Expert-level proficiency in Python and systems programming languages like C++ or Rust

  • Cloud Platforms

    Hands-on experience with AWS, GCP, or Azure for ML workloads

  • Language Model Understanding

    Strong intuitions about how large language models work and their training requirements

Education

  • Computer Science Degree

    Bachelor's or Master's degree in Computer Science, Engineering, or related technical field

  • Alternative Experience

    Equivalent practical experience in distributed systems and ML infrastructure

Experience

  • Infrastructure Experience

    5+ years of experience building production-scale distributed systems

  • ML Systems Experience

    3+ years of experience with machine learning infrastructure and training systems

  • End-to-End Delivery

    Proven track record of architecting and shipping complex systems with high ownership

  • Research Environment

    Experience working in fast-paced research environments with rapid iteration requirements

Skills

Required

  • Distributed Systems

    Deep expertise in building scalable distributed systems architecture

  • Python Programming

    Advanced Python skills for ML infrastructure development

  • ML Infrastructure

    Experience with training pipelines, model serving, and MLOps

  • System Architecture

    Ability to design end-to-end systems with high performance requirements

  • Language Models

    Understanding of transformer architectures and training dynamics

Preferred

  • Reinforcement Learning

    Nice to have

    Experience with RL systems and human feedback integration

  • CUDA Programming

    Nice to have

    GPU programming experience for training optimization

  • Research Background

    Nice to have

    Experience working in AI/ML research environments

  • Code Generation

    Nice to have

    Familiarity with code generation models and their unique challenges

  • Data Engineering

    Nice to have

    Large-scale data processing and ETL pipeline experience

Tech stack

Languages

PythonC++CUDA

Frameworks

PyTorchTensorFlowRayKubernetes

Databases

PostgreSQLRedisS3

Tools

DockerMLflowWeights & BiasesApache Spark

Other

AWS/GCPNCCLHorovodApache Kafka

Compensation

Pay and benefits.

Base·USD 180,000 – 300,000

Equity·Stock options

Benefits

  • Equity Package

    Significant equity stake in a fast-growing AI company with strong venture backing

  • Health Insurance

    Comprehensive medical, dental, and vision insurance coverage

  • Office Environment

    Beautiful offices in North Beach San Francisco and Manhattan with well-stocked libraries

  • Learning Budget

    Professional development budget for conferences, courses, and technical resources

  • Flexible PTO

    Unlimited paid time off policy to maintain work-life balance

  • Retirement Benefits

    401(k) plan with company matching contributions

  • Relocation Support

    Relocation assistance for moving to San Francisco or New York offices

Process

Interview steps.

  1. 01

    Initial Screen

    Phone or video call with hiring manager to discuss background and role fit

  2. 02

    Technical Phone Interview

    45-minute technical discussion covering distributed systems and ML infrastructure

  3. 03

    System Design Interview

    Design a large-scale ML training or inference system relevant to Cursor's needs

  4. 04

    Coding Interview

    Live coding session focused on algorithms and data structures

  5. 05

    Research Collaboration Interview

    Discussion with research team about supporting ML research workflows

  6. 06

    Final Interview

    Culture fit and leadership discussion with senior team members

  7. 07

    Reference Checks

    Professional references contacted before final offer

Full posting

Original listing.

Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.

We're in-person with cozy offices in North Beach, San Francisco and Manhattan, New York, replete with well-stocked libraries.

Research Engineer

Cursor is building the future of coding. We train frontier coding agents and scale RL on real user data to make them increasingly effective.

About the role

We’re looking for Research Engineers to build the training, inference, and data systems behind our frontier coding models. You’ll work directly with researchers to make progress repeatable and iteration fast.

What you’ll do

  • Build our distributed training, inference, and RL infrastructure

  • Write libraries to simplify how researchers do large-scale data jobs

  • Architect the systems that turn Cursor user data into effective training data

You may be a fit if

  • You have a strong infrastructure/distributed systems background

  • You are able to architect and ship end-to-end with high ownership

  • You have strong intuitions about how language models work

  • You’re excited to learn more about ML

Redirects to Cursor's application page.

Other roles

More at Cursor.

View all 36 roles