Software Engineer, ML Infrastructure

ML Infrastructure Engineer · Senior · Full Time

SF / NYUSD 160k – 250k6mo ago
Apply for this role

Opens Cursor's application page

Role

What you'll do.

Cursor is seeking a Software Engineer to join their ML Infrastructure team, working on large-scale compute, storage, and software infrastructure to support AI-powered coding model development. The role involves collaborating with ML researchers to build high-performance GPU clusters, improve training throughput, and develop systems for workload scheduling and data movement in both SF and NY offices.

Responsibilities

  • ML Training Infrastructure: Collaborate with ML researchers to improve throughput and reliability of large-scale model training
  • GPU Cluster Management: Work with OEMs and cloud providers to plan and build cutting-edge GPU infrastructure
  • Compute Scalability: Improve density and scalability of compute environments for increasingly large RL workloads
  • Infrastructure Automation: Create software and systems to automate building, monitoring, and running GPU clusters
  • Workload Orchestration: Build workload scheduling and data movement systems to support growing training footprint
  • System Reliability: Ensure high availability and performance of distributed ML infrastructure
  • Cross-functional Collaboration: Work closely with ML engineers and researchers to enable their work through infrastructure improvements

Qualifications

What we look for.

Technical

  • Systems Programming

    Strong background in systems and infrastructure-focused software engineering

  • Programming Languages

    Proficiency in Python, TypeScript, Rust, and Go for infrastructure development

  • Distributed Systems

    Experience with distributed storage and networking infrastructure

  • Linux Administration

    Deep knowledge of Linux systems across cloud and bare metal environments

  • Large-scale Infrastructure

    Exposure to systems managing thousands of nodes with significant resource footprints

  • Infrastructure as Code

    Production experience with IaC and configuration management across hosts and Kubernetes

Education

  • Computer Science Degree

    Bachelor's or Master's degree in Computer Science, Engineering, or related technical field

  • Systems Engineering Background

    Strong foundation in computer systems, networking, and distributed computing

Experience

  • Infrastructure Engineering

    5+ years of experience in large-scale infrastructure development

  • Distributed Systems

    Proven experience building and maintaining distributed computing systems

  • ML Infrastructure

    Experience supporting machine learning workloads and training pipelines

  • Cloud and Bare Metal

    Experience with both cloud platforms and bare metal infrastructure management

Skills

Required

  • Python

    Advanced proficiency for ML infrastructure and automation

  • TypeScript

    Strong skills for developer tooling and web interfaces

  • Rust

    Systems programming for high-performance infrastructure

  • Go

    Distributed systems and cloud-native development

  • Linux Systems

    Deep knowledge of Linux administration and networking

  • Distributed Storage

    Experience with large-scale storage systems

  • Kubernetes

    Container orchestration and cluster management

  • Infrastructure as Code

    Terraform, Ansible, or similar tools for automation

Preferred

  • NVIDIA GPU Operations

    Nice to have

    Experience with Blackwell and Hopper-class GPU hardware

  • InfiniBand/RoCE

    Nice to have

    High-performance networking for GPU clusters

  • Ray Framework

    Nice to have

    Distributed computing for ML workloads

  • SLURM

    Nice to have

    Workload management and job scheduling

  • ML Training Pipelines

    Nice to have

    Experience optimizing machine learning training workflows

  • Bare Metal Infrastructure

    Nice to have

    Direct hardware management and optimization

Tech stack

Languages

PythonTypeScriptRustGo

Frameworks

RayKubernetes

Databases

Distributed Storage Systems

Tools

SLURMInfrastructure-as-CodeConfiguration ManagementNVIDIA GPUsInfiniBand/RoCE

Other

Linux SystemsCloud PlatformsBare Metal Infrastructure

Compensation

Pay and benefits.

Base·USD 160,000 – 250,000

Equity·Stock options

Benefits

  • Equity Package

    Significant equity stake in a fast-growing AI company

  • Health Insurance

    Comprehensive medical, dental, and vision coverage

  • In-Person Culture

    Cozy offices in North Beach SF and Manhattan with well-stocked libraries

  • Learning Resources

    Access to cutting-edge research and development in AI/ML

  • Flat Organization

    Direct impact and minimal bureaucracy in a talent-dense team

  • Professional Development

    Opportunity to work with state-of-the-art ML infrastructure

Process

Interview steps.

  1. 01

    Initial Screen

    Phone/video call with hiring manager to discuss background and role alignment

  2. 02

    Technical Interview

    Systems design and infrastructure architecture discussion with engineering team

  3. 03

    Coding Assessment

    Live coding session focused on systems programming and distributed systems

  4. 04

    Cultural Fit

    Conversation about values, working style, and team collaboration

  5. 05

    Final Interview

    Leadership interview covering career goals and company vision alignment

  6. 06

    Reference Check

    Verification of experience and performance with previous employers

Full posting

Original listing.

Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.

We're in-person with cozy offices in North Beach, San Francisco and Manhattan, New York, replete with well-stocked libraries.

About the role

The ML Infrastructure team builds large-scale compute, storage, and software infrastructure to support Cursor’s work building the world’s best agentic coding model. We’re looking for strong engineers who are interested in building high-performance infrastructure and the software to support it. This role works closely with ML researchers and engineers to enable their work through improvements to our training framework, systems reliability/performance, and developer experience.

What you’ll do

  • Collaborate with ML researchers to improve the throughput and reliability of training

  • Work with OEMs, cloud service providers, and others to plan and build cutting-edge GPU infrastructure

  • Improve the density and scalability of compute environments to enable increasingly large RL workloads

  • Create software and systems to automate building, monitoring, and running GPU clusters

  • Build workload scheduling and data movement systems to support Cursor’s growing training footprint

You may be a fit if

  • A strong background in systems and infrastructure-focused software engineering, particularly in Python, Typescript, Rust, and Golang

  • Experience with distributed storage and networking infrastructure, particularly on Linux systems across cloud and bare metal environments

  • Exposure to large-scale systems and their unique challenges, ideally across thousands of nodes with significant resource footprints.

  • Production use of infrastructure-as-code and configuration management, across hosts and Kubernetes

Nice to have

  • Operational exposure to Nvidia GPUs with Infiniband or RoCE, particularly with Blackwell and Hopper-class hardware

  • Exposure to Ray, Slurm, or other common compute and runtime schedulers

Redirects to Cursor's application page.

Other roles

More at Cursor.

View all 36 roles