Software Engineer, ML Infrastructure
ML Infrastructure Engineer · Senior · Full Time
Opens Cursor's application page
Role
What you'll do.
Cursor is seeking a Software Engineer to join their ML Infrastructure team, working on large-scale compute, storage, and software infrastructure to support AI-powered coding model development. The role involves collaborating with ML researchers to build high-performance GPU clusters, improve training throughput, and develop systems for workload scheduling and data movement in both SF and NY offices.
Responsibilities
- ML Training Infrastructure: Collaborate with ML researchers to improve throughput and reliability of large-scale model training
- GPU Cluster Management: Work with OEMs and cloud providers to plan and build cutting-edge GPU infrastructure
- Compute Scalability: Improve density and scalability of compute environments for increasingly large RL workloads
- Infrastructure Automation: Create software and systems to automate building, monitoring, and running GPU clusters
- Workload Orchestration: Build workload scheduling and data movement systems to support growing training footprint
- System Reliability: Ensure high availability and performance of distributed ML infrastructure
- Cross-functional Collaboration: Work closely with ML engineers and researchers to enable their work through infrastructure improvements
Qualifications
What we look for.
Technical
Systems Programming
Strong background in systems and infrastructure-focused software engineering
Programming Languages
Proficiency in Python, TypeScript, Rust, and Go for infrastructure development
Distributed Systems
Experience with distributed storage and networking infrastructure
Linux Administration
Deep knowledge of Linux systems across cloud and bare metal environments
Large-scale Infrastructure
Exposure to systems managing thousands of nodes with significant resource footprints
Infrastructure as Code
Production experience with IaC and configuration management across hosts and Kubernetes
Education
Computer Science Degree
Bachelor's or Master's degree in Computer Science, Engineering, or related technical field
Systems Engineering Background
Strong foundation in computer systems, networking, and distributed computing
Experience
Infrastructure Engineering
5+ years of experience in large-scale infrastructure development
Distributed Systems
Proven experience building and maintaining distributed computing systems
ML Infrastructure
Experience supporting machine learning workloads and training pipelines
Cloud and Bare Metal
Experience with both cloud platforms and bare metal infrastructure management
Skills
Required
Python
Advanced proficiency for ML infrastructure and automation
TypeScript
Strong skills for developer tooling and web interfaces
Rust
Systems programming for high-performance infrastructure
Go
Distributed systems and cloud-native development
Linux Systems
Deep knowledge of Linux administration and networking
Distributed Storage
Experience with large-scale storage systems
Kubernetes
Container orchestration and cluster management
Infrastructure as Code
Terraform, Ansible, or similar tools for automation
Preferred
NVIDIA GPU Operations
Nice to haveExperience with Blackwell and Hopper-class GPU hardware
InfiniBand/RoCE
Nice to haveHigh-performance networking for GPU clusters
Ray Framework
Nice to haveDistributed computing for ML workloads
SLURM
Nice to haveWorkload management and job scheduling
ML Training Pipelines
Nice to haveExperience optimizing machine learning training workflows
Bare Metal Infrastructure
Nice to haveDirect hardware management and optimization
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 160,000 – 250,000
Equity·Stock options
Benefits
Equity Package
Significant equity stake in a fast-growing AI company
Health Insurance
Comprehensive medical, dental, and vision coverage
In-Person Culture
Cozy offices in North Beach SF and Manhattan with well-stocked libraries
Learning Resources
Access to cutting-edge research and development in AI/ML
Flat Organization
Direct impact and minimal bureaucracy in a talent-dense team
Professional Development
Opportunity to work with state-of-the-art ML infrastructure
Process
Interview steps.
- 01
Initial Screen
Phone/video call with hiring manager to discuss background and role alignment
- 02
Technical Interview
Systems design and infrastructure architecture discussion with engineering team
- 03
Coding Assessment
Live coding session focused on systems programming and distributed systems
- 04
Cultural Fit
Conversation about values, working style, and team collaboration
- 05
Final Interview
Leadership interview covering career goals and company vision alignment
- 06
Reference Check
Verification of experience and performance with previous employers
Full posting
Original listing.
Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.
We're in-person with cozy offices in North Beach, San Francisco and Manhattan, New York, replete with well-stocked libraries.
About the role
The ML Infrastructure team builds large-scale compute, storage, and software infrastructure to support Cursor’s work building the world’s best agentic coding model. We’re looking for strong engineers who are interested in building high-performance infrastructure and the software to support it. This role works closely with ML researchers and engineers to enable their work through improvements to our training framework, systems reliability/performance, and developer experience.
What you’ll do
Collaborate with ML researchers to improve the throughput and reliability of training
Work with OEMs, cloud service providers, and others to plan and build cutting-edge GPU infrastructure
Improve the density and scalability of compute environments to enable increasingly large RL workloads
Create software and systems to automate building, monitoring, and running GPU clusters
Build workload scheduling and data movement systems to support Cursor’s growing training footprint
You may be a fit if
A strong background in systems and infrastructure-focused software engineering, particularly in Python, Typescript, Rust, and Golang
Experience with distributed storage and networking infrastructure, particularly on Linux systems across cloud and bare metal environments
Exposure to large-scale systems and their unique challenges, ideally across thousands of nodes with significant resource footprints.
Production use of infrastructure-as-code and configuration management, across hosts and Kubernetes
Nice to have
Operational exposure to Nvidia GPUs with Infiniband or RoCE, particularly with Blackwell and Hopper-class hardware
Exposure to Ray, Slurm, or other common compute and runtime schedulers
Redirects to Cursor's application page.
Other roles
More at Cursor.
Field Engineer - India
Mid
Engineering Manager, Agent & Product Security
Manager
Engineering Manager, ML
Manager
Regional Director, Field Engineering, Healthcare
Manager
RVP, Field Engineering, Public Sector
VP