# Software Engineer, ML Platform
**Company:** [Cursor](https://scaleengineer.com/companies/cursor)
Join Cursor's ML Platform team as a Software Engineer to build core infrastructure that transforms product usage into better AI models. This role focuses on distributed systems engineering across telemetry, data pipelines, observability, and GPU cluster management for a mission-driven organization automating professional code development. You'll work in a flat, talent-dense environment shipping platform primitives that directly impact ML researchers and product engineers.
**Role:** Backend Engineer
**Seniority:** Senior
**Locations:** San Francisco
**Salary:** 180000–280000 USD
[Apply](https://jobs.ashbyhq.com/cursor/167f0e93-6915-4d56-803a-be89d1441fb5)
Canonical: https://scaleengineer.com/jobs/cursor/software-engineer-ml-platform
---
## Responsibilities

- Design and Build Core Platform Systems: Architect, develop, and maintain distributed systems that serve as the foundation for ML research operations. Design scalable infrastructure including data ingestion pipelines, orchestration systems, and monitoring tools that handle high-volume production workloads with sub-second latencies where required.
- Partner with Research and Product Teams: Collaborate closely with ML researchers, product engineers, and data scientists to understand infrastructure pain points and translate recurring challenges into durable, reusable platform primitives. Participate in technical design reviews and infrastructure planning sessions to ensure alignment with research roadmaps.
- Own System Reliability and Performance: Take end-to-end ownership of reliability, performance optimization, and developer experience for platform systems in your domain. Establish SLOs, implement monitoring and alerting, manage incident response, and continuously measure system health metrics to maintain production excellence.
- Ship Iteratively with High Ownership: Execute rapid development cycles in a flat organizational structure where individual impact is immediate and measurable. Define success metrics for infrastructure improvements, ship features incrementally, gather feedback from platform users, and iterate on designs based on real-world usage patterns.
- Operate Production Distributed Systems at Scale: Manage large-scale distributed infrastructure across GPU fleets, Kubernetes clusters, and cloud/bare-metal environments. Monitor system performance, troubleshoot production issues, optimize resource utilization, and ensure systems remain responsive under varying load conditions.

## Requirements

### education

- {"name":"Computer Science Foundation","description":"Bachelor's degree in Computer Science, Computer Engineering, or equivalent practical experience demonstrating mastery of core CS concepts including algorithms, data structures, and systems design."}

### technical

- {"name":"Distributed Systems and Infrastructure Engineering","description":"Proven expertise designing and operating production distributed systems at scale. Deep understanding of system architecture patterns, consensus algorithms, fault tolerance, and trade-offs in distributed computing. Experience with production-grade data pipelines, message queues, or streaming platforms."}
- {"name":"Production Experience with Large-Scale Infrastructure","description":"Hands-on experience owning and operating high-volume production systems. Background in at least one of: event ingestion and product analytics pipelines, data infrastructure platforms (Spark, Flink, Ray), ML pipeline orchestration, GPU cluster scheduling, or similar large-scale systems."}
- {"name":"Linux, Cloud, and Container Orchestration","description":"Strong proficiency with Linux systems administration, cloud platforms (AWS, GCP, Azure), and modern container orchestration tools. Practical experience with Kubernetes, Docker, and infrastructure-as-code tooling. Familiarity with both cloud-native and bare-metal deployment scenarios."}
- {"name":"Systems Programming and Backend Development","description":"Strong foundation in backend software engineering with demonstrated ability to write reliable, performant production code. Experience with at least one systems-level language (Go, Rust, C++) or proficiency in Python for systems infrastructure. Understanding of performance profiling, bottleneck identification, and optimization techniques."}

### experience

- {"name":"Ownership of Production Distributed Systems","description":"Multiple years of experience with hands-on ownership and operation of distributed systems handling significant throughput or complexity. Examples include: data ingestion platforms processing millions of events/second, machine learning data pipelines, distributed job scheduling systems, or similar infrastructure-level projects."}
- {"name":"Infrastructure and Platform Engineering","description":"Background building developer platforms or infrastructure systems that other engineers depend on daily. Experience understanding and solving for developer experience, reliability requirements, and operational complexity in shared platform systems."}
- {"name":"Collaboration with Research or ML Teams","description":"Experience working closely with researchers, data scientists, or other technical stakeholders to understand requirements and translate business/research problems into technical infrastructure solutions. Ability to communicate across different technical backgrounds."}

## Skills

### required

- {"name":"Distributed Systems Design","description":"Architecture and design of fault-tolerant distributed systems with focus on consistency models, scalability, and operational complexity."}
- {"name":"Data Pipeline Infrastructure","description":"Building and operating high-volume data pipelines, ETL systems, or streaming platforms that process production data at scale."}
- {"name":"Kubernetes and Container Orchestration","description":"Deep hands-on experience deploying, configuring, and troubleshooting Kubernetes clusters and container-based infrastructure."}
- {"name":"Performance Optimization","description":"Identifying and eliminating performance bottlenecks in distributed systems through profiling, benchmarking, and architectural improvements."}
- {"name":"Operational Excellence","description":"Experience establishing SLOs, implementing monitoring and alerting, managing incident response, and maintaining production systems reliably."}
- {"name":"Backend Development","description":"Strong proficiency in backend programming languages (Go, Rust, Python, Java, or C++) for building and maintaining production systems."}

### preferred

- {"name":"Event Ingestion and Analytics Pipelines","description":"Direct experience with product analytics, event collection systems, or OpenTelemetry/distributed tracing frameworks at scale."}
- {"name":"Apache Spark, Flink, or Ray","description":"Production experience with large-scale data processing frameworks used for ML data preparation and experimentation."}
- {"name":"GPU Cluster Scheduling and ML Infrastructure","description":"Experience with GPU resource management, job scheduling systems, or machine learning compute infrastructure (SLURM, Kubeflow, Ray clusters)."}
- {"name":"Observability and Debugging Tools","description":"Building or working extensively with monitoring systems, visualization tools, or debugging infrastructure for complex systems."}
- {"name":"Machine Learning Research Infrastructure","description":"Familiarity with ML workflows, experiment tracking, dataset management, or tools commonly used by ML researchers and data scientists."}
- {"name":"Bare Metal and Cloud Infrastructure","description":"Experience managing infrastructure across both cloud providers and on-premises bare metal environments for maximum flexibility and cost optimization."}

## Tech stack

### tools

- {"name":"Docker","description":"Container runtime for packaging and deploying services across the ML Platform infrastructure."}
- {"name":"Prometheus","description":"Metrics collection and monitoring system for observability and alerting on platform systems."}
- {"name":"Grafana","description":"Visualization and dashboarding platform for monitoring infrastructure and system health."}
- {"name":"Git and GitHub","description":"Version control and collaboration tools for code management and CI/CD workflows."}
- {"name":"Terraform or Helm","description":"Infrastructure-as-code tools for managing Kubernetes deployments and cloud resource provisioning."}

### others

- {"name":"OpenTelemetry","description":"Open-source observability standards for distributed tracing and instrumentation of complex systems."}
- {"name":"gRPC and Protocol Buffers","description":"High-performance RPC framework and data serialization format for service-to-service communication."}
- {"name":"Linux System Administration","description":"Deep operating system knowledge for performance tuning, networking, and system-level optimization."}
- {"name":"Cloud Platforms (AWS, GCP, Azure)","description":"Production experience with at least one major cloud provider's compute, storage, and networking services."}

### databases

- {"name":"PostgreSQL","description":"Primary relational database for structured data storage and metadata management in platform systems."}
- {"name":"Apache Cassandra or Similar","description":"Time-series or high-throughput databases suitable for storing events, metrics, and observability data at scale."}
- {"name":"Elasticsearch","description":"Distributed search and analytics engine for log aggregation, tracing data, and observability use cases."}

### languages

- {"name":"Go","description":"Primary systems language for building high-performance backend services and infrastructure components."}
- {"name":"Python","description":"Essential for data pipeline work, ML integration, and infrastructure automation commonly used in ML Platform engineering."}
- {"name":"Rust","description":"Used for performance-critical systems requiring memory safety and concurrent performance guarantees."}

### frameworks

- {"name":"Kubernetes","description":"Container orchestration platform for managing distributed ML workloads and research compute infrastructure."}
- {"name":"Apache Spark","description":"Large-scale data processing framework for building distributed ML data pipelines and ETL systems."}
- {"name":"Ray","description":"Distributed computing framework optimized for ML workloads, experiment execution, and cluster resource management."}
- {"name":"Apache Flink","description":"Stream processing framework for real-time data pipelines and event processing at scale."}

## Benefits

### benefits

- {"name":"Equity and Stock Options","description":"Significant equity stake as part of compensation package, allowing you to participate in company upside as Cursor scales its AI-powered development platform."}
- {"name":"Comprehensive Health Coverage","description":"Full medical, dental, and vision insurance benefits covering you and your family with competitive plan options."}
- {"name":"401(k) Retirement Planning","description":"Company-sponsored retirement savings plan with employer matching to support long-term financial security."}
- {"name":"Unlimited PTO and Wellness","description":"Flexible time-off policy combined with wellness stipends to support mental and physical health."}
- {"name":"In-Person Collaboration Spaces","description":"Access to well-stocked offices in San Francisco (North Beach), Palo Alto, and Manhattan with cozy libraries and collaboration environments designed for productive engineering work."}
- {"name":"Professional Development","description":"Learning budget and opportunities to attend conferences, take courses, and develop expertise in emerging infrastructure technologies."}

## Compensation

- **max:** 280000
- **min:** 180000
- **currency:** USD
- **stockOptions:** true

## Interview process

### steps

- {"name":"Application Review and Screening","description":"Initial review of resume, background, and relevant infrastructure engineering experience. The team looks for demonstrated expertise in distributed systems, production-scale deployments, and platform engineering mindset."}
- {"name":"Technical Phone Screening (30-45 minutes)","description":"Focused conversation with a team member about your systems engineering background. Discussion covers specific projects you've owned, technical challenges you've solved, your approach to reliability and performance optimization, and why you're interested in ML Platform infrastructure."}
- {"name":"Architecture and Design Discussion (60 minutes)","description":"Second technical interview diving into system design. You may be asked to discuss how you would design a telemetry pipeline, data platform, observability system, or GPU cluster scheduling solution. Focus is on your architectural thinking, trade-off analysis, and communication of complex technical concepts."}
- {"name":"Systems and Operations Deep Dive (60 minutes)","description":"Third technical interview exploring operational aspects. Discussion of how you've handled production incidents, optimized system performance, designed for reliability, and managed observability. May include questions about specific infrastructure tools and platforms you've used."}
- {"name":"Onsite Interview and Project Work","description":"Full day at one of Cursor's offices (San Francisco, Palo Alto, or Manhattan). You'll work on a small infrastructure project with team members, present your approach and solutions, discuss ideas with the ML Platform team, and meet key stakeholders. This is your opportunity to understand the team dynamics, technical depth, and flat organizational structure firsthand."}
- {"name":"Offer and Negotiation","description":"Following successful interviews, the team extends an offer including base salary, equity, and benefits. Cursor values transparency and is open to negotiation on compensation structure to ensure mutual satisfaction."}

## Full description
Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.

## About the role

As a Software Engineer on **ML Platform** at Cursor, you'll build the infrastructure that turns real product usage into better models — and keeps research moving fast on large GPU fleets. ML Platform is organized into four teams. Depending on your background, you may join any of them:

* **Telemetry** — Own the collection and serving path that turns real product use into a record research can trust; without slowing the product, and under a small, explicit policy. Client-side or high-volume ingestion experience is a plus.
* **ML Data Platform** — Build the shared environments and pipeline substrate researchers extend, so new experiments don’t fork their own stack.
* **Observability** — Make it easy for researchers to start, watch, and debug their own runs.
* **ML DevX and Systems** — Shorten the path from idea to a trusted run on the research fleet.

We're looking for strong distributed-systems and infrastructure engineers who want to sit next to research and ship platform primitives that move the product.

_We're in-person with cozy offices in North Beach, San Francisco, Palo Alto, and Manhattan, New York, complete with well-stocked libraries._

## What you’ll do

* Design, build, and operate core platform systems used daily by ML researchers and product engineers
* Partner closely with research to turn recurring pain into durable infrastructure
* Own reliability, performance, and developer experience for the systems in your lane
* Ship iteratively in a flat, high-ownership environment. Measure impact, then raise the bar

## You may be a fit if

* You have a strong background in systems / infrastructure software engineering and enjoy building platforms other engineers depend on
* You've owned production distributed systems at meaningful scale (ingestion, data pipelines, scheduling/orchestration, or similar)
* You're comfortable across Linux, cloud and/or bare metal, and modern orchestration (Kubernetes, Ray, or equivalent)
* You like working closely with ML researchers and product engineers
* You thrive where ownership is high and the feedback loop is short

## **Especially strong backgrounds by team**

* **Telemetry:** event ingestion, product analytics pipelines, OpenTelemetry / tracing, reliable data APIs
* **Product Data Platform:** data frameworks, Spark / Flink / Ray, ML dataset and training-data infrastructure
* **Observability:** experiment / run monitoring, debug and eval tooling, agent-friendly observability UX
* **ML DevX and Systems:** GPU / cluster scheduling, job queues, node health, research compute developer experience

## Applying

If there appears to be a fit, we'll reach out to schedule 2-3 short technicals. After, we'll schedule an onsite in our office, where you'll work on a small project, discuss ideas, and meet the team.
