Software Engineer, ML Platform
Backend Engineer · Senior · Full Time
Opens Cursor's application page
Role
What you'll do.
Join Cursor's ML Platform team as a Software Engineer to build core infrastructure that transforms product usage into better AI models. This role focuses on distributed systems engineering across telemetry, data pipelines, observability, and GPU cluster management for a mission-driven organization automating professional code development. You'll work in a flat, talent-dense environment shipping platform primitives that directly impact ML researchers and product engineers.
Responsibilities
- Design and Build Core Platform Systems: Architect, develop, and maintain distributed systems that serve as the foundation for ML research operations. Design scalable infrastructure including data ingestion pipelines, orchestration systems, and monitoring tools that handle high-volume production workloads with sub-second latencies where required.
- Partner with Research and Product Teams: Collaborate closely with ML researchers, product engineers, and data scientists to understand infrastructure pain points and translate recurring challenges into durable, reusable platform primitives. Participate in technical design reviews and infrastructure planning sessions to ensure alignment with research roadmaps.
- Own System Reliability and Performance: Take end-to-end ownership of reliability, performance optimization, and developer experience for platform systems in your domain. Establish SLOs, implement monitoring and alerting, manage incident response, and continuously measure system health metrics to maintain production excellence.
- Ship Iteratively with High Ownership: Execute rapid development cycles in a flat organizational structure where individual impact is immediate and measurable. Define success metrics for infrastructure improvements, ship features incrementally, gather feedback from platform users, and iterate on designs based on real-world usage patterns.
- Operate Production Distributed Systems at Scale: Manage large-scale distributed infrastructure across GPU fleets, Kubernetes clusters, and cloud/bare-metal environments. Monitor system performance, troubleshoot production issues, optimize resource utilization, and ensure systems remain responsive under varying load conditions.
Qualifications
What we look for.
Technical
Distributed Systems and Infrastructure Engineering
Proven expertise designing and operating production distributed systems at scale. Deep understanding of system architecture patterns, consensus algorithms, fault tolerance, and trade-offs in distributed computing. Experience with production-grade data pipelines, message queues, or streaming platforms.
Production Experience with Large-Scale Infrastructure
Hands-on experience owning and operating high-volume production systems. Background in at least one of: event ingestion and product analytics pipelines, data infrastructure platforms (Spark, Flink, Ray), ML pipeline orchestration, GPU cluster scheduling, or similar large-scale systems.
Linux, Cloud, and Container Orchestration
Strong proficiency with Linux systems administration, cloud platforms (AWS, GCP, Azure), and modern container orchestration tools. Practical experience with Kubernetes, Docker, and infrastructure-as-code tooling. Familiarity with both cloud-native and bare-metal deployment scenarios.
Systems Programming and Backend Development
Strong foundation in backend software engineering with demonstrated ability to write reliable, performant production code. Experience with at least one systems-level language (Go, Rust, C++) or proficiency in Python for systems infrastructure. Understanding of performance profiling, bottleneck identification, and optimization techniques.
Education
Computer Science Foundation
Bachelor's degree in Computer Science, Computer Engineering, or equivalent practical experience demonstrating mastery of core CS concepts including algorithms, data structures, and systems design.
Experience
Ownership of Production Distributed Systems
Multiple years of experience with hands-on ownership and operation of distributed systems handling significant throughput or complexity. Examples include: data ingestion platforms processing millions of events/second, machine learning data pipelines, distributed job scheduling systems, or similar infrastructure-level projects.
Infrastructure and Platform Engineering
Background building developer platforms or infrastructure systems that other engineers depend on daily. Experience understanding and solving for developer experience, reliability requirements, and operational complexity in shared platform systems.
Collaboration with Research or ML Teams
Experience working closely with researchers, data scientists, or other technical stakeholders to understand requirements and translate business/research problems into technical infrastructure solutions. Ability to communicate across different technical backgrounds.
Skills
Required
Distributed Systems Design
Architecture and design of fault-tolerant distributed systems with focus on consistency models, scalability, and operational complexity.
Data Pipeline Infrastructure
Building and operating high-volume data pipelines, ETL systems, or streaming platforms that process production data at scale.
Kubernetes and Container Orchestration
Deep hands-on experience deploying, configuring, and troubleshooting Kubernetes clusters and container-based infrastructure.
Performance Optimization
Identifying and eliminating performance bottlenecks in distributed systems through profiling, benchmarking, and architectural improvements.
Operational Excellence
Experience establishing SLOs, implementing monitoring and alerting, managing incident response, and maintaining production systems reliably.
Backend Development
Strong proficiency in backend programming languages (Go, Rust, Python, Java, or C++) for building and maintaining production systems.
Preferred
Event Ingestion and Analytics Pipelines
Nice to haveDirect experience with product analytics, event collection systems, or OpenTelemetry/distributed tracing frameworks at scale.
Apache Spark, Flink, or Ray
Nice to haveProduction experience with large-scale data processing frameworks used for ML data preparation and experimentation.
GPU Cluster Scheduling and ML Infrastructure
Nice to haveExperience with GPU resource management, job scheduling systems, or machine learning compute infrastructure (SLURM, Kubeflow, Ray clusters).
Observability and Debugging Tools
Nice to haveBuilding or working extensively with monitoring systems, visualization tools, or debugging infrastructure for complex systems.
Machine Learning Research Infrastructure
Nice to haveFamiliarity with ML workflows, experiment tracking, dataset management, or tools commonly used by ML researchers and data scientists.
Bare Metal and Cloud Infrastructure
Nice to haveExperience managing infrastructure across both cloud providers and on-premises bare metal environments for maximum flexibility and cost optimization.
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 180,000 – 280,000
Equity·Stock options
Benefits
Equity and Stock Options
Significant equity stake as part of compensation package, allowing you to participate in company upside as Cursor scales its AI-powered development platform.
Comprehensive Health Coverage
Full medical, dental, and vision insurance benefits covering you and your family with competitive plan options.
401(k) Retirement Planning
Company-sponsored retirement savings plan with employer matching to support long-term financial security.
Unlimited PTO and Wellness
Flexible time-off policy combined with wellness stipends to support mental and physical health.
In-Person Collaboration Spaces
Access to well-stocked offices in San Francisco (North Beach), Palo Alto, and Manhattan with cozy libraries and collaboration environments designed for productive engineering work.
Professional Development
Learning budget and opportunities to attend conferences, take courses, and develop expertise in emerging infrastructure technologies.
Process
Interview steps.
- 01
Application Review and Screening
Initial review of resume, background, and relevant infrastructure engineering experience. The team looks for demonstrated expertise in distributed systems, production-scale deployments, and platform engineering mindset.
- 02
Technical Phone Screening (30-45 minutes)
Focused conversation with a team member about your systems engineering background. Discussion covers specific projects you've owned, technical challenges you've solved, your approach to reliability and performance optimization, and why you're interested in ML Platform infrastructure.
- 03
Architecture and Design Discussion (60 minutes)
Second technical interview diving into system design. You may be asked to discuss how you would design a telemetry pipeline, data platform, observability system, or GPU cluster scheduling solution. Focus is on your architectural thinking, trade-off analysis, and communication of complex technical concepts.
- 04
Systems and Operations Deep Dive (60 minutes)
Third technical interview exploring operational aspects. Discussion of how you've handled production incidents, optimized system performance, designed for reliability, and managed observability. May include questions about specific infrastructure tools and platforms you've used.
- 05
Onsite Interview and Project Work
Full day at one of Cursor's offices (San Francisco, Palo Alto, or Manhattan). You'll work on a small infrastructure project with team members, present your approach and solutions, discuss ideas with the ML Platform team, and meet key stakeholders. This is your opportunity to understand the team dynamics, technical depth, and flat organizational structure firsthand.
- 06
Offer and Negotiation
Following successful interviews, the team extends an offer including base salary, equity, and benefits. Cursor values transparency and is open to negotiation on compensation structure to ensure mutual satisfaction.
Full posting
Original listing.
Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.
About the role
As a Software Engineer on ML Platform at Cursor, you'll build the infrastructure that turns real product usage into better models — and keeps research moving fast on large GPU fleets. ML Platform is organized into four teams. Depending on your background, you may join any of them:
Telemetry — Own the collection and serving path that turns real product use into a record research can trust; without slowing the product, and under a small, explicit policy. Client-side or high-volume ingestion experience is a plus.
ML Data Platform — Build the shared environments and pipeline substrate researchers extend, so new experiments don’t fork their own stack.
Observability — Make it easy for researchers to start, watch, and debug their own runs.
ML DevX and Systems — Shorten the path from idea to a trusted run on the research fleet.
We're looking for strong distributed-systems and infrastructure engineers who want to sit next to research and ship platform primitives that move the product.
We're in-person with cozy offices in North Beach, San Francisco, Palo Alto, and Manhattan, New York, complete with well-stocked libraries.
What you’ll do
Design, build, and operate core platform systems used daily by ML researchers and product engineers
Partner closely with research to turn recurring pain into durable infrastructure
Own reliability, performance, and developer experience for the systems in your lane
Ship iteratively in a flat, high-ownership environment. Measure impact, then raise the bar
You may be a fit if
You have a strong background in systems / infrastructure software engineering and enjoy building platforms other engineers depend on
You've owned production distributed systems at meaningful scale (ingestion, data pipelines, scheduling/orchestration, or similar)
You're comfortable across Linux, cloud and/or bare metal, and modern orchestration (Kubernetes, Ray, or equivalent)
You like working closely with ML researchers and product engineers
You thrive where ownership is high and the feedback loop is short
Especially strong backgrounds by team
Telemetry: event ingestion, product analytics pipelines, OpenTelemetry / tracing, reliable data APIs
Product Data Platform: data frameworks, Spark / Flink / Ray, ML dataset and training-data infrastructure
Observability: experiment / run monitoring, debug and eval tooling, agent-friendly observability UX
ML DevX and Systems: GPU / cluster scheduling, job queues, node health, research compute developer experience
Applying
If there appears to be a fit, we'll reach out to schedule 2-3 short technicals. After, we'll schedule an onsite in our office, where you'll work on a small project, discuss ideas, and meet the team.
Redirects to Cursor's application page.
Other roles
More at Cursor.
Software Engineer, RL Data
Mid
Software Engineer, Pretraining
Senior
Field Engineer, Life Sciences
Mid
Field Engineer - India
Mid
Engineering Manager, Agent & Product Security
Manager