# Engineering Manager, ML
**Company:** [Cursor](https://scaleengineer.com/companies/cursor)
Lead a high-impact ML infrastructure team at Cursor, where you'll own the training, testing, and evaluation systems that power the world's leading AI-assisted coding platform. This role sits at the intersection of infrastructure and model behavior, requiring both strong technical leadership and hands-on systems expertise to build scalable ML training pipelines, robust evaluation frameworks, and optimized environments for model iteration at scale.
**Role:** Engineering Manager
**Seniority:** Manager
**Locations:** San Francisco
**Salary:** 185000–280000 USD
[Apply](https://jobs.ashbyhq.com/cursor/3752886b-a6e2-402c-ba2a-47d0659ff335)
Canonical: https://scaleengineer.com/jobs/cursor/engineering-manager-ml
---
## Responsibilities

- ML Infrastructure Leadership & Technical Direction: Set and execute technical strategy for model training, testing, and evaluation infrastructure at scale. Make critical architectural decisions that balance latency, quality, and cost tradeoffs. Own the codebase alongside your team, conducting technical reviews and debugging complex distributed systems issues that span both infrastructure and model behavior boundaries.
- Reinforcement Learning & Evaluation Infrastructure: Design and implement rollout infrastructure enabling researchers to conduct large-scale RL experiments efficiently. Build evaluation pipelines that catch regressions before production deployment and provide rapid, trustworthy signal on experimental changes. Create reproducible, sandboxed training and testing environments optimized for iteration speed.
- Team Leadership & Engineering Development: Source, interview, and hire exceptional infrastructure engineers aligned with Cursor's mission and values. Develop engineers through mentorship, code review, coaching, and strategic project assignments. Build a small, talent-dense team capable of operating with high autonomy in ambiguous environments while maintaining rigorous engineering standards.
- Cross-Functional Collaboration & Communication: Partner closely with research teams to translate model-level requirements into concrete infrastructure solutions. Communicate fluently across both research and engineering domains, identifying technical ownership boundaries and ensuring clarity on whether issues stem from systems design or model behavior. Collaborate with product and research to prioritize infrastructure investments.
- Measurement & Rigor: Establish rigorous measurement frameworks for infrastructure quality and team progress, especially in areas where shipping code doesn't guarantee impact. Instrument systems for observability, implement monitoring for reliability and performance under production load. Drive data-driven decision-making about infrastructure investments and prioritization.

## Requirements

### education

- {"name":"Computer Science or Related Field","description":"Bachelor's degree in Computer Science, Computer Engineering, Mathematics, or equivalent professional experience demonstrating advanced technical depth in systems design and software engineering."}

### technical

- {"name":"ML Infrastructure & Training Frameworks","description":"Production experience building or leading infrastructure for model training, evaluation, or serving systems. Deep knowledge of distributed training frameworks, experiment tracking, model checkpoint management, and optimization for large-scale ML workloads."}
- {"name":"Distributed Systems & Reliability Engineering","description":"Strong fundamentals in distributed systems, including consistency, fault tolerance, and performance optimization. Proven ability to design systems that perform reliably under real production load, not just in theory. Experience with containerization, orchestration platforms, or microservices architecture."}
- {"name":"Software Engineering Excellence","description":"Demonstrated ability to write clean, maintainable production code and conduct thorough technical code reviews. Comfortable with infrastructure-as-code principles, CI/CD pipelines, and modern development workflows. Proficiency in at least one systems-level programming language (e.g., Python, Go, Rust, C++)."}
- {"name":"Debugging Complex Systems","description":"Strong debugging skills for production systems spanning multiple layers of abstraction. Ability to trace issues from application code through infrastructure to identify root causes. Experience with observability tools, profiling, and performance analysis."}

### experience

- {"name":"ML Infrastructure Leadership","description":"Led engineering teams building infrastructure for training, evaluating, or serving machine learning models in production environments. Track record of shipping infrastructure that unblocked research or production teams and improved iteration velocity."}
- {"name":"Team Building & Development","description":"Proven track record hiring and developing infrastructure engineers who have grown into senior roles. Experience mentoring engineers through both technical skill development and career growth. Demonstrated ability to build high-performing, autonomous teams."}
- {"name":"Production Systems at Scale","description":"Hands-on experience owning or building production systems handling significant scale or complexity. Deep understanding of reliability, performance characteristics, and operational challenges in production environments."}
- {"name":"Cross-Functional Technical Communication","description":"Experience translating between research and engineering domains, explaining technical tradeoffs and architectural decisions to non-infrastructure specialists. Ability to communicate with researchers fluently about model behavior and systems behavior."}

## Skills

### required

- {"name":"ML Model Training Infrastructure","description":"Production experience with distributed training systems, experiment management, model checkpointing, and training pipeline orchestration for large-scale machine learning workloads."}
- {"name":"Distributed Systems Design","description":"Core expertise in designing and operating distributed systems including data consistency, fault tolerance, load balancing, and performance optimization under real-world constraints."}
- {"name":"Infrastructure-as-Code & DevOps","description":"Proficiency with containerization (Docker), orchestration platforms, CI/CD systems, and infrastructure automation. Experience managing deployment pipelines and production infrastructure."}
- {"name":"Technical Leadership","description":"Ability to set technical direction, make architectural decisions, conduct thorough code reviews, and stay hands-on with implementation alongside the team."}
- {"name":"Team Leadership & Hiring","description":"Experience sourcing, interviewing, and hiring high-quality engineers. Track record of developing engineers through mentorship and strategic project assignments."}

### preferred

- {"name":"Reinforcement Learning Infrastructure","description":"Hands-on experience building or maintaining infrastructure specifically for reinforcement learning training, including rollout systems, reward computation, and experiment orchestration at scale."}
- {"name":"Evaluation & Testing Frameworks","description":"Experience designing evaluation pipelines for machine learning models, building regression detection systems, or creating comprehensive testing infrastructure for model behavior validation."}
- {"name":"Simulated Environments & Sandboxing","description":"Experience building and maintaining sandboxed or simulated environments for model training or testing, including reproducibility, isolation, and performance optimization."}
- {"name":"AI/Coding Tools Ecosystem","description":"Familiarity with AI-assisted coding tools, language models, or code completion systems. Experience integrating with or building on top of modern AI development tools."}
- {"name":"Observability & Monitoring","description":"Deep expertise in instrumentation, logging, tracing, and monitoring systems. Experience building observability platforms that provide actionable insights into system behavior and performance."}
- {"name":"Cost Optimization for ML Infrastructure","description":"Experience optimizing infrastructure costs for machine learning systems, including compute resource allocation, efficient data movement, and cost-aware architectural decisions."}

## Tech stack

### tools

- {"name":"Git & Version Control","description":"Essential for managing infrastructure code, collaborative development, and CI/CD pipelines. Deep understanding of branching strategies and code review workflows."}
- {"name":"Observability & Monitoring Tools","description":"Prometheus, Grafana, Datadog, New Relic, or similar platforms for infrastructure monitoring, alerting, and debugging production systems."}
- {"name":"CI/CD Platforms","description":"GitHub Actions, GitLab CI, Jenkins, or similar for automating testing, building, and deployment of infrastructure and model training pipelines."}
- {"name":"Experiment Tracking","description":"MLflow, Weights & Biases, Neptune, or similar platforms for tracking experiments, metrics, and model versions across training runs."}

### others

- {"name":"Cloud Platforms (AWS/GCP/Azure)","description":"Production experience running distributed training and ML infrastructure on cloud platforms. Understanding of compute instances, storage, networking, and cost optimization."}
- {"name":"Infrastructure Architecture & Design","description":"Strong foundation in system design, scalability patterns, and architectural tradeoffs specific to ML workloads including latency, throughput, and cost optimization."}
- {"name":"Production Debugging & Profiling","description":"Tools and techniques for profiling, tracing, and debugging production systems. Experience using strace, perf, flamegraphs, or similar profiling tools."}

### databases

- {"name":"Time-Series Databases","description":"Monitoring and metrics collection for infrastructure observability. Examples include Prometheus, InfluxDB, or similar systems for tracking performance metrics."}
- {"name":"Data Lakes or Warehouses","description":"Storage and query systems for large-scale training data, evaluation results, and experimental metrics. Experience with S3, GCS, Snowflake, or similar platforms valuable."}

### languages

- {"name":"Python","description":"Primary language for ML infrastructure, training scripts, and orchestration systems. Preferred for writing infrastructure code that interfaces with model training pipelines."}
- {"name":"Go or Rust","description":"Systems-level language for high-performance infrastructure components, distributed systems services, or resource-constrained environments. Valuable for building robust, efficient infrastructure."}

### frameworks

- {"name":"PyTorch or TensorFlow","description":"Industry-standard deep learning frameworks. Experience with their distributed training capabilities, mixed precision training, and integration with training infrastructure."}
- {"name":"Ray","description":"Distributed computing framework for scalable machine learning. Useful for distributed training, reinforcement learning experiments, and hyperparameter tuning at scale."}
- {"name":"Kubernetes","description":"Container orchestration for managing distributed training clusters, job scheduling, resource management, and infrastructure scalability."}

## Benefits

### benefits

- {"name":"Equity & Ownership","description":"Meaningful equity stake in Cursor as an early-stage company with significant growth potential in the AI developer tools space. Direct impact on the company's success."}
- {"name":"Cutting-Edge Technology & Impact","description":"Lead infrastructure powering the world's leading AI-assisted coding platform. Work at the intersection of distributed systems and machine learning on problems with significant real-world impact."}
- {"name":"Small, Talent-Dense Team","description":"Work in a flat organizational structure with exceptionally talented engineers and researchers. High autonomy, rapid decision-making, and direct impact on company direction."}
- {"name":"Technical Leadership Opportunity","description":"Rare opportunity to remain deeply technical while leading infrastructure strategy. Set architectural direction and stay hands-on with code and debugging alongside your team."}
- {"name":"Mentorship & Growth","description":"Develop and mentor exceptional infrastructure engineers. Build a team culture focused on learning, shipping, and creative problem-solving."}

## Compensation

- **max:** 280000
- **min:** 185000
- **currency:** USD
- **stockOptions:** true

## Interview process

### steps

- {"name":"Initial Screening Conversation","description":"Preliminary call with Cursor's recruiting team to discuss your background in ML infrastructure and team leadership experience. Focus on your technical depth and experience scaling machine learning systems."}
- {"name":"Technical Deep Dive","description":"Conversation with current infrastructure team members or engineering leadership. Expect discussion of past projects, architectural decisions, debugging experiences, and how you've handled complex distributed systems challenges."}
- {"name":"Leadership & Vision Interview","description":"Discussion with Cursor's leadership about your approach to team building, technical direction-setting, and cross-functional collaboration. Emphasis on your vision for ML infrastructure at scale and how you'd approach ambiguity."}
- {"name":"Research & Engineering Collaboration","description":"Conversation with members of Cursor's research team to assess your ability to translate between research and engineering domains, understand model training dynamics, and make infrastructure tradeoff decisions."}
- {"name":"Executive Alignment","description":"Final conversation with senior leadership to discuss long-term vision, team scaling strategy, and alignment on Cursor's mission to automate coding through cutting-edge AI and infrastructure."}

## Full description
Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.

## About the Role

You will lead a team of engineers building the infrastructure used to train, test, and evaluate our models. This is one of the few places at Cursor where infrastructure and model behavior meet directly: when something breaks, it's rarely obvious whether it's a systems bug or the model doing exactly what it was trained to do, and your team has to be good at telling the difference before they can fix it.

You'll set technical direction for how we train and evaluate models at scale, stay close enough to the code to debug alongside your team, and work daily with researchers to turn tradeoffs in latency, quality, and cost into infrastructure that actually gets built. We're hiring across a range of scope for this role, depending on experience and the size of problem you're ready to own.

## Example projects include..

* Building the rollout infrastructure that lets researchers run RL experiments at scale without fighting the plumbing.
* Designing eval pipelines that catch regressions before they ship, and give researchers fast, trustworthy signal on whether a change actually helped.
* Owning the environments in which models are trained and tested: sandboxed, reproducible, and fast enough that iteration speed isn't the bottleneck.
* Bringing rigor to how the team measures quality and progress, in places where "did it ship" isn't the same as "did it work?"
* Partnering with research to translate model-level tradeoffs (latency, quality, cost) into concrete infrastructure decisions.
* Hiring and growing the team: sourcing, interviewing, and closing exceptional infrastructure engineers, while developing your engineers through coaching, mentorship, and high-leverage project assignments.

## You may be a fit if

* You've led engineering teams building infrastructure that trains, evaluates, or serves ML models in production.
* You have strong infrastructure and distributed systems fundamentals: you know what reliability and performance look like under real load, not just in a design doc.
* You genuinely want to stay technical: you're comfortable writing code, reviewing PRs with depth, and using tools like Cursor itself to move fast.
* You’re comfortable operating in ambiguity: you ask the right questions, make sound decisions with incomplete information, and help the team find a path forward.
* You have a track record of hiring and developing engineers who are better than you were at their stage.
* You can talk fluently with researchers about model behavior and with engineers about systems design, and you know when a problem is actually the other team's.
* Bonus: hands-on experience with RL training infrastructure, eval frameworks, or building and maintaining simulated environments for model training or testing.
