# Engineering Manager, Evals
**Company:** [Cursor](https://scaleengineer.com/companies/cursor)
Cursor is seeking an Engineering Manager for their Evals team to lead critical evaluation systems that measure and improve AI coding agent quality. The ideal candidate will drive the development of comprehensive evaluation tools, align cross-functional teams, and build robust metrics that enhance product development and model training processes.
**Role:** Engineering Manager
**Seniority:** Manager
**Locations:** San Francisco
**Salary:** 250000–350000 USD
[Apply](https://jobs.ashbyhq.com/cursor/74a6ac48-d85f-45a0-9775-3cdb8b713e1a)
Canonical: https://scaleengineer.com/jobs/cursor/engineering-manager-evals
---
## Responsibilities

- Eval Strategy: Set comprehensive evaluation roadmap defining measurement criteria, strategic importance, and how evaluation signals drive product and training decisions
- Team Leadership: Lead and develop a high-impact team of engineers and researchers focused on creating evaluation datasets and developer-friendly tools
- CursorBench Development: Guide the evolution of CursorBench to accurately reflect real developer workflows and expand evaluation capabilities
- Quality Metrics: Define precise online quality signals and transform potential regressions into robust performance guardrails
- Integration Management: Integrate evaluation processes into decision-making workflows for product launches, deployments, and model training cycles

## Requirements

### education

- {"name":"Advanced Degree","description":"Bachelor's or Master's degree in Computer Science, Software Engineering, or related technical field preferred"}

### technical

- {"name":"Evaluation Systems","description":"Proven experience building and operating evaluation or measurement systems (AI evaluations, experimentation platforms, relevance metrics)"}
- {"name":"Data Analysis","description":"Strong data acumen with ability to collaborate effectively with data scientists and researchers"}
- {"name":"AI/ML Trends","description":"Deep understanding of emerging AI research and industry trends in model and agent behavior"}

### experience

- {"name":"Engineering Leadership","description":"Extensive experience leading engineering teams that ship complex production systems"}
- {"name":"Cross-Functional Alignment","description":"Proven ability to align research, product, data, and infrastructure teams around quality metrics and processes"}

## Skills

### required

- {"name":"People Leadership","description":"Strong people management, coaching, and team development skills"}
- {"name":"Strategic Planning","description":"Ability to develop and execute comprehensive technical strategies"}

### preferred

- {"name":"AI Research","description":"Background or deep interest in AI/ML research and emerging technology trends"}
- {"name":"Measurement Frameworks","description":"Experience designing robust evaluation and measurement frameworks"}

## Tech stack

### tools

- {"name":"CursorBench","description":"Internal evaluation platform for AI coding agents"}

### others

- {"name":"AI Evaluation Tools","description":"Familiarity with AI benchmarking and evaluation platforms"}

### databases

- {"name":"Data Analysis Databases","description":"Experience with analytical databases for tracking evaluation metrics"}

### languages

- {"name":"Python","description":"Primary language for data analysis and AI evaluation tools"}

### frameworks

- {"name":"Machine Learning Frameworks","description":"Likely experience with TensorFlow, PyTorch, or similar ML frameworks"}

## Benefits

### benefits

- {"name":"Competitive Compensation","description":"Highly competitive salary package for top engineering talent"}
- {"name":"Equity Options","description":"Startup equity package to share in long-term company growth"}
- {"name":"Professional Development","description":"Opportunities for continuous learning and cutting-edge AI research exposure"}
- {"name":"Innovative Work Environment","description":"Flat organizational structure with emphasis on creativity and impactful work"}

## Compensation

- **max:** 350000
- **min:** 250000
- **currency:** USD
- **stockOptions:** true

## Interview process

### steps

- {"name":"Initial Screening","description":"HR phone screen to assess basic qualifications and cultural fit"}
- {"name":"Technical Leadership Interview","description":"In-depth discussion of engineering management experience and leadership philosophy"}
- {"name":"AI/ML Technical Interview","description":"Deep dive into candidate's understanding of AI evaluation methodologies and research trends"}
- {"name":"Team Fit Interview","description":"Meeting with potential team members to assess collaborative potential"}
- {"name":"Final Executive Interview","description":"Conversation with senior leadership to align on strategic vision and leadership approach"}

## Full description
Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.

### About the Role

As an **Engineering Manager** on the **Evals** team at Cursor, you’ll lead the group responsible for creating high-signal evaluation datasets for coding agents and building the tools engineers use to write and run them. The team also owns online evaluation systems that track agent quality in production, and the close integration between online and offline evaluations.

The evaluation systems that this team builds, including [CursorBench](https://cursor.com/blog/cursorbench), are critical in the development of our coding models and the [quality of our Cursor agents](https://cursor.com/blog/continually-improving-agent-harness). Your impact will compound across every Cursor product and every Cursor model by making quality measurable, comparable, and easy to improve.

### **What you’ll do**

* Set the eval roadmap end-to-end—what we measure, why it matters, and how signals turn into shipping + training decisions.
* Lead and grow a high-impact team of engineers and researchers building eval datasets and developer-friendly tools to write and run evals.
* Guide the next generation of [**CursorBench**](https://cursor.com/blog/cursorbench) so it continues to reflect real developer workflows at Cursor, and expand it with new evals that measure other properties developers value.
* Define crisp online quality signals and turn regressions into robust guardrails.
* Integrate evals into decision-making cadence for launches, deploys, and model training loops.

### You may be a fit if

* You’ve led engineering teams shipping production systems and have strong people leadership and coaching skills.
* You can align research, product, data, and infrastructure on what “good” means—and turn that into durable metrics, processes, and release/training rituals.
* You have good taste and strong opinions on model and agent behaviors, and you stay up-to-date on emerging research and industry trends.
* You have strong data acumen, and can collaborate effectively with data scientists and researchers.
* You’ve built and operated evaluation or measurement systems (e.g., AI evals, experimentation platforms, ranking/relevance, search quality, or reliability instrumentation).

#LI-DNI
