Engineering Manager, Evals

Engineering Manager · Manager · Full Time

San FranciscoUSD 250k – 350k1mo ago
Apply for this role

Opens Cursor's application page

Role

What you'll do.

Cursor is seeking an Engineering Manager for their Evals team to lead critical evaluation systems that measure and improve AI coding agent quality. The ideal candidate will drive the development of comprehensive evaluation tools, align cross-functional teams, and build robust metrics that enhance product development and model training processes.

Responsibilities

  • Eval Strategy: Set comprehensive evaluation roadmap defining measurement criteria, strategic importance, and how evaluation signals drive product and training decisions
  • Team Leadership: Lead and develop a high-impact team of engineers and researchers focused on creating evaluation datasets and developer-friendly tools
  • CursorBench Development: Guide the evolution of CursorBench to accurately reflect real developer workflows and expand evaluation capabilities
  • Quality Metrics: Define precise online quality signals and transform potential regressions into robust performance guardrails
  • Integration Management: Integrate evaluation processes into decision-making workflows for product launches, deployments, and model training cycles

Qualifications

What we look for.

Technical

  • Evaluation Systems

    Proven experience building and operating evaluation or measurement systems (AI evaluations, experimentation platforms, relevance metrics)

  • Data Analysis

    Strong data acumen with ability to collaborate effectively with data scientists and researchers

  • AI/ML Trends

    Deep understanding of emerging AI research and industry trends in model and agent behavior

Education

  • Advanced Degree

    Bachelor's or Master's degree in Computer Science, Software Engineering, or related technical field preferred

Experience

  • Engineering Leadership

    Extensive experience leading engineering teams that ship complex production systems

  • Cross-Functional Alignment

    Proven ability to align research, product, data, and infrastructure teams around quality metrics and processes

Skills

Required

  • People Leadership

    Strong people management, coaching, and team development skills

  • Strategic Planning

    Ability to develop and execute comprehensive technical strategies

Preferred

  • AI Research

    Nice to have

    Background or deep interest in AI/ML research and emerging technology trends

  • Measurement Frameworks

    Nice to have

    Experience designing robust evaluation and measurement frameworks

Tech stack

Languages

Python

Frameworks

Machine Learning Frameworks

Databases

Data Analysis Databases

Tools

CursorBench

Other

AI Evaluation Tools

Compensation

Pay and benefits.

Base·USD 250,000 – 350,000

Equity·Stock options

Benefits

  • Competitive Compensation

    Highly competitive salary package for top engineering talent

  • Equity Options

    Startup equity package to share in long-term company growth

  • Professional Development

    Opportunities for continuous learning and cutting-edge AI research exposure

  • Innovative Work Environment

    Flat organizational structure with emphasis on creativity and impactful work

Process

Interview steps.

  1. 01

    Initial Screening

    HR phone screen to assess basic qualifications and cultural fit

  2. 02

    Technical Leadership Interview

    In-depth discussion of engineering management experience and leadership philosophy

  3. 03

    AI/ML Technical Interview

    Deep dive into candidate's understanding of AI evaluation methodologies and research trends

  4. 04

    Team Fit Interview

    Meeting with potential team members to assess collaborative potential

  5. 05

    Final Executive Interview

    Conversation with senior leadership to align on strategic vision and leadership approach

Full posting

Original listing.

Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.

About the Role

As an Engineering Manager on the Evals team at Cursor, you’ll lead the group responsible for creating high-signal evaluation datasets for coding agents and building the tools engineers use to write and run them. The team also owns online evaluation systems that track agent quality in production, and the close integration between online and offline evaluations.

The evaluation systems that this team builds, including CursorBench, are critical in the development of our coding models and the quality of our Cursor agents. Your impact will compound across every Cursor product and every Cursor model by making quality measurable, comparable, and easy to improve.

What you’ll do

  • Set the eval roadmap end-to-end—what we measure, why it matters, and how signals turn into shipping + training decisions.

  • Lead and grow a high-impact team of engineers and researchers building eval datasets and developer-friendly tools to write and run evals.

  • Guide the next generation of CursorBench so it continues to reflect real developer workflows at Cursor, and expand it with new evals that measure other properties developers value.

  • Define crisp online quality signals and turn regressions into robust guardrails.

  • Integrate evals into decision-making cadence for launches, deploys, and model training loops.

You may be a fit if

  • You’ve led engineering teams shipping production systems and have strong people leadership and coaching skills.

  • You can align research, product, data, and infrastructure on what “good” means—and turn that into durable metrics, processes, and release/training rituals.

  • You have good taste and strong opinions on model and agent behaviors, and you stay up-to-date on emerging research and industry trends.

  • You have strong data acumen, and can collaborate effectively with data scientists and researchers.

  • You’ve built and operated evaluation or measurement systems (e.g., AI evals, experimentation platforms, ranking/relevance, search quality, or reliability instrumentation).

#LI-DNI

Redirects to Cursor's application page.

Other roles

More at Cursor.

View all 36 roles