Software Engineer, RL Data

ML Engineer · Mid · Full Time

San FranciscoUSD 180k – 280k2d ago
Apply for this role

Opens Cursor's application page

Role

What you'll do.

Join Cursor's RL Data team as a Software Engineer to design training tasks, reward functions, and environments for frontier coding agents. You'll work with real user data to create datasets and evaluation systems that directly improve agent performance, combining software engineering fundamentals with data systems expertise in a flat, talent-dense organization focused on automating coding.

Responsibilities

  • Design and Iterate Training Task Sets: Create curated task datasets that teach specific agent capabilities, such as code generation quality, debugging, or complex programming patterns. Analyze traces and evaluation metrics to identify what's working and iterate on task design based on model performance data, ensuring continuous improvement in agent effectiveness.
  • Analyze Agent Traces and Identify Failure Modes: Conduct deep analysis of model behavior from training traces and inference logs to discover failure patterns, unexpected behaviors, or edge cases. Build systematic approaches to surface and reproduce these modes, converting ad-hoc findings into scalable datasets that help the model learn from mistakes.
  • Develop Reward Functions and Evaluation Systems: Design reward functions that accurately measure task performance and align with Cursor's goals for agent effectiveness. Create robust evaluation systems that provide meaningful signals about model progress, enabling data-driven decisions about training effectiveness and dataset quality.
  • Build Reusable Data Infrastructure: Transform one-off scripts and recipes into production-ready systems that other teams can leverage. Abstract reward logic, environment setup, and data quality checks into maintainable, well-documented components that scale across the organization and improve data team productivity.
  • Collaborate with Research Teams: Partner with researchers to validate that datasets are actually teaching the intended capabilities, not just improving benchmark scores. Conduct analyses to understand model learning dynamics, propose data improvements, and ensure training datasets align with real-world deployment requirements.
  • Own Data Quality and Validation: Establish and maintain high standards for training data quality, including validation pipelines, outlier detection, and data lineage tracking. Ensure datasets are clean, representative, and properly documented for reproducibility and future reference.
  • Scale Real User Data Integration: Work on systems to incorporate real user data into RL training pipelines, ensuring privacy, quality, and relevance. Design scalable data collection and processing infrastructure that enables frontier coding agents to learn from actual user interactions and coding patterns.

Qualifications

What we look for.

Technical

  • Software Engineering Excellence

    Demonstrated ability to write production-quality code with strong fundamentals in software design, testing, debugging, and code organization. Experience shipping code and maintaining systems in production environments.

  • Data Systems Expertise

    Solid background in data infrastructure, ETL pipelines, or distributed data processing. Proficiency with data storage systems, query optimization, and designing systems that handle large-scale datasets efficiently.

  • Python and Scientific Computing

    Advanced Python programming skills with experience using scientific computing libraries like NumPy, Pandas, or similar tools for data manipulation and analysis at scale.

  • Problem Solving Under Ambiguity

    Comfort with ill-defined problems and ability to research, decompose complex challenges into concrete, measurable components with clear success criteria.

Education

  • Bachelor's Degree in Computer Science or Related Field

    Foundation in computer science, mathematics, physics, or engineering with strong algorithmic thinking and computational problem-solving skills.

Experience

  • 3-5 Years of Software Engineering

    Solid professional experience building production systems, with proven ability to own features end-to-end and deliver high-quality engineering solutions in team environments.

  • Data Infrastructure or ML Systems Background

    Background in data engineering, ML infrastructure, backend systems, or distributed systems. Experience with training machine learning models, optimization, or large-scale data processing is valuable.

  • Analytical Mindset with Shipping Velocity

    Track record of moving quickly while maintaining high code quality. Ability to balance perfectionism with pragmatism, knowing when to iterate versus when to over-engineer.

Skills

Required

  • Software Engineering Fundamentals

    Strong ability to write clean, efficient, and maintainable code with proficiency in version control, testing, and code design patterns essential for production systems.

  • Data Systems and Pipeline Design

    Experience building, maintaining, and optimizing data pipelines and systems that process large-scale datasets, with understanding of data quality, validation, and lineage.

  • Python Programming

    Proficiency in Python for data processing, scripting, and system development, which is foundational for RL data infrastructure and model training workflows.

  • Machine Learning Concepts

    Understanding of ML fundamentals including training loops, loss functions, evaluation metrics, and how datasets impact model behavior and performance.

  • Problem Decomposition

    Ability to break down complex, ambiguous problems into measurable, concrete tasks and define appropriate success metrics and evaluation criteria.

Preferred

  • Reinforcement Learning Experience

    Nice to have

    Prior experience with RL algorithms, training environments, reward design, or agent development is valuable but not required for this role.

  • Distributed Systems

    Nice to have

    Background working with distributed computing, parallel processing, or scalable infrastructure for handling large-scale machine learning workloads.

  • Infrastructure and DevOps

    Nice to have

    Experience with data infrastructure, experiment tracking systems, or backend services that support ML operations and scalability.

  • Behavioral Analysis and Data Exploration

    Nice to have

    Skill in analyzing complex system behavior, reading logs and traces, and discovering patterns or failure modes in real-world agent performance.

  • AI/ML Model Training

    Nice to have

    Experience training large language models, coding models, or other frontier AI systems at scale with understanding of dataset quality impact.

Tech stack

Languages

PythonSQL

Frameworks

PyTorchRay RLlib or OpenAI Gym

Databases

Data Warehousing Systems

Tools

Git and Version ControlExperiment Tracking ToolsJupyter Notebooks

Other

RESTful APIsDocker and Containerization

Compensation

Pay and benefits.

Base·USD 180,000 – 280,000

Equity·Stock options

Benefits

  • Equity Compensation

    Significant stock option grants as an early-stage employee at Cursor, a well-funded AI startup working on frontier technology with potential for substantial upside.

  • Health and Wellness

    Comprehensive health insurance coverage including medical, dental, and vision plans for employees and their families.

  • Professional Development

    Access to learning resources, conference attendance budgets, and opportunities to work on cutting-edge AI research and machine learning systems.

  • Flexible Work Arrangement

    Remote-friendly work environment with flexibility to work from San Francisco or distributed, enabling work-life balance while collaborating with a talent-dense team.

  • Collaborative Culture

    Flat organizational structure with direct access to leadership, emphasis on spirited technical debate, and an environment that values creative problem-solving and shipping fast.

  • Meaningful Impact

    Direct influence on frontier coding agent capabilities that shape the future of AI-assisted software development, working on problems at the frontier of reinforcement learning and code generation.

Process

Interview steps.

  1. 01

    Initial Screening

    Phone or video screening with a recruiter to assess background, experience with data systems or RL, and general fit with Cursor's mission and culture focused on building productive tools.

  2. 02

    Technical Problem Discussion

    Conversation with a senior engineer covering data systems design, problem decomposition, and how you approach building reliable infrastructure. Expect discussions around real-world tradeoffs in data pipeline design.

  3. 03

    RL and Data Fundamentals

    Technical discussion with the RL Data team lead exploring your understanding of reinforcement learning concepts, task design, evaluation metrics, and how you think about dataset quality and agent performance measurement.

  4. 04

    System Design and Collaboration

    Work through a realistic scenario involving designing a training task, defining rewards, and building infrastructure to scale it. This assesses how you think about user-facing problems and collaborate across teams.

  5. 05

    Leadership and Culture Fit

    Conversation with Cursor leadership exploring your approach to shipping code, handling ambiguity, and working in a flat organization. Emphasis on truth-seeking, creative problem-solving, and working with a high-performance team.

Full posting

Original listing.

Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.

Software Engineer, Reinforcement learning

Cursor is building the future of coding. We train frontier coding agents and scale RL on real user data to make them increasingly effective.

About the role

As a Software Engineer on the RL Data team at Cursor, you'll create the tasks, rewards, and environments that train our coding agents. The team owns the data that goes into training: what the model is asked to do, how we score it, and the setups it learns in.

What you’ll do

  • Designing a task set that teaches a specific agent capability, then iterating on it from traces and evals until the model actually gets better.

  • Reading a pile of agent traces, finding a failure mode or a surprising behavior, and building a system that surfaces more of the same.

  • Turning a one-off recipe into something other teams can reuse: better rewards, cleaner environments, tighter data quality.

  • Partnering with research on whether a dataset is actually teaching the thing we think it is.

You may be a fit if

  • You write careful, fast code and have strong software engineering fundamentals.

  • You like setting tasks: breaking a fuzzy capability into something concrete you can measure.

  • You have an infra, data, or distributed systems background. RL experience is a plus, not a requirement.

  • You enjoy looking at messy real-world agent behavior and turning it into a dataset or a tool.

Redirects to Cursor's application page.

Other roles

More at Cursor.

View all 39 roles