Software Engineer, RL Data
ML Engineer · Mid · Full Time
Opens Cursor's application page
Role
What you'll do.
Join Cursor's RL Data team as a Software Engineer to design training tasks, reward functions, and environments for frontier coding agents. You'll work with real user data to create datasets and evaluation systems that directly improve agent performance, combining software engineering fundamentals with data systems expertise in a flat, talent-dense organization focused on automating coding.
Responsibilities
- Design and Iterate Training Task Sets: Create curated task datasets that teach specific agent capabilities, such as code generation quality, debugging, or complex programming patterns. Analyze traces and evaluation metrics to identify what's working and iterate on task design based on model performance data, ensuring continuous improvement in agent effectiveness.
- Analyze Agent Traces and Identify Failure Modes: Conduct deep analysis of model behavior from training traces and inference logs to discover failure patterns, unexpected behaviors, or edge cases. Build systematic approaches to surface and reproduce these modes, converting ad-hoc findings into scalable datasets that help the model learn from mistakes.
- Develop Reward Functions and Evaluation Systems: Design reward functions that accurately measure task performance and align with Cursor's goals for agent effectiveness. Create robust evaluation systems that provide meaningful signals about model progress, enabling data-driven decisions about training effectiveness and dataset quality.
- Build Reusable Data Infrastructure: Transform one-off scripts and recipes into production-ready systems that other teams can leverage. Abstract reward logic, environment setup, and data quality checks into maintainable, well-documented components that scale across the organization and improve data team productivity.
- Collaborate with Research Teams: Partner with researchers to validate that datasets are actually teaching the intended capabilities, not just improving benchmark scores. Conduct analyses to understand model learning dynamics, propose data improvements, and ensure training datasets align with real-world deployment requirements.
- Own Data Quality and Validation: Establish and maintain high standards for training data quality, including validation pipelines, outlier detection, and data lineage tracking. Ensure datasets are clean, representative, and properly documented for reproducibility and future reference.
- Scale Real User Data Integration: Work on systems to incorporate real user data into RL training pipelines, ensuring privacy, quality, and relevance. Design scalable data collection and processing infrastructure that enables frontier coding agents to learn from actual user interactions and coding patterns.
Qualifications
What we look for.
Technical
Software Engineering Excellence
Demonstrated ability to write production-quality code with strong fundamentals in software design, testing, debugging, and code organization. Experience shipping code and maintaining systems in production environments.
Data Systems Expertise
Solid background in data infrastructure, ETL pipelines, or distributed data processing. Proficiency with data storage systems, query optimization, and designing systems that handle large-scale datasets efficiently.
Python and Scientific Computing
Advanced Python programming skills with experience using scientific computing libraries like NumPy, Pandas, or similar tools for data manipulation and analysis at scale.
Problem Solving Under Ambiguity
Comfort with ill-defined problems and ability to research, decompose complex challenges into concrete, measurable components with clear success criteria.
Education
Bachelor's Degree in Computer Science or Related Field
Foundation in computer science, mathematics, physics, or engineering with strong algorithmic thinking and computational problem-solving skills.
Experience
3-5 Years of Software Engineering
Solid professional experience building production systems, with proven ability to own features end-to-end and deliver high-quality engineering solutions in team environments.
Data Infrastructure or ML Systems Background
Background in data engineering, ML infrastructure, backend systems, or distributed systems. Experience with training machine learning models, optimization, or large-scale data processing is valuable.
Analytical Mindset with Shipping Velocity
Track record of moving quickly while maintaining high code quality. Ability to balance perfectionism with pragmatism, knowing when to iterate versus when to over-engineer.
Skills
Required
Software Engineering Fundamentals
Strong ability to write clean, efficient, and maintainable code with proficiency in version control, testing, and code design patterns essential for production systems.
Data Systems and Pipeline Design
Experience building, maintaining, and optimizing data pipelines and systems that process large-scale datasets, with understanding of data quality, validation, and lineage.
Python Programming
Proficiency in Python for data processing, scripting, and system development, which is foundational for RL data infrastructure and model training workflows.
Machine Learning Concepts
Understanding of ML fundamentals including training loops, loss functions, evaluation metrics, and how datasets impact model behavior and performance.
Problem Decomposition
Ability to break down complex, ambiguous problems into measurable, concrete tasks and define appropriate success metrics and evaluation criteria.
Preferred
Reinforcement Learning Experience
Nice to havePrior experience with RL algorithms, training environments, reward design, or agent development is valuable but not required for this role.
Distributed Systems
Nice to haveBackground working with distributed computing, parallel processing, or scalable infrastructure for handling large-scale machine learning workloads.
Infrastructure and DevOps
Nice to haveExperience with data infrastructure, experiment tracking systems, or backend services that support ML operations and scalability.
Behavioral Analysis and Data Exploration
Nice to haveSkill in analyzing complex system behavior, reading logs and traces, and discovering patterns or failure modes in real-world agent performance.
AI/ML Model Training
Nice to haveExperience training large language models, coding models, or other frontier AI systems at scale with understanding of dataset quality impact.
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 180,000 – 280,000
Equity·Stock options
Benefits
Equity Compensation
Significant stock option grants as an early-stage employee at Cursor, a well-funded AI startup working on frontier technology with potential for substantial upside.
Health and Wellness
Comprehensive health insurance coverage including medical, dental, and vision plans for employees and their families.
Professional Development
Access to learning resources, conference attendance budgets, and opportunities to work on cutting-edge AI research and machine learning systems.
Flexible Work Arrangement
Remote-friendly work environment with flexibility to work from San Francisco or distributed, enabling work-life balance while collaborating with a talent-dense team.
Collaborative Culture
Flat organizational structure with direct access to leadership, emphasis on spirited technical debate, and an environment that values creative problem-solving and shipping fast.
Meaningful Impact
Direct influence on frontier coding agent capabilities that shape the future of AI-assisted software development, working on problems at the frontier of reinforcement learning and code generation.
Process
Interview steps.
- 01
Initial Screening
Phone or video screening with a recruiter to assess background, experience with data systems or RL, and general fit with Cursor's mission and culture focused on building productive tools.
- 02
Technical Problem Discussion
Conversation with a senior engineer covering data systems design, problem decomposition, and how you approach building reliable infrastructure. Expect discussions around real-world tradeoffs in data pipeline design.
- 03
RL and Data Fundamentals
Technical discussion with the RL Data team lead exploring your understanding of reinforcement learning concepts, task design, evaluation metrics, and how you think about dataset quality and agent performance measurement.
- 04
System Design and Collaboration
Work through a realistic scenario involving designing a training task, defining rewards, and building infrastructure to scale it. This assesses how you think about user-facing problems and collaborate across teams.
- 05
Leadership and Culture Fit
Conversation with Cursor leadership exploring your approach to shipping code, handling ambiguity, and working in a flat organization. Emphasis on truth-seeking, creative problem-solving, and working with a high-performance team.
Full posting
Original listing.
Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.
Software Engineer, Reinforcement learning
Cursor is building the future of coding. We train frontier coding agents and scale RL on real user data to make them increasingly effective.
About the role
As a Software Engineer on the RL Data team at Cursor, you'll create the tasks, rewards, and environments that train our coding agents. The team owns the data that goes into training: what the model is asked to do, how we score it, and the setups it learns in.
What you’ll do
Designing a task set that teaches a specific agent capability, then iterating on it from traces and evals until the model actually gets better.
Reading a pile of agent traces, finding a failure mode or a surprising behavior, and building a system that surfaces more of the same.
Turning a one-off recipe into something other teams can reuse: better rewards, cleaner environments, tighter data quality.
Partnering with research on whether a dataset is actually teaching the thing we think it is.
You may be a fit if
You write careful, fast code and have strong software engineering fundamentals.
You like setting tasks: breaking a fuzzy capability into something concrete you can measure.
You have an infra, data, or distributed systems background. RL experience is a plus, not a requirement.
You enjoy looking at messy real-world agent behavior and turning it into a dataset or a tool.
Redirects to Cursor's application page.
Other roles
More at Cursor.
Software Engineer, Pretraining
Senior
Field Engineer, Life Sciences
Mid
Field Engineer - India
Mid
Engineering Manager, Agent & Product Security
Manager
Engineering Manager, ML
Manager