Software Engineer, Data Infrastructure
Backend Engineer · Senior · Full Time
Opens Cursor's application page
Role
What you'll do.
Cursor is seeking a Software Engineer for Data Infrastructure to own the full lifecycle of data systems that power their AI-powered coding tool. This role involves building and maintaining scalable data pipelines, telemetry systems, and storage infrastructure that supports daily product releases and model improvements. The ideal candidate has deep experience with Spark, distributed systems, and production data infrastructure at scale.
Responsibilities
- Data Pipeline Architecture: Design and implement scalable data pipelines that process telemetry, prompts, completions, and agent run data from daily product releases
- System Migration and Redesign: Evaluate existing systems and determine when to patch versus rebuild, shipping replacements while maintaining operational continuity
- Privacy-Compliant Data Handling: Implement data retention and usage policies that respect Privacy Mode settings and organizational configurations
- Performance Optimization: Debug and resolve performance issues across client instrumentation, streaming systems, storage layers, and model-facing workflows
- Schema Evolution Management: Design and implement schema validation and evolution strategies to prevent silent degradation across multiple data consumers
- Cost Management: Implement data retention policies, compression strategies, and storage optimization to control infrastructure costs
- Instrumentation Gap Resolution: Identify and resolve telemetry gaps, implement monitoring contracts, and build dashboards for early detection of issues
- Cross-Team Collaboration: Work closely with product and model teams to understand data requirements and prioritize infrastructure improvements by business impact
Qualifications
What we look for.
Technical
Apache Spark Expertise
Deep production experience with Spark (Databricks or open-source), including optimization and troubleshooting at scale
Distributed Systems Experience
Hands-on ownership of large-scale data pipelines and storage systems with proven scalability track record
Ray Data Production Experience
Real-world experience deploying and managing Ray Data for distributed data processing workloads
Performance Debugging Skills
Ability to diagnose and resolve performance issues across multiple layers including compute, storage, networking, and application code
Data Modeling Expertise
Strong understanding of data modeling principles, schema design, and long-term system maintainability
Education
Computer Science Degree
Bachelor's degree in Computer Science, Engineering, or equivalent practical experience in software development
Experience
Large-Scale Systems
5+ years building and operating production data infrastructure systems at significant scale
Data Pipeline Ownership
End-to-end ownership experience including design, implementation, deployment, and ongoing operations
System Architecture Decisions
Proven track record of making sound technical decisions about when to refactor versus rebuild existing systems
Skills
Required
Apache Spark
Expert-level proficiency in Spark for large-scale data processing, including performance tuning and optimization
Ray Data
Production experience with Ray Data framework for distributed ML data processing workflows
Python/Scala
Strong programming skills in languages commonly used for data infrastructure development
Distributed Systems
Deep understanding of distributed computing principles, consistency models, and fault tolerance
Data Pipeline Design
Ability to architect robust, scalable data pipelines with proper error handling and monitoring
Performance Debugging
Systematic approach to identifying and resolving performance bottlenecks across complex distributed systems
Preferred
ClickHouse
Nice to haveExperience running or scaling ClickHouse for real-time analytics workloads
dbt
Nice to haveFamiliarity with dbt for data transformation and analytics engineering workflows
Dagster
Nice to haveExperience with Dagster or similar orchestration tools for data pipeline management
Cloud Platforms
Nice to haveExperience with AWS, GCP, or Azure for cloud-based data infrastructure deployment
Kubernetes
Nice to haveContainer orchestration experience for deploying and managing data processing workloads
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 180,000 – 280,000
Equity·Stock options
Benefits
In-Person Work Environment
Cozy offices in North Beach, San Francisco and Manhattan, New York with well-stocked libraries
Equity Package
Competitive equity compensation as part of total compensation package
Professional Development
Work with cutting-edge AI technology and contribute to groundbreaking automation tools
Collaborative Culture
Small, talent-dense team with flat organization structure encouraging spirited debate and creative problem-solving
Office Amenities
Well-appointed offices with comprehensive libraries and comfortable work environments
Process
Interview steps.
- 01
Initial Screen
Application review and initial fit assessment based on technical background and experience
- 02
Technical Interviews
2-3 short technical interviews focusing on data infrastructure, distributed systems, and problem-solving approach
- 03
Onsite Interview
Full-day onsite visit including hands-on project work, technical discussions, and team meetings to assess cultural fit and technical depth
Full posting
Original listing.
Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.
We're in-person with cozy offices in North Beach, San Francisco and Manhattan, New York, replete with well-stocked libraries.
About the Role
Cursor ships daily. Every release leaves signals behind: telemetry, prompts, completions, agent runs, sessions. Those signals power model improvement, evals, and experimentation. Data infrastructure is what turns them into something teams can trust.
A lot of systems here started simple so we could move fast. Over time, the constraints change and the “good enough” version becomes the bottleneck. This role owns the full ladder: patch what should be patched, redesign what should be redesigned, ship the replacement, and operate it.
Privacy guarantees are part of correctness. What we can retain and use depends on Privacy Mode and org configuration, and getting that wrong breaks a product promise. We choose work by business impact: what blocks product and model teams today, and what will block them next month.
Sample projects include...
A core pipeline started as a pragmatic reuse of infrastructure built for something else. It works, but it cannot guarantee properties downstream consumers now need (for example, point-in-time consistency). You design and ship the replacement while keeping the existing system running.
A new product surface ships without instrumentation. You talk to the team, define what needs to be captured, and wire it through before the absence becomes anyone else’s problem.
Eval coverage drops. You trace it to an instrumentation gap introduced weeks ago by a product change nobody flagged. You fix the gap, add a contract so it cannot recur, and ship the dashboard that would have caught it earlier.
Multiple consumers depend on overlapping data. You design schema evolution and validation so changes in one place do not silently degrade the others.
Storage costs rise faster than usage. You decide what is worth keeping, implement retention and compression, and delete what is not.
What we're looking for
We’re looking for someone who has built real systems at scale and cares about correctness, cost, and ergonomics.
Strong signals include:
Deep experience with Spark (Databricks or open-source Spark both count)
Production experience with Ray Data
Hands-on ownership of large data pipelines and storage systems
Comfort debugging performance issues across client instrumentation, streaming, storage, and model-facing workflows, as well as, compute, storage, and networking layers
Clear thinking about data modeling and long-term maintainability
You have good judgment about when to patch and when to rebuild
Nice to have
Experience running or scaling ClickHouse
Familiarity with dbt, Dagster, or similar orchestration and modeling tools
We're in-person with cozy offices in North Beach, San Francisco and Manhattan, New York, replete with well-stocked libraries.
Applying
If there appears to be a fit, we'll reach to schedule 2-3 short technicals. After, we'll schedule an onsite in our office, where you'll work on a small project, discuss ideas, and meet the team.
Redirects to Cursor's application page.
Other roles
More at Cursor.
Field Engineer - India
Mid
Engineering Manager, Agent & Product Security
Manager
Engineering Manager, ML
Manager
RVP, Field Engineering, Public Sector
VP
Regional Director, Field Engineering, Healthcare
Manager