Software Engineer, Data Infrastructure

Backend Engineer · Senior · Full Time

SF / NYUSD 180k – 280k6mo ago
Apply for this role

Opens Cursor's application page

Role

What you'll do.

Cursor is seeking a Software Engineer for Data Infrastructure to own the full lifecycle of data systems that power their AI-powered coding tool. This role involves building and maintaining scalable data pipelines, telemetry systems, and storage infrastructure that supports daily product releases and model improvements. The ideal candidate has deep experience with Spark, distributed systems, and production data infrastructure at scale.

Responsibilities

  • Data Pipeline Architecture: Design and implement scalable data pipelines that process telemetry, prompts, completions, and agent run data from daily product releases
  • System Migration and Redesign: Evaluate existing systems and determine when to patch versus rebuild, shipping replacements while maintaining operational continuity
  • Privacy-Compliant Data Handling: Implement data retention and usage policies that respect Privacy Mode settings and organizational configurations
  • Performance Optimization: Debug and resolve performance issues across client instrumentation, streaming systems, storage layers, and model-facing workflows
  • Schema Evolution Management: Design and implement schema validation and evolution strategies to prevent silent degradation across multiple data consumers
  • Cost Management: Implement data retention policies, compression strategies, and storage optimization to control infrastructure costs
  • Instrumentation Gap Resolution: Identify and resolve telemetry gaps, implement monitoring contracts, and build dashboards for early detection of issues
  • Cross-Team Collaboration: Work closely with product and model teams to understand data requirements and prioritize infrastructure improvements by business impact

Qualifications

What we look for.

Technical

  • Apache Spark Expertise

    Deep production experience with Spark (Databricks or open-source), including optimization and troubleshooting at scale

  • Distributed Systems Experience

    Hands-on ownership of large-scale data pipelines and storage systems with proven scalability track record

  • Ray Data Production Experience

    Real-world experience deploying and managing Ray Data for distributed data processing workloads

  • Performance Debugging Skills

    Ability to diagnose and resolve performance issues across multiple layers including compute, storage, networking, and application code

  • Data Modeling Expertise

    Strong understanding of data modeling principles, schema design, and long-term system maintainability

Education

  • Computer Science Degree

    Bachelor's degree in Computer Science, Engineering, or equivalent practical experience in software development

Experience

  • Large-Scale Systems

    5+ years building and operating production data infrastructure systems at significant scale

  • Data Pipeline Ownership

    End-to-end ownership experience including design, implementation, deployment, and ongoing operations

  • System Architecture Decisions

    Proven track record of making sound technical decisions about when to refactor versus rebuild existing systems

Skills

Required

  • Apache Spark

    Expert-level proficiency in Spark for large-scale data processing, including performance tuning and optimization

  • Ray Data

    Production experience with Ray Data framework for distributed ML data processing workflows

  • Python/Scala

    Strong programming skills in languages commonly used for data infrastructure development

  • Distributed Systems

    Deep understanding of distributed computing principles, consistency models, and fault tolerance

  • Data Pipeline Design

    Ability to architect robust, scalable data pipelines with proper error handling and monitoring

  • Performance Debugging

    Systematic approach to identifying and resolving performance bottlenecks across complex distributed systems

Preferred

  • ClickHouse

    Nice to have

    Experience running or scaling ClickHouse for real-time analytics workloads

  • dbt

    Nice to have

    Familiarity with dbt for data transformation and analytics engineering workflows

  • Dagster

    Nice to have

    Experience with Dagster or similar orchestration tools for data pipeline management

  • Cloud Platforms

    Nice to have

    Experience with AWS, GCP, or Azure for cloud-based data infrastructure deployment

  • Kubernetes

    Nice to have

    Container orchestration experience for deploying and managing data processing workloads

Tech stack

Languages

PythonScalaSQL

Frameworks

Apache SparkRay DatadbtDagster

Databases

ClickHousePostgreSQLRedis

Tools

DatabricksKubernetesDockerApache KafkaTerraformGrafana

Other

AWSApache ParquetProtocol Buffers

Compensation

Pay and benefits.

Base·USD 180,000 – 280,000

Equity·Stock options

Benefits

  • In-Person Work Environment

    Cozy offices in North Beach, San Francisco and Manhattan, New York with well-stocked libraries

  • Equity Package

    Competitive equity compensation as part of total compensation package

  • Professional Development

    Work with cutting-edge AI technology and contribute to groundbreaking automation tools

  • Collaborative Culture

    Small, talent-dense team with flat organization structure encouraging spirited debate and creative problem-solving

  • Office Amenities

    Well-appointed offices with comprehensive libraries and comfortable work environments

Process

Interview steps.

  1. 01

    Initial Screen

    Application review and initial fit assessment based on technical background and experience

  2. 02

    Technical Interviews

    2-3 short technical interviews focusing on data infrastructure, distributed systems, and problem-solving approach

  3. 03

    Onsite Interview

    Full-day onsite visit including hands-on project work, technical discussions, and team meetings to assess cultural fit and technical depth

Full posting

Original listing.

Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.

We're in-person with cozy offices in North Beach, San Francisco and Manhattan, New York, replete with well-stocked libraries.

About the Role

Cursor ships daily. Every release leaves signals behind: telemetry, prompts, completions, agent runs, sessions. Those signals power model improvement, evals, and experimentation. Data infrastructure is what turns them into something teams can trust.

A lot of systems here started simple so we could move fast. Over time, the constraints change and the “good enough” version becomes the bottleneck. This role owns the full ladder: patch what should be patched, redesign what should be redesigned, ship the replacement, and operate it.

Privacy guarantees are part of correctness. What we can retain and use depends on Privacy Mode and org configuration, and getting that wrong breaks a product promise. We choose work by business impact: what blocks product and model teams today, and what will block them next month.

Sample projects include...

  • A core pipeline started as a pragmatic reuse of infrastructure built for something else. It works, but it cannot guarantee properties downstream consumers now need (for example, point-in-time consistency). You design and ship the replacement while keeping the existing system running.

  • A new product surface ships without instrumentation. You talk to the team, define what needs to be captured, and wire it through before the absence becomes anyone else’s problem.

  • Eval coverage drops. You trace it to an instrumentation gap introduced weeks ago by a product change nobody flagged. You fix the gap, add a contract so it cannot recur, and ship the dashboard that would have caught it earlier.

  • Multiple consumers depend on overlapping data. You design schema evolution and validation so changes in one place do not silently degrade the others.

  • Storage costs rise faster than usage. You decide what is worth keeping, implement retention and compression, and delete what is not.

What we're looking for

We’re looking for someone who has built real systems at scale and cares about correctness, cost, and ergonomics.

Strong signals include:

  • Deep experience with Spark (Databricks or open-source Spark both count)

  • Production experience with Ray Data

  • Hands-on ownership of large data pipelines and storage systems

  • Comfort debugging performance issues across client instrumentation, streaming, storage, and model-facing workflows, as well as, compute, storage, and networking layers

  • Clear thinking about data modeling and long-term maintainability

  • You have good judgment about when to patch and when to rebuild

Nice to have

  • Experience running or scaling ClickHouse

  • Familiarity with dbt, Dagster, or similar orchestration and modeling tools

We're in-person with cozy offices in North Beach, San Francisco and Manhattan, New York, replete with well-stocked libraries.

Applying

If there appears to be a fit, we'll reach to schedule 2-3 short technicals. After, we'll schedule an onsite in our office, where you'll work on a small project, discuss ideas, and meet the team.

Redirects to Cursor's application page.

Other roles

More at Cursor.

View all 36 roles