Software Engineer, Pretraining
Data Engineer · Senior · Full Time
Opens Cursor's application page
Role
What you'll do.
Cursor is seeking a Software Engineer for Pretraining to build data systems behind frontier coding models' initial training. You'll work on large-scale web crawling, distributed data platform infrastructure, and data quality pipelines, transforming raw internet-scale data into high-quality datasets that power model training. This role requires deep expertise in distributed systems, data engineering, and infrastructure optimization with high ownership and collaboration with AI agents.
Responsibilities
- Build and Scale Web Crawling Infrastructure: Design, build, and scale distributed web crawling systems that discover, schedule, fetch, and parse high-quality documents across the open web for frontier model pretraining. Implement sophisticated URL seeding and scoring algorithms, optimize fair host scheduling to ensure crawl capacity lands on the most valuable sources, and automate delivery of crawl datasets into downstream data pipelines.
- Develop Data Quality Pipelines: Architect and own high-throughput, fully telemetered data pipelines capable of processing frontier-scale data with complete end-to-end traceability. Build systems with proactive monitoring and alerting that surface data quality drift before training runs commence. Design and implement classification, ranking, filtering, and cleaning models that operate at extreme throughput without becoming critical path bottlenecks.
- Engineer Data Platform Infrastructure: Develop the foundational platform and orchestration layer that transforms raw web, code, multimodal, and acquired data into production-ready training datasets. Own pipeline reliability, observability, and reproducibility at scale. Create comprehensive data lineage tracking, freshness monitoring, and health signals that enable researchers to trust data quality and iterate rapidly on new data sources and mixtures.
- Optimize Data Quality Through Experimentation: Design and execute scaling-ladder experiments on data mixtures, repeatability, and quality depth. Transform qualitative data assessments into quantitative evidence through rigorous experimental methodology. Partner with training teams to establish causal links between data characteristics and downstream model performance metrics including loss reduction and capability improvements.
- Debug and Harden Complex Distributed Systems: Independently diagnose and resolve complex system failures in crawling and data infrastructure end-to-end. Improve system availability, recovery patterns, and ingestion lag through systematic performance analysis. Defeat antibot mechanisms, improve HTML and document parsing extractors, and harden infrastructure against edge cases that prevent clean content acquisition.
- Cross-Functional Technical Partnership: Collaborate tightly with Data Acquisition, Data Quality, Data Platform, and frontier training teams to close the loop between infrastructure improvements and model capability gains. Work alongside AI agents in debugging and optimization tasks. Translate data pipeline enhancements into measurable improvements in model loss curves and evaluation performance metrics.
Qualifications
What we look for.
Technical
Distributed Systems Architecture
Deep expertise designing, implementing, and operating large-scale distributed systems at the infrastructure level. Strong intuitions about consensus algorithms, load balancing, fault tolerance, and graceful degradation. Experience building systems that must operate continuously at extreme scale with minimal latency variability.
Data Pipeline and ETL Engineering
Production experience building and scaling high-throughput data pipelines, ETL workflows, and data transformation systems. Proficiency with orchestration frameworks, data lineage tracking, and monitoring patterns for complex data flows. Experience debugging performance bottlenecks and optimizing throughput at scale.
Web Crawling or Search Infrastructure
Experience with or strong understanding of large-scale web crawling, distributed fetching systems, or search engine infrastructure. Familiarity with challenges including antibot detection evasion, HTML parsing at scale, URL scheduling and prioritization algorithms, and crawl infrastructure reliability patterns.
Performance-Critical Systems Programming
Proven ability to write performance-critical production code that must meet strict latency and throughput requirements. Strong proficiency in languages commonly used for systems programming and data infrastructure. Experience profiling, benchmarking, and optimizing hot paths in latency-sensitive systems.
Observability and Monitoring
Strong capability building comprehensive observability into production systems through strategic instrumentation, metric design, and alerting patterns. Experience identifying signal from noise in high-volume telemetry data and using observability to drive root cause analysis of system issues.
Education
Computer Science Fundamentals
Strong foundation in core computer science concepts including algorithms, data structures, distributed systems theory, and system design principles. This can be demonstrated through formal education, self-study, or practical experience shipping production systems.
Experience
Staff-Level Infrastructure Ownership
Demonstrated history of unusually fast career progression with staff-level or equivalent ownership responsibilities achieved within a few years. Track record of shipping significant infrastructure systems end-to-end with high autonomy and measurable impact on organizational efficiency or capability.
Large-Scale Data Systems
Production experience operating or building data systems that handle petabyte-scale or internet-scale datasets. Familiarity with challenges of data quality assurance, reproducibility, and lineage tracking at massive scale. Understanding of how data quality and characteristics directly impact downstream model training or business outcomes.
Autonomous Problem-Solving
Proven ability to independently architect solutions, debug complex systems without hand-holding, and drive problems to resolution. Comfort working with incomplete specifications and ambiguous requirements. Experience collaborating effectively with cross-functional teams while maintaining technical autonomy.
Skills
Required
Python
Production Python expertise for data processing pipelines, system scripting, and infrastructure tooling. Common choice for data engineering work at frontier AI companies.
Go or Rust
Experience with compiled, systems-level languages for building performance-critical infrastructure components, especially crawling systems and high-throughput data processing kernels.
SQL
Strong SQL proficiency for data transformation, analysis, and pipeline development. Essential for working with structured data and designing efficient data transformations.
Distributed Systems Design
Ability to reason about consistency models, fault tolerance, scalability, and system architecture decisions in distributed computing contexts. Familiarity with concepts like sharding, replication, and consensus.
Data Pipeline Orchestration
Working knowledge of workflow orchestration platforms such as Airflow, Prefect, Dagster, or equivalent systems for scheduling, monitoring, and managing complex data workflows at scale.
Preferred
Web Crawling Stack
Nice to haveExperience with Scrapy, Selenium, or other web crawling frameworks. Familiarity with HTTP/HTML parsing libraries and challenges of robustness at scale.
Kubernetes and Container Orchestration
Nice to haveExperience deploying and scaling containerized workloads in Kubernetes or similar container orchestration systems for managing large-scale compute jobs.
Machine Learning Training Infrastructure
Nice to haveUnderstanding of how training data flows into machine learning systems, data loader patterns, and the relationship between data quality and model training efficiency.
Data Warehousing Concepts
Nice to haveFamiliarity with data warehouse design, columnar storage, OLAP patterns, and how data organization impacts query performance for analytics and monitoring.
Message Queue Systems
Nice to haveExperience with Kafka, Pulsar, or similar distributed message brokers for building high-throughput, decoupled data pipelines and real-time data processing.
Experimentation and A/B Testing Infrastructure
Nice to haveExperience building or working within frameworks for designing, executing, and analyzing large-scale experiments to measure causality and validate hypotheses.
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 180,000 – 280,000
Equity·Stock options
Benefits
Equity Compensation
Competitive stock option grants providing meaningful ownership in Cursor, an AI company at the forefront of coding automation. Equity vests over a standard four-year schedule with a one-year cliff.
Health and Wellness Coverage
Comprehensive health insurance including medical, dental, and vision coverage. Cursor supports proactive wellness through gym memberships and mental health resources.
Remote-First Work Environment
Flexibility to work remotely with optional access to the San Francisco headquarters. Cursor's flat organizational structure and talent-dense team enable effective collaboration regardless of location.
Professional Development
Opportunities to work on frontier AI infrastructure challenges, collaborating with leading researchers and engineers. Access to learning resources, conference attendance support, and mentorship from experienced infrastructure engineers.
Unlimited PTO
Flexible vacation policy respecting work-life balance while enabling unlimited time off for rest, travel, and personal priorities.
Hardware and Equipment
Provision of high-performance computing resources and development equipment necessary for effective systems engineering and infrastructure work.
Process
Interview steps.
- 01
Initial Screen
Casual conversation with a recruiting or engineering team member to discuss background, career trajectory, and interest in frontier AI infrastructure. Focus on understanding your experience with large-scale distributed systems and data engineering.
- 02
Technical Depth Interview
Detailed technical discussion with a senior engineer about your experience building and operating large-scale systems. Expect questions about specific systems you've built, architectural decisions you've made, debugging approaches, and your intuitions about distributed systems tradeoffs.
- 03
System Design Interview
Collaborative session designing solutions to problems relevant to data pipelines and crawling infrastructure. You'll discuss approaches to building scalable systems, handling failures, ensuring data quality, and monitoring system health. No coding required; focus on architectural thinking and problem-solving methodology.
- 04
Team and Values Fit
Conversation with multiple team members to assess cultural alignment and collaboration style. Cursor values truth-seeking engineers who enjoy spirited debate, are passionate about shipping code, and thrive in flat, talent-dense organizations. Discuss your experience working autonomously, iterating rapidly, and learning from failure.
- 05
Offer and Negotiation
Upon passing technical and cultural interviews, compensation discussion covering base salary, equity grants, benefits, and start date. Cursor is transparent about compensation ranges and open to negotiation based on experience and market conditions.
Full posting
Original listing.
Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.
About the role
We’re looking for Software Engineers to build the data systems behind our frontier coding models’ initial training. You’ll work on large-scale crawling, data platform, and pipeline infrastructure, turning raw dumps into the datasets our models train on, and making iteration with researchers fast and reliable.
The Data Quality team owns the entire road between raw internet-scale data and the tokens that train frontier models. This team makes sure the right data, in the right form, hits the training clusters on time and at the quality bar required to push the scaling curve. This is done by building our own models, our own high-performance pipelines, and by running the experiments that prove the data is actually stellar.
The Data Platform Team owns the infrastructure and pipelines that transform raw data dumps into training-ready datasets. This team improves the speed, reliability, and developer experience of our initial training data pipelines so researchers can quickly experiment with new data sources, quality filters, taxonomies, multimodal data, and data mixes that improve model performance.
The Crawling team owns large-scale web crawling and parsing that feeds the top of the funnel for initial training. They discover, schedule, fetch, and parse public web content so high-quality documents become the raw scrapes that Data Quality and Data Platform turn into tokens and mixes. This is deep distributed systems work with real ownership: host coverage and prioritization, fetch success under antibot and trap content, HTML/document parsing quality, and reliability of the crawl infrastructure that must continuously supply every downstream data pipeline.
What you’ll do
on the Data Quality Team:
Build and own high-throughput, fully telemetered data pipelines that process frontier-scale data with end-to-end traceability. If something breaks or drifts, your systems will tell us before the training run does.
Train and ship models that classify, rank, filter, clean, and identify data at extreme throughput. These models have to be both accurate and fast enough to sit in the critical path without becoming the bottleneck.
Design and run scaling-ladder experiments on data-mixture, repeatability, and quality depth that turn “this dataset feels good” into hard evidence the training team can trust.
Partner tightly with Data Acquisition to hunt down missing or low-quality sources, and with the training teams to close the loop on what actually moves loss and downstream evals.
Treat data quality as a systems problem and a research problem. You will write performance-critical code one week and design careful experiments the next.
on the Data Platform Team:
Build the platform that turns raw web, code, multimodal, and acquired data into training-ready datasets for frontier pretraining runs.
Own the pipelines, orchestration, and tooling that make pretraining data iteration fast, reliable, observable, and reproducible at scale.
Create clear signals for data quality, lineage, freshness, and pipeline health so researchers can trust what goes into each run.
Partner with initial training, crawling, data quality, and acquisition teams to turn new data ideas into measurable improvements in loss, evals, and model capability.
on the Crawling Team:
Build and scale the web crawling systems that discover, schedule, fetch, and parse high-quality documents across the open web for initial training.
Improve URL seeding, scoring, and fair host scheduling so crawl capacity lands on the hosts and pages that matter most for model quality.
Raise crawl success and parsing quality — defeating antibot failures, improving extractors, and capturing content we previously could not get cleanly.
Debug and harden complex crawl infrastructure end-to-end for availability, recovery, and ingestion lag, and automate delivery of crawl datasets into the data pipeline.
Work independently (and alongside AI agents) and partner with Data Quality and Data Platform so new coverage shows up as better tokens in training runs.
You may be a fit if
You have a strong infrastructure or data platform background, and ideally a spike of outlier depth somewhere (crawling/search infra is a plus, not a hard requirement)
You are a high-slope engineer who has moved unusually fast — for example, staff-level ownership within a few years — or you bring deep domain experience
You are able to architect and ship end-to-end with high ownership, debug complex systems independently, and work alongside AI agents
You have strong intuitions about large-scale distributed systems
You’re excited to learn how pre-training data shapes model quality, and want the ownership and visibility that comes with building systems that feed frontier training runs
#LI-DNI
Redirects to Cursor's application page.
Other roles
More at Cursor.
Software Engineer, RL Data
Mid
Field Engineer, Life Sciences
Mid
Field Engineer - India
Mid
Engineering Manager, Agent & Product Security
Manager
Engineering Manager, ML
Manager