Data Engineer

Mid · Full Time · Remote

Remote, US · RemoteUSD 83k – 207k1d ago
Apply for this role

Opens Benchling's application page

Role

What you'll do.

Benchling is seeking a Data Engineer to join its AI & Data Engineering (AIDE) team, responsible for building and operating production-grade data pipelines, Snowflake warehouse infrastructure, and analytics platforms that support the entire company's AI initiatives and cross-functional stakeholders. This role requires 3+ years of experience with cloud data warehouse systems, strong SQL and Python expertise, and proficiency with modern data orchestration tools like dbt and Airflow in an enterprise SaaS environment.

Responsibilities

  • Own Core Data Pipelines End to End: Build and operate production-grade ELT pipelines that move data from Benchling's product platform, Salesforce, and third-party systems into Snowflake. Model data using dbt with comprehensive testing, monitoring, and schema versioning to ensure reliability and scalability. These pipelines serve as the foundational infrastructure that the entire company builds upon, requiring software engineering rigor and performance optimization.
  • Build Data Foundation for AI Initiatives: Partner closely with AIDE's AI engineering team to ensure governed, trustworthy, and high-quality datasets are available for agentic AI tooling and internal AI applications. Enable AI-assisted workflows by curating and structuring data that meets accuracy, freshness, and compliance requirements for machine learning model consumption.
  • Own Data Governance and Pipeline Health: Maintain Snowflake access controls through role-based access control (RBAC), monitor data quality metrics, enforce PII-handling and data-access policies, and optimize warehouse costs and query performance. Establish data governance frameworks and implement automated quality checks to ensure data integrity across all enterprise systems.
  • Contribute to Platform Architecture Strategy: Collaborate with the data and AI engineering team on strategic decisions including cloud data warehouse architecture, semantic layer and metrics store design, and scaling data infrastructure. Provide technical input on system design decisions that impact reliability, performance, and cost efficiency across the organization.
  • Support Multi-Departmental Stakeholders: Translate ambiguous data requirements from GTM, Customer Success, Product, Finance, and other departments into scoped, buildable solutions. Provide technical leadership and communication across the organization to ensure data availability and quality supports critical business intelligence and analytics needs.

Qualifications

What we look for.

Technical

  • Production Data Pipeline Development

    3+ years of professional experience designing, building, and operating production-grade data pipelines with complete ownership of ingestion, transformation, and cloud data warehouse modeling. Demonstrated expertise in ELT architecture and real-world experience scaling pipelines to meet enterprise performance and reliability standards.

  • SQL and Python Proficiency

    Advanced SQL skills for complex data transformations, window functions, and query optimization. Production-grade Python experience for data processing, pipeline orchestration logic, and data quality testing. Ability to write maintainable, well-tested code following software engineering best practices.

  • Cloud Data Warehouse Expertise

    Production experience with Snowflake or comparable modern cloud data warehouse platforms such as BigQuery or Redshift. Deep understanding of data warehouse architecture, query optimization, partitioning strategies, and cost management in cloud-based analytics systems.

  • Data Modeling and dbt Framework

    Strong hands-on experience with data modeling methodologies including dimensional modeling, facts and dimensions, and slowly changing dimensions. Proficiency with dbt (data build tool) for building transformation workflows, version control, documentation, and testing within modern data stacks.

  • Data Orchestration and Workflow Tools

    Comfort with orchestration platforms such as Apache Airflow for scheduling, monitoring, and managing complex data job dependencies. Experience managing cron-based jobs, error handling, retry logic, and orchestration monitoring in production environments.

  • Software Engineering Practices for Data Systems

    Application of software engineering best practices to data infrastructure including version control (Git), code review processes, CI/CD pipelines, and automated testing frameworks. Experience managing data schemas, backward compatibility, and safe deployment patterns for data systems.

  • Cloud Infrastructure and DevOps

    Hands-on experience with AWS, GCP, or Azure cloud infrastructure supporting production data pipelines. Comfort with infrastructure-as-code, containerization, networking, and managing production system deployments and monitoring.

  • Data Privacy, Governance, and Compliance

    Working knowledge of data privacy frameworks, data governance practices, PII handling, RBAC implementation, and data quality testing frameworks. Understanding of regulatory compliance requirements and secure data management best practices in enterprise environments.

Education

  • Computer Science, Data Engineering, or Related Field

    Bachelor's degree in Computer Science, Data Engineering, Mathematics, Statistics, or related technical field preferred, though equivalent professional experience is highly valued. Strong foundation in computer science principles, algorithms, and database systems.

Experience

  • 3+ Years Production Data Engineering

    Minimum 3 years of professional experience building, deploying, and maintaining production data pipelines in enterprise or high-scale environments. Track record of owning critical data infrastructure that supports multiple internal stakeholders across different departments such as Sales, Customer Success, Product, and Finance.

  • Enterprise SaaS or Biotech Background

    Background in enterprise SaaS, life sciences, or biotech companies preferred. Experience with complex data integration from product platforms, CRM systems, and third-party tools. Familiarity with life science workflows and scientific data management advantageous but not required.

  • Cross-Functional Collaboration Experience

    Proven experience supporting many internal stakeholders across multiple departments rather than a single customer or team. Strong track record of translating business requirements into data solutions and communicating technical concepts to non-technical audiences.

  • Fast-Paced, Early-Stage Team Environment

    Comfort working in small, fast-moving, still-forming teams actively defining their own processes and practices. Experience collaborating in startups or newly established teams where roles are flexible and processes evolve rapidly. Openness to AI integration and learning about emerging technologies.

Skills

Required

  • SQL

    Advanced SQL for complex queries, data transformations, window functions, CTEs, and query optimization. Writing efficient queries that perform well at scale in Snowflake or comparable data warehouses.

  • Python

    Production-grade Python for data processing, ETL script development, and orchestration logic. Experience with pandas, testing frameworks, and writing maintainable, well-documented code.

  • dbt (Data Build Tool)

    Hands-on experience building transformation workflows with dbt, managing data lineage, implementing data quality tests, and following dbt best practices. Experience with version control integration and documentation generation.

  • Snowflake

    Production experience with Snowflake data warehouse including database design, query optimization, role-based access control (RBAC), and cost management. Understanding of Snowflake architecture and performance tuning.

  • Apache Airflow

    Hands-on experience orchestrating data pipelines with Apache Airflow. Understanding of DAGs, task dependencies, monitoring, error handling, and scaling Airflow deployments in production.

  • Git Version Control

    Proficiency with Git for code version control, managing branches, pull requests, code reviews, and CI/CD integration. Experience working in collaborative development environments with peer review processes.

  • Cloud Infrastructure (AWS/GCP/Azure)

    Working knowledge of major cloud platforms and services. Experience with cloud-based compute, storage, networking, and infrastructure needed to support production data pipelines.

  • Data Modeling Methodologies

    Strong understanding of data modeling approaches including star schema, dimensional modeling, and normalization. Experience designing scalable warehouse schemas that support analytics and reporting needs.

  • Data Quality and Testing

    Implementation of automated data quality frameworks, validation testing, schema monitoring, and data integrity checks. Experience with tools and patterns that ensure data reliability at scale.

Preferred

  • Product Analytics and BI Tools

    Nice to have

    Familiarity with modern business intelligence platforms such as Sigma, Omni, Looker, or Tableau deployed in self-service analytics models. Experience with product behavioral data and event instrumentation.

  • GTM Analytics and Salesforce

    Nice to have

    Experience with go-to-market analytics and Salesforce integration. Knowledge of sales pipeline data, customer lifecycle analytics, and GTM data models.

  • AI/ML Observability and Telemetry

    Nice to have

    Exposure to AI usage telemetry, LLM observability data, and supporting machine learning systems with curated datasets. Understanding of AI model monitoring and performance tracking.

  • Metrics Layer and Semantic Modeling

    Nice to have

    Experience building or maintaining a metrics layer or semantic layer such as dbt metrics, Cube, or similar platforms. Knowledge of defining enterprise metrics and ensuring consistency across analytics tools.

  • Event-Driven Architecture

    Nice to have

    Experience with product event taxonomy governance, event instrumentation, and event-driven data pipelines. Familiarity with real-time data processing patterns.

  • Life Sciences Domain Knowledge

    Nice to have

    Background in life sciences, biotechnology, or biotech research. Understanding of scientific workflows, experimental data management, and domain-specific data challenges.

  • Data Integration Tools

    Nice to have

    Experience with data integration and ELT tools such as Fivetran, Stitch, or similar platforms for moving data from source systems into data warehouses.

Compensation

Pay and benefits.

Base·USD 83,000 – 207,000

Equity·Stock options

Benefits

  • Equity and Stock Options

    Competitive equity stake in Benchling with stock options as part of the compensation package, providing significant upside participation in the company's success as a well-funded life sciences technology leader.

  • Comprehensive Health and Wellness

    Full medical, dental, and vision coverage with subsidized premiums. Mental health support, therapy services, and wellness programs designed to support employee wellbeing and work-life balance.

  • Remote Work Flexibility

    Fully remote position with flexibility to work from anywhere, eliminating commute time and enabling geographic flexibility. Access to Benchling's distributed team infrastructure and collaboration tools.

  • Professional Development

    Learning and development budget for courses, conferences, certifications, and training. Access to internal resources for continuous skill development, including AI literacy programs and domain expertise in life sciences.

  • Competitive Time Off

    Generous paid time off, flexible vacation policies, and sabbatical opportunities. Support for work-life integration and mental health through unlimited PTO frameworks or similar policies.

  • Retirement Planning

    401(k) retirement plan with employer matching contributions. Financial planning resources and retirement planning support.

Process

Interview steps.

  1. 01

    Initial Recruiter Screening

    Introductory conversation with Benchling's recruiting team to discuss background, career trajectory, and interest in the Data Engineer role. Assessment of baseline qualifications and alignment with the AI & Data Engineering (AIDE) team's mission of supporting company-wide data and AI initiatives.

  2. 02

    AI-Focused Capabilities Assessment

    Brief AI-focused exercise or discussion designed to understand how you think about and use AI tools to drive impact in your role. Benchling values AI fluency as foundational to all work; candidates should come prepared to discuss AI tools they use, their workflows, and how they leverage AI to solve data and engineering challenges.

  3. 03

    Technical Screening Interview

    Deep-dive technical conversation with a senior data engineer covering data pipeline architecture, SQL optimization, dbt best practices, Snowflake data warehouse design, and orchestration patterns. Expect discussions around system design, trade-offs, and real-world production data engineering scenarios.

  4. 04

    Data Architecture and Strategy Discussion

    Conversation with data team leadership about architectural decisions, scaling challenges, metrics layer design, and governance frameworks. Discussion of how you approach stakeholder management across multiple departments and translate business requirements into data infrastructure decisions.

  5. 05

    Cross-Functional Alignment Interview

    Conversation with stakeholders from GTM, Product, Finance, or Customer Success teams that rely on data infrastructure. Assessment of communication style, ability to translate technical concepts, and track record of supporting diverse internal customers with different data needs.

  6. 06

    Team and Culture Fit Discussion

    Final conversation with AIDE team members to assess cultural alignment, comfort in fast-moving early-stage team environments, and openness to AI integration and learning. Discussion of team processes, how the team is defining its own practices, and collaborative work style.

Full posting

Original listing.

We are rebuilding biotech for the AI era.

When a breakthrough is delayed, the world waits. Getting a molecule from discovery to patients, or a crop from lab to field, involves thousands of slow, manual, disconnected steps. AI has the potential to change this, compressing decades of R&D work into years. But that only happens when clean, structured scientific data and AI are built into how science gets done.

Benchling is the AI platform for biotech R&D. Scientists use Benchling to design experiments, capture structured data, and run AI agents and models directly in their workflows. Over 200,000 scientists around the world trust Benchling to power their most important work, from academic labs to Sanofi, Moderna, and more than half of the world's top 50 biopharma.

We’re building an AI scientist for our customers. We can’t do that if we haven’t built the muscle ourselves. AI fluency is the foundation we build on; it's core to how we work, and we're committed to helping every new hire integrate it into their day-to-day. As part of our interview process, you'll complete a brief AI-focused exercise or discussion so we can understand how you think about and use AI to drive impact in your role. Feel free to reference any tools, platforms, or workflows you use today.

ROLE OVERVIEW

Biotechnology is rewriting life as we know it, from the medicines we take, to the crops we grow, the materials we wear, and the household goods that we rely on every day. But moving at the new speed of science requires better technology. Benchling's mission is to unlock the power of biotechnology. The world's most innovative biotech companies use Benchling's R&D Cloud to power the development of breakthrough products and accelerate time to milestone and market. Come help us bring modern software to modern science.

Benchling is building AI & Data Engineering (AIDE), a small, autonomous team within our Security & IT organization. AIDE owns three things: internal AI tooling, adoption, and AI-assisted workflows across the company; cross-functional and company-wide agentic AI applications that no single department owns; and the enterprise data engineering, analytics architecture, and source-of-truth datasets that everything above depends on. AIDE’s data and analytics functions grew out of our former Data, Analytics & Systems (DAS) team, and this role carries forward DAS's original charter: building and running the data pipelines, warehouse, and analytics infrastructure that the entire company relies on for trustworthy answers.

This is a data engineering role — we want someone who builds and operates reliable, production-grade data pipelines and warehouse infrastructure, not a data scientist focused on modeling or analysis.

This role exists because AIDE's data function supports the whole company — GTM, Customer Success, Product, Finance, and beyond — not just one team, and the team needs to grow to support these initiatives as we expand the team’s scope and portfolio. You'll own core pipelines end to end (ingestion, transformation, warehouse, and the BI/analytics layer on top), partner with the rest of the data team on the team's data architecture, and help build the trusted data foundation that AIDE's AI-adoption and agentic AI work increasingly depends on.

Check out our engineering blog for examples of past work across Benchling.

 

RESPONSIBILITIES

  • Own core data pipelines end to end: Build and operate the ELT pipeline that moves data from Benchling's product, Salesforce, and third-party systems into Snowflake, modeled with dbt, and built to production standards — testing, monitoring, schema versioning — that hold up as usage scales. This is infrastructure the rest of the company builds on, not a one-off project.

  • Build the data foundation for AIDE's AI initiatives: Partner with AIDE's AI engineering side to make governed, trustworthy data available for the agentic AI tooling and internal AI applications the team ships.

  • Own data governance and pipeline health: Maintain Snowflake access controls (RBAC), monitor data quality, uphold PII-handling and data-access policy, and manage warehouse cost and performance as usage grows.

  • Contribute to platform strategy: Weigh in on bigger structural decisions — warehouse architecture, semantic layer/metrics store design— alongside the rest of the data and AI engineering team.

 

QUALIFICATIONS

  • 3+ years of professional experience building and operating production data pipelines — ingestion, transformation, and modeling data into a cloud data warehouse.

  • Strong SQL and Python skills; hands on experience with data modeling methodologies and tools, preferable with dbt.

  • Experience applying software engineering practices to data systems — version control, code review, CI/CD, automated testing — and comfort working with cloud infrastructure (AWS or similar) supporting production pipelines.

  • Experience with Snowflake or a comparable modern cloud data warehouse in production.

  • Comfort with orchestration tooling (Airflow or similar) for scheduled data jobs.

  • Track record supporting many stakeholders across departments such as Sales, CS, Product, Finance, rather than a single internal customer.

  • Understanding of data privacy, governance, quality, and testing frameworks and best practices.

  • Strong communication skills; comfortable translating ambiguous requests from non-technical stakeholders into a scoped, buildable data solution.

  • Comfortable in a small, fast-moving, still-forming team — AIDE only stood up in its current form in mid-2026 and is actively defining its own processes.

  • Interest in learning more about life science (prior knowledge is not required).

NICE TO HAVE

  • Familiarity with product behavioral data and a modern BI tool (Sigma, Omni, Looker, Tableau) deployed in a self-service model.

  • Experience with product/usage analytics instrumentation and event-taxonomy governance.

  • Familiarity with GTM analytics tools such as Salesforce.

  • Exposure to AI-usage telemetry, LLM observability data, or supporting AI/ML tooling with curated data.

  • Background in enterprise SaaS, life sciences, or biotech.

  • Experience building or maintaining a metrics layer. 

 

#LI-Remote

#BI-Remote

#LI-CG1

Benchling welcomes everyone.

We believe diversity enriches our team so we hire people with a wide range of identities, backgrounds, and experiences.

We are an equal opportunity employer. That means we don’t discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. We also consider for employment qualified applicants with arrest and conviction records, consistent with applicable federal, state and local law, including but not limited to the San Francisco Fair Chance Ordinance.

Redirects to Benchling's application page.

Other roles

More at Benchling.

View all 10 roles