Data Engineer
Mid · Full Time
Opens Benchling's application page
Role
What you'll do.
Benchling seeks a Data Engineer to build and operate production-grade data pipelines and cloud warehouse infrastructure for the AI & Data Engineering (AIDE) team. You'll own end-to-end ELT processes across Snowflake, dbt, and orchestration tools while establishing the trusted data foundation supporting company-wide analytics, AI initiatives, and cross-functional stakeholders. This role requires 3+ years of professional data pipeline experience with strong SQL, Python, and software engineering practices applied to data systems.
Responsibilities
- Build and Operate End-to-End ELT Pipelines: Design, build, and maintain production-grade data pipelines that move data from Benchling's product platform, Salesforce, and third-party systems into Snowflake. Implement data modeling with dbt following production standards including comprehensive testing, continuous monitoring, schema versioning, and performance optimization to support company-wide infrastructure at scale.
- Establish Data Foundation for AI Initiatives: Collaborate with AIDE's AI engineering team to design and deliver governed, trustworthy datasets for agentic AI tooling and internal AI applications. Ensure data quality, documentation, and accessibility requirements are met to support machine learning and AI workload execution across the organization.
- Manage Data Governance and Warehouse Operations: Oversee Snowflake role-based access control (RBAC) implementation, monitor data quality metrics, enforce PII handling and data-access policies, and optimize warehouse cost and performance as usage grows. Maintain data lineage tracking and compliance documentation for enterprise data governance standards.
- Contribute to Data Architecture and Strategy: Participate in architectural decision-making regarding warehouse design patterns, semantic layer and metrics store implementation, data modeling frameworks, and technology stack evolution. Partner with cross-functional data and AI engineering teams to shape the data platform roadmap and best practices.
- Support Multi-Stakeholder Data Needs: Partner with GTM, Customer Success, Product, Finance, and other departments to translate ambiguous data requirements into scoped, buildable solutions. Provide technical guidance on data availability, quality, and governance to internal business customers across the organization.
Qualifications
What we look for.
Technical
SQL and Python Proficiency
Advanced SQL skills for complex query optimization, window functions, and performance tuning. Strong Python experience for data transformation logic, pipeline orchestration, and systems integration.
dbt (Data Build Tool) Expertise
Hands-on production experience with dbt for data modeling, transformation logic, testing frameworks, and CI/CD integration. Familiarity with dbt best practices, incremental models, macros, and documentation generation.
Cloud Data Warehouse Operation
3+ years of professional experience building and operating Snowflake or equivalent modern cloud data warehouse in production environments. Understanding of warehouse architecture, query optimization, cost management, and performance tuning.
Data Pipeline Orchestration
Practical experience with Apache Airflow or comparable orchestration frameworks for scheduling, monitoring, and managing complex data workflows. Understanding of DAG (directed acyclic graph) design patterns and error handling.
Software Engineering Practices for Data Systems
Ability to apply production software engineering methodologies to data infrastructure including version control (Git), code review processes, CI/CD pipelines, automated testing frameworks, and infrastructure-as-code principles.
Cloud Infrastructure and DevOps
Working knowledge of AWS, GCP, or Azure cloud infrastructure supporting data pipelines. Comfort with cloud IAM, networking, cost optimization, and monitoring tools relevant to data platform operations.
Data Modeling and Normalization
Strong understanding of data modeling methodologies including dimensional modeling, star schema design, normalization principles, and slowly changing dimensions. Experience translating business requirements into optimized data models.
Data Privacy, Governance, and Quality
Knowledge of data privacy frameworks (GDPR, CCPA), data governance best practices, PII identification and masking, data lineage tracking, and quality assurance frameworks for enterprise data systems.
Education
Bachelor's Degree in Computer Science, Engineering, or Related Field
Formal education in computer science, software engineering, data science, or equivalent technical discipline demonstrating foundational knowledge of systems design and software development principles.
Experience
3+ Years Production Data Pipeline Development
Professional experience building, deploying, and maintaining production data pipelines across ingestion, transformation, and modeling layers into cloud data warehouses at scale.
Multi-Stakeholder Data Support
Track record supporting diverse internal business customers across multiple departments (Sales, Customer Success, Product, Finance) rather than supporting a single team or use case.
Production Data Warehouse Administration
Experience managing production cloud data warehouse infrastructure including schema design, access controls, performance monitoring, cost optimization, and reliability operations.
Skills
Required
SQL
Advanced SQL querying, optimization, and performance tuning for complex analytical and operational queries.
Python
Proficient Python programming for data transformation, pipeline development, and systems integration.
dbt
Production-level data build tool expertise for data modeling, testing, and transformation logic.
Snowflake
Cloud data warehouse architecture, query optimization, cost management, and access control implementation.
Apache Airflow
Data pipeline orchestration, DAG design, monitoring, and error handling.
Data Modeling
Dimensional modeling, schema design, and normalization frameworks for analytical databases.
Git and Version Control
Source code version control, branching strategies, and collaborative development workflows.
Data Governance and Privacy
PII identification, data access controls, compliance frameworks, and quality assurance practices.
Preferred
BI Tools (Sigma, Omni, Looker, Tableau)
Nice to haveExperience with modern business intelligence platforms and self-service analytics implementation.
Product and Usage Analytics
Nice to haveFamiliarity with product behavioral data, event taxonomy governance, and product analytics instrumentation.
Salesforce Analytics
Nice to haveExperience with Salesforce ecosystem and GTM analytics tools for sales and customer success teams.
AI and ML Observability
Nice to haveExposure to AI usage telemetry, LLM observability data, or supporting AI/ML systems with curated data pipelines.
Metrics Layer Implementation
Nice to haveExperience building or maintaining semantic layers or metrics stores (dbt metrics, LookML, Tableau calculations).
SaaS or Life Sciences Domain
Nice to haveBackground in enterprise SaaS companies or life sciences/biotech industry providing contextual understanding of domain challenges.
CI/CD Pipeline Configuration
Nice to haveExperience setting up and maintaining continuous integration and deployment pipelines for data systems.
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 153,000 – 207,000
Equity·Stock options
Benefits
Equity and Stock Options
Meaningful equity grants aligning your success with company growth as a key part of total compensation.
Health, Dental, and Vision Insurance
Comprehensive medical, dental, and vision coverage for employees and their families.
401(k) Retirement Plan
Employer-sponsored retirement savings plan with matching contributions.
Flexible Paid Time Off
Flexible PTO policy supporting work-life balance and personal well-being.
Hybrid Work Flexibility
Flexible hybrid arrangement requiring 3 days per week on-site (Monday, Tuesday, Thursday) with remote work flexibility.
Professional Development and Learning
Opportunities to expand technical skills, participate in industry conferences, and develop expertise in life sciences and AI technologies.
AI Tool Integration
Access to and encouragement to use AI tools and platforms to enhance productivity and innovation in your day-to-day work.
Process
Interview steps.
- 01
Initial Screening Call
Introductory conversation with recruiter to discuss background, motivation, and high-level alignment with the Data Engineer role and AIDE team mission.
- 02
Technical Deep Dive Interview
Conversation with data engineering team members covering SQL/Python proficiency, dbt experience, data modeling approaches, production pipeline architecture, and troubleshooting complex data platform scenarios.
- 03
AI Fluency Discussion or Exercise
Benchling prioritizes AI fluency across the organization. As part of this assessment, you'll complete a brief AI-focused exercise or participate in a discussion demonstrating how you think about and use AI tools (ChatGPT, Copilot, etc.) to drive impact in data engineering. Feel free to reference any AI tools or workflows you currently use.
- 04
Data Architecture and Strategy Round
Conversation with senior data team members or AI & Data Engineering leadership discussing warehouse architecture decisions, metrics layer design, data governance approaches, and how you'd approach scaling data infrastructure for a growing organization.
- 05
Cross-Functional Stakeholder Perspective
Conversation or case study exercise exploring how you've managed competing requirements from multiple internal stakeholders, translated ambiguous data requests into scoped solutions, and provided technical guidance to non-technical partners.
- 06
Final Leadership Interview
Conversation with AIDE team leadership or broader engineering management discussing long-term vision, team collaboration in a fast-moving environment, and how you'd approach learning about life sciences and biotech context.
Full posting
Original listing.
We are rebuilding biotech for the AI era.
When a breakthrough is delayed, the world waits. Getting a molecule from discovery to patients, or a crop from lab to field, involves thousands of slow, manual, disconnected steps. AI has the potential to change this, compressing decades of R&D work into years. But that only happens when clean, structured scientific data and AI are built into how science gets done.
Benchling is the AI platform for biotech R&D. Scientists use Benchling to design experiments, capture structured data, and run AI agents and models directly in their workflows. Over 200,000 scientists around the world trust Benchling to power their most important work, from academic labs to Sanofi, Moderna, and more than half of the world's top 50 biopharma.
We’re building an AI scientist for our customers. We can’t do that if we haven’t built the muscle ourselves. AI fluency is the foundation we build on; it's core to how we work, and we're committed to helping every new hire integrate it into their day-to-day. As part of our interview process, you'll complete a brief AI-focused exercise or discussion so we can understand how you think about and use AI to drive impact in your role. Feel free to reference any tools, platforms, or workflows you use today.
ROLE OVERVIEW
Biotechnology is rewriting life as we know it, from the medicines we take, to the crops we grow, the materials we wear, and the household goods that we rely on every day. But moving at the new speed of science requires better technology. Benchling's mission is to unlock the power of biotechnology. The world's most innovative biotech companies use Benchling's R&D Cloud to power the development of breakthrough products and accelerate time to milestone and market. Come help us bring modern software to modern science.
Benchling is building AI & Data Engineering (AIDE), a small, autonomous team within our Security & IT organization. AIDE owns three things: internal AI tooling, adoption, and AI-assisted workflows across the company; cross-functional and company-wide agentic AI applications that no single department owns; and the enterprise data engineering, analytics architecture, and source-of-truth datasets that everything above depends on. AIDE’s data and analytics functions grew out of our former Data, Analytics & Systems (DAS) team, and this role carries forward DAS's original charter: building and running the data pipelines, warehouse, and analytics infrastructure that the entire company relies on for trustworthy answers.
This is a data engineering role — we want someone who builds and operates reliable, production-grade data pipelines and warehouse infrastructure, not a data scientist focused on modeling or analysis.
This role exists because AIDE's data function supports the whole company — GTM, Customer Success, Product, Finance, and beyond — not just one team, and the team needs to grow to support these initiatives as we expand the team’s scope and portfolio. You'll own core pipelines end to end (ingestion, transformation, warehouse, and the BI/analytics layer on top), partner with the rest of the data team on the team's data architecture, and help build the trusted data foundation that AIDE's AI-adoption and agentic AI work increasingly depends on.
Check out our engineering blog for examples of past work across Benchling.
RESPONSIBILITIES
Own core data pipelines end to end: Build and operate the ELT pipeline that moves data from Benchling's product, Salesforce, and third-party systems into Snowflake, modeled with dbt, and built to production standards — testing, monitoring, schema versioning — that hold up as usage scales. This is infrastructure the rest of the company builds on, not a one-off project.
Build the data foundation for AIDE's AI initiatives: Partner with AIDE's AI engineering side to make governed, trustworthy data available for the agentic AI tooling and internal AI applications the team ships.
Own data governance and pipeline health: Maintain Snowflake access controls (RBAC), monitor data quality, uphold PII-handling and data-access policy, and manage warehouse cost and performance as usage grows.
Contribute to platform strategy: Weigh in on bigger structural decisions — warehouse architecture, semantic layer/metrics store design— alongside the rest of the data and AI engineering team.
QUALIFICATIONS
3+ years of professional experience building and operating production data pipelines — ingestion, transformation, and modeling data into a cloud data warehouse.
Strong SQL and Python skills; hands on experience with data modeling methodologies and tools, preferable with dbt.
Experience applying software engineering practices to data systems — version control, code review, CI/CD, automated testing — and comfort working with cloud infrastructure (AWS or similar) supporting production pipelines.
Experience with Snowflake or a comparable modern cloud data warehouse in production.
Comfort with orchestration tooling (Airflow or similar) for scheduled data jobs.
Track record supporting many stakeholders across departments such as Sales, CS, Product, Finance, rather than a single internal customer.
Understanding of data privacy, governance, quality, and testing frameworks and best practices.
Strong communication skills; comfortable translating ambiguous requests from non-technical stakeholders into a scoped, buildable data solution.
Comfortable in a small, fast-moving, still-forming team — AIDE only stood up in its current form in mid-2026 and is actively defining its own processes.
Interest in learning more about life science (prior knowledge is not required).
NICE TO HAVE
Familiarity with product behavioral data and a modern BI tool (Sigma, Omni, Looker, Tableau) deployed in a self-service model.
Experience with product/usage analytics instrumentation and event-taxonomy governance.
Familiarity with GTM analytics tools such as Salesforce.
Exposure to AI-usage telemetry, LLM observability data, or supporting AI/ML tooling with curated data.
Background in enterprise SaaS, life sciences, or biotech.
Experience building or maintaining a metrics layer.
HOW WE WORK
We offer a flexible hybrid work arrangement that prioritizes in-office collaboration. Employees are expected to be on-site 3 days per week (Monday, Tuesday, and Thursday).
#LI-Hybrid
#BI-Hybrid
#LI-CG1
Benchling welcomes everyone.
We believe diversity enriches our team so we hire people with a wide range of identities, backgrounds, and experiences.
We are an equal opportunity employer. That means we don’t discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. We also consider for employment qualified applicants with arrest and conviction records, consistent with applicable federal, state and local law, including but not limited to the San Francisco Fair Chance Ordinance.
Redirects to Benchling's application page.
Other roles
More at Benchling.
Data Engineer
Mid
Software Engineer, Platform (Developer Experience)
Mid
Software Engineer, Platform (Developer Experience)
Mid
Software Engineer, Model Evaluation and Improvement
Mid
Software Engineer, Full Stack (Document Canvas)
Mid