Staff Data Engineer
Data Engineer · Staff · Full Time · Remote
Opens Hims & Hers's application page
Role
What you'll do.
Staff Data Engineer at Hims & Hers, a leading telehealth and digital health platform, responsible for architecting and operating critical data platform infrastructure that powers analytics for millions of patients. This hands-on Staff role drives cross-squad technical initiatives across BigQuery, dbt, Airflow, Kafka, and Databricks while mentoring senior engineers and establishing data governance standards in a regulated healthcare environment.
Responsibilities
- Lead Multi-Squad Platform Initiatives: Serve as DRI (Directly Responsible Individual) for high-complexity, multi-sprint platform initiatives spanning the Data Platform Engineering team, including Fivetran connector buildouts, Databricks Lakehouse migration workstreams, event streaming infrastructure implementations, and engineering standards adoption across nine engineers. Own end-to-end delivery from conception through production deployment with cross-functional coordination.
- Design and Maintain Production Data Pipeline Infrastructure: Architect, build, and operate production-grade data ingestion pipelines and platform infrastructure spanning the full modern data stack—BigQuery, dbt, Airflow on Astronomer, Confluent Kafka, Databricks, and Fivetran. Own the Bronze and Silver layer implementations that Analytics Engineering, Data Science, and business intelligence teams depend on daily for patient care decisions and operational analytics.
- Engineer Event-Driven and Streaming Data Systems: Design, implement, and operate real-time event streaming pipelines using Kafka, PySpark, and Databricks Structured Streaming in production environments. Define scaling strategies, establish cost guardrails, implement consumer lag alerting mechanisms, and create comprehensive runbooks before production deployment to ensure reliability at enterprise scale.
- Establish Data Contracts and Schema Governance: Own end-to-end data contracts, schema governance policies, and service-level agreements (SLAs) for ingestion and raw-to-cleansed data transformations. Implement Schema Registry governance across Kafka topics and manage schema evolution patterns to prevent downstream breakage and ensure data consistency across disparate systems.
- Implement Data Quality and Observability: Own data quality for all pipelines through dbt tests, anomaly detection frameworks, and schema change monitoring with automatic drift alerting. Establish observable pipelines from inception—write quality gates as code, implement Datadog monitoring and alerting before production deployment, and participate in on-call rotation for Tier 1 operational ownership.
- Operate Systems with Reliability Focus: Establish key performance indicators (KPIs) and service-level objectives (SLOs) for data systems. Implement comprehensive Datadog monitoring and alerting as infrastructure-as-code, participate in on-call rotation, own Tier 1 operational tickets and runbooks, and drive systemic efficiency improvements across platform-owned pipelines and infrastructure through root cause analysis.
- Manage Data Integration and Activation Layer: Own end-to-end integration and data activation pipelines through Fivetran connectors and Hightouch reverse ETL platforms. Manage infrastructure-as-code provisioning using Terraform/OpenTofu, production monitoring, schema change governance, connector health monitoring, and operational excellence across data movement tools in healthcare compliance environment.
- Enable Downstream Data Teams: Support Analytics Engineers, Data Scientists, and ML engineers by building platform capabilities and data pipelines that unblock their development roadmaps. Partner with legal, security, and DevOps teams on compliance controls and infrastructure hardening while maintaining clear separation of concerns—platform layer owned by Data Platform Engineering, transformation and model readiness by Analytics Engineering.
- Mentor Senior Data Engineers: Guide Senior Data Engineers through design reviews, code reviews, and collaborative pair programming sessions. Help senior engineers expand from squad-level scope to cross-squad technical leadership responsibilities. Drive adoption of engineering standards including testing practices, CI/CD patterns, observability-as-code, and Schema Registry governance across the team.
- Drive Engineering Standards and Best Practices: Contribute to and drive adoption of platform engineering standards, testing practices, and CI/CD patterns across the Data Platform Engineering organization. Participate in Architecture Review Committee (ARC) reviews for changes with cross-team or cost implications. Establish and document standards for production-grade pipeline development and infrastructure-as-code practices.
Qualifications
What we look for.
Technical
Data Platform Architecture and Design
Demonstrated expertise designing and operating production-grade data ingestion pipelines, platform infrastructure, and modern data stack systems. Proficiency in multi-cloud environments spanning GCP (BigQuery) and AWS (EKS, data services). Strong understanding of Bronze/Silver/Gold medallion architecture patterns, data contracts, and schema governance at enterprise scale.
Real-Time Data Processing and Streaming
Hands-on experience building and operating production Kafka or Confluent Kafka clusters including producers, consumers, schema evolution, Schema Registry governance, and consumer lag management. Proficiency with PySpark and Databricks Structured Streaming for large-scale stream processing. Experience implementing CDC (Change Data Capture) patterns for real-time ingestion from operational databases.
BigQuery and Databricks Proficiency
Production experience administering and governing dbt in BigQuery or Databricks environments including CI/CD configuration, testing standards, documentation governance, and platform-level schema governance. Hands-on dbt implementation experience for ingestion-layer (Bronze/Silver) transformations. Experience with Databricks Delta Lake, Databricks Workflows, and Unity Catalog governance framework.
Orchestration and Workflow Management
Proven experience building and operating Airflow DAGs at scale on Astronomer. Deep understanding of task-level orchestration patterns, DAG reliability engineering, dynamic DAG generation, multi-priority scheduling strategies, and performance optimization. Experience with SLA management and alerting for workflow dependencies.
Data Integration Tools and ETL Platforms
Hands-on experience with Fivetran or equivalent connector platforms including infrastructure-as-code provisioning, schema change handling, connector health monitoring, and operational troubleshooting. Experience with Hightouch or equivalent reverse ETL platforms for data activation and downstream system synchronization.
Infrastructure-as-Code and DevOps Practices
Strong proficiency with Terraform or OpenTofu for infrastructure provisioning and management. Treat infrastructure changes with same rigor as application code changes. Experience with CI/CD pipelines, version control integration, and GitOps workflows. Comfort with AWS EKS orchestration for Airflow and other containerized data services.
Data Quality, Testing, and Observability
Production experience implementing comprehensive data quality frameworks using dbt tests, anomaly detection tools, and schema validation. Proficiency with Datadog or equivalent observability platforms for metrics, logging, and alerting-as-code. Experience establishing data SLOs and implementing monitoring for data freshness, completeness, and accuracy metrics.
Programming and SQL Competency
Strong Python skills for production-grade data pipeline development, testing, and code review. Expert-level SQL proficiency for complex analytical queries, window functions, and performance optimization. Ability to write, review, and raise quality standards for production data engineering code.
Healthcare Data Compliance and Governance
Familiarity with regulated healthcare environment requirements including HIPAA/PHI (Protected Health Information) handling, role-based access controls (RBAC), audit logging, encryption standards, and data residency requirements. Understanding of data compliance frameworks and healthcare-specific data governance practices.
Education
Bachelor's Degree in Computer Science or Related Field
Bachelor's degree in Computer Science, Software Engineering, Data Science, Mathematics, or related technical discipline. Equivalent professional experience may substitute for formal education.
Experience
8+ Years Data Platform Engineering
Minimum eight years of professional experience designing, building, and operating data pipelines, data infrastructure, and data platform systems. Track record of owning complex, mission-critical data systems from architecture through production deployment and operational maintenance. Experience scaling data systems to support hundreds or thousands of concurrent users and millions of data points.
Staff or Senior Level Data Engineering Role
Proven experience at Staff, Principal, or Senior level data engineering roles with cross-squad or cross-team scope. Demonstrated ability to drive large technical initiatives end-to-end, mentor junior engineers, and establish technical standards. Track record of unblocking downstream teams and driving platform-level improvements.
Regulated Industry or Healthcare Experience
Prior experience working in regulated industries such as healthcare, fintech, or pharmaceutical sectors. Familiarity with healthcare data handling, HIPAA compliance requirements, audit logging practices, and regulated environment best practices. Experience at direct-to-consumer healthcare, telehealth, or similarly regulated companies is a plus.
Multi-Cloud Infrastructure Management
Production experience operating data systems across multiple cloud providers, specifically GCP and AWS. Comfort with cloud cost optimization, multi-cloud architecture patterns, and cloud-native data services including BigQuery, Redshift, and managed Kafka offerings.
Skills
Required
BigQuery
Production-grade expertise with Google BigQuery including data warehouse design, query optimization, cost management, and advanced features like BQML and BigQuery DataTransfer Service integration.
dbt (data build tool)
Advanced dbt proficiency including CI/CD pipeline integration, testing frameworks, documentation-as-code, YAML configuration, Jinja templating, and governance at platform scale. Experience with both ingestion and transformation-layer dbt implementations.
Apache Airflow (Astronomer)
Advanced Airflow proficiency including DAG design patterns, dynamic DAG generation, task dependencies, branching logic, sensor implementations, SLA management, and performance tuning on Astronomer platform.
Confluent Kafka
Production Kafka expertise including cluster administration, topic management, producer/consumer implementations, schema evolution, Confluent Schema Registry governance, consumer group management, and lag monitoring.
Databricks
Comprehensive Databricks platform experience including Delta Lake, Databricks SQL, Databricks Workflows, Unity Catalog governance, cluster optimization, and PySpark for large-scale transformations.
Terraform or OpenTofu
Advanced infrastructure-as-code proficiency with Terraform or OpenTofu including module design, state management, CI/CD integration, version control workflows, and cross-cloud provisioning.
Python
Production-grade Python development for data engineering including package management, testing frameworks, code review standards, performance optimization, and integration with data platforms.
SQL
Expert SQL proficiency including complex window functions, subqueries, performance optimization, explain plans, query tuning, and advanced analytical queries for data warehouse environments.
Fivetran
Hands-on Fivetran experience including connector configuration, transformation logic, monitoring and alerting, API-based automation, and operational troubleshooting for data integration scenarios.
Change Data Capture (CDC)
Production experience implementing CDC patterns for real-time data ingestion including log-based CDC, query-based CDC, and CDC integration with streaming platforms like Kafka.
Preferred
PySpark and SparkSQL
Nice to haveAdvanced PySpark expertise for large-scale distributed data processing, DataFrame operations, optimization strategies, and Spark SQL for complex analytical transformations.
Hightouch
Nice to haveExperience with Hightouch or equivalent reverse ETL platforms for data activation, downstream system synchronization, and business process automation through data movement.
MLOps and Feature Engineering
Nice to haveExperience supporting ML engineering teams through production data pipelines, feature stores, model training datasets, experiment tracking infrastructure, and ML-specific data governance.
Looker LookML
Nice to haveFamiliarity with Looker LookML or equivalent BI semantic layer tools for defining metrics, dimensions, explores, and governance at the presentation layer.
Go Programming
Nice to haveExperience with Go programming language for building Kafka service components, microservices, and data infrastructure tooling.
GDPR and EU Data Compliance
Nice to haveFamiliarity with UK and GDPR data compliance requirements distinct from US HIPAA, including data residency, processing agreements, and cross-border data transfer restrictions.
Datadog Observability
Nice to haveProduction experience with Datadog for infrastructure monitoring, log aggregation, APM (Application Performance Monitoring), and alerting-as-code for data pipelines.
Compensation
Pay and benefits.
Base·USD 170,000 – 200,000
Equity·Stock options
Full posting
Original listing.
Hims & Hers is the leading health and wellness platform, on a mission to help the world feel great through the power of better health. We are redefining healthcare by putting the customer first and delivering access to care that is affordable, accessible, and personal, from diagnosis to treatment to delivery. No two people are the same, so we provide access to personalized care designed for results. By normalizing health & wellness challenges and innovating on their solutions, we’re making better health outcomes easier to achieve.
Hims & Hers is a public company, traded on the NYSE under the ticker symbol “HIMS.” To learn more about the brand and offerings, you can visit hims.com/about and hims.com/how-it-works . For information on the company’s outstanding benefits, culture, and its talent-first flexible/remote work approach, see below and visit www.hims.com/careers-professionals.
About the Role:
We're looking for a Staff Data Engineer to join the Data Platform Engineering team at Hims & Hers as a key technical driver for our most critical platform initiatives. Your scope spans multiple squads: you will drive shared architectural decisions, enhance cross-team reliability, and improve the overall developer experience for a team of nine engineers building the infrastructure that powers patient care for millions of Hims & Hers subscribers.
This is a hands-on execution role. You will own large, complex deliverables end-to-end - from design through production - across our full stack: BigQuery, dbt, Airflow on Astronomer, Confluent Kafka, Databricks, Fivetran, and Terraform/OpenTofu. You will be the DRI (Directly Responsible Individual) for cross-squad initiatives and the engineer other Senior DEs look to for technical direction and growth.
You Will:
Serve as DRI for high-complexity, multi-sprint platform initiatives - Fivetran connector buildouts, Databricks Lakehouse migration workstreams, event streaming infrastructure, lower environment implementation, and engineering standards adoption
Architect, build, and maintain production-grade ingestion pipelines and platform infrastructure - from source connectivity through Bronze/Silver layers - that Analytics Engineering, Data Science, and business teams build on daily
Design, implement, and operate event-driven and streaming data pipelines using Kafka, PySpark, and Databricks Structured Streaming - including defining scaling strategies, cost guardrails, consumer lag alerting, and runbooks before those services reach production
Own the ingestion and raw-to-cleansed layer (Bronze to Silver) data contracts, schema governance, and SLAs
Own data quality for pipelines you build: write dbt tests, wire anomaly detection, validate schemas, and alert on data drift - pipelines ship with quality gates, not after them
Own the reliability of systems you build: establish KPIs and SLOs, implement Datadog monitoring and alerting as code, participate in the on-call rotation, and own Tier 1 operational tickets and runbooks for systems under your domain
Own the integration and data activation layer - Fivetran connectors and Hightouch reverse ETL pipeline connectors - end-to-end from IaC provisioning to production monitoring and schema change governance
Support Analytics Engineers, Data Scientists, and ML engineers by building platform capabilities and data pipelines that unblock their roadmap; partner with legal, security, and DevOps on compliance controls and IaC hardening as needed. DE's responsibility is the platform layer and data delivery; transformation logic and model readiness for serving are owned by Analytics Engineering
Identify and resolve systemic inefficiencies across DPE-owned pipelines and infrastructure - root cause, not just symptom
Mentor Senior Data Engineers through design reviews, code reviews, and pairing; help them grow from squad-level to cross-squad scope
Contribute to and drive adoption of engineering standards - testing practices, CI/CD patterns, observability-as-code, Schema Registry governance - and participate in ARC reviews for changes with cross-team or cost impact
You Have:
8+ years of professional experience designing, building, and operating data pipelines and platform infrastructure
Experience with CDC (Change Data Capture) patterns for real-time ingestion.
Experience with Flink for stream processing
Experience governing and administering dbt in a production BigQuery or Databricks environment - CI/CD configuration, testing standards, documentation standards, and platform-level schema governance. Hands-on dbt experience for ingestion-layer (Bronze/Silver) pipelines
Experience building and operating Airflow DAGs at scale - task-level orchestration patterns, DAG reliability, and multi-priority scheduling
Experience building event streaming pipelines using Kafka or Confluent Kafka - producers, consumers, schema evolution, Schema Registry governance, and consumer lag management
Multi-cloud fluency across GCP and AWS - both are required day-to-day: BigQuery runs on GCP, Airflow runs on AWS EKS
Experience owning data quality for production pipelines - dbt tests, anomaly detection, alerting on schema changes and data drift
Experience with Fivetran or equivalent connector platform - IaC provisioning, schema change handling, and connector health monitoring
Experience with the Databricks platform - Delta Lake, Databricks Workflows, and Unity Catalog
Familiarity with data compliance in a regulated environment - HIPAA/PHI handling, access controls, and audit logging
Infrastructure-as-code experience - Terraform or equivalent; you treat infrastructure changes like code changes
Strong Python and SQL skills; comfortable writing, reviewing, and raising the bar on production-grade pipeline code
Strong design instincts: you take ambiguous requirements, write clear solution designs, and ship to production with minimal rework
Preferred Qualifications:
PySpark/SparkSQL for large-scale data processing
Experience with Hightouch or equivalent reverse ETL platform
Experience with MLOps - supporting ML engineers with data pipelines for model training, feature stores, or experimentation
Familiarity with Looker LookML or equivalent BI serving layer
Go experience for Kafka service development
Experience at a direct-to-consumer healthcare, telehealth, or similarly regulated company
Familiarity with UK/GDPR data compliance requirements distinct from US HIPAA
Our Benefits (there are more but here are some highlights):
Competitive salary & equity compensation for full-time roles
Unlimited PTO, company holidays, and quarterly mental health days
Comprehensive health benefits including medical, dental & vision, and parental leave
Employee Stock Purchase Program (ESPP)
401k benefits with employer matching contribution
Offsite team retreats
We are committed to building a workforce that reflects diverse perspectives and prioritizes ethics, wellness, and a strong sense of belonging. If you're excited about this role, we encourage you to apply—even if you're not sure if your background or experience is a perfect match.
Hims considers all qualified applicants for employment, including applicants with arrest or conviction records, in accordance with the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance, the California Fair Chance Act, and any similar state or local fair chance laws.
It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.
Hims & Hers is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If you need assistance or an accommodation due to a disability, please contact us at [email protected] and describe the needed accommodation. Your privacy is important to us, and any information you share will only be used for the legitimate purpose of considering your request for accommodation. Hims & Hers gives consideration to all qualified applicants without regard to any protected status, including disability. Please do not send resumes to this email address.
To learn more about how we collect, use, retain, and disclose Personal Information, please visit our Global Candidate Privacy Statement.
Redirects to Hims & Hers's application page.
Other roles
More at Hims & Hers.
Principal Data Engineer
Principal
Sr. Software Engineer, Patient Platform (Backend)
Senior
Software Engineer II
Mid
Staff Forward Deployed Engineer
Staff
Principal Engineer (Fullstack/Backend)
Principal