Principal Data Engineer

Data Engineer · Principal · Full Time · Remote

US Remote · RemoteUSD 220k – 260k2d ago
Apply for this role

Opens Hims & Hers's application page

Role

What you'll do.

Principal Data Engineer at Hims & Hers - the leading telehealth platform serving millions of patients. As the most senior individual contributor on the Data Platform Engineering team, you'll architect and own the long-term vision for cloud-native data infrastructure spanning GCP BigQuery, Apache Airflow, Kafka/Confluent, and Databricks. This role requires 15+ years of data platform architecture experience, deep expertise in multi-cloud environments, and demonstrated ability to align technical architecture with business goals across an entire engineering organization while mentoring senior engineers and driving critical cost and compliance initiatives in a regulated healthcare environment.

Responsibilities

  • Data Platform Architecture Leadership: Own the long-term technical architecture for Data Platform Engineering across ingestion, orchestration, event streaming, and infrastructure that enables transformation and serving. Drive highest-stakes architectural decisions including CDC-based streaming platform design (Kafka to Flink to BigQuery), orchestration platform evaluation, and lower environment strategy built from scratch to support millions of patients across telehealth, prescription, and wellness products.
  • Architecture Review and Technical Governance: Chair Architecture Review Committee (ARC) decisions and act as primary technical DRI for cross-team, multi-system, and cost-impacting changes. Establish and enforce engineering standards and production readiness criteria across all DPE-owned systems including testing requirements, CI/CD patterns, observability-as-code, logging standards, data contracts, and Schema Registry governance.
  • Data Quality and Observability Architecture: Own comprehensive data quality frameworks and observability architecture including dbt anomaly detection, schema validation, data drift alerting, and platform standards that ensure downstream consumers can trust data integrity. Establish platform standards for production-ready emerging streaming and CDC capabilities.
  • Self-Service Analytics Platform Strategy: Define technical strategy for self-service analytics by determining platform capabilities that enable Analytics Engineering independence, implementing guardrails to prevent downstream breakage, and architecting solutions to reduce DPE's bottleneck over time while maintaining governance and data quality standards.
  • Data Platform Tooling Evaluation and Governance: Own evaluation, onboarding, and ongoing governance of DPE-managed tooling including Fivetran, Confluent, and equivalent platforms. Manage contract negotiations, cost tracking, deprecation decisions, and ensure tools align with organizational data strategy and scalability requirements.
  • Data Sharing, Access Control, and Reverse ETL: Own data sharing and egress patterns including access provisioning, cross-team data contracts, reverse ETL implementations (Hightouch), and governed consumption paths for internal and external consumers. Ensure compliance with HIPAA and GDPR requirements while enabling efficient data activation.
  • Cloud Cost Governance and Optimization: Drive cost governance for platform infrastructure spanning BigQuery slot reservations, query optimization, partition strategies, orchestration rightsizing, and cloud spend accountability across the full DPE stack. Own accountability for optimizing GCP and AWS spend while maintaining performance and reliability.
  • Incident Response and Operational Excellence: Lead incident response for platform-level P1/P2 incidents as technical escalation point. Facilitate blameless root cause analyses, drive systemic fixes that prevent recurrence, and establish operational procedures that enhance platform reliability and resilience for millions of patients.
  • Technical Documentation and Standards: Produce exemplary technical artifacts including architecture decision records, solution design documents, and RFCs that create alignment and establish reference standards for the team. Document patterns and practices that enable knowledge transfer and consistency across the data engineering discipline.
  • Senior Engineer Mentorship and Technical Growth: Mentor and elevate Staff and Senior Data Engineers through hands-on design reviews, code reviews, and pairing sessions. Raise the technical ceiling and establish a culture of architectural excellence, design thinking, and continuous learning across the data engineering organization.
  • Cross-Functional Partnership and Compliance: Partner cross-functionally with ML/Data Science, legal/security/compliance, and DevOps teams to deliver platform capabilities that are ML-ready, compliant with HIPAA and GDPR regulations, and hardened at the infrastructure layer for a healthcare environment serving millions of patients.
  • Hands-On Execution on Critical Path Work: Contribute directly to highest-impact initiatives including net-new streaming platform development, lower environments strategy, and other critical path projects. Balance architectural leadership with hands-on engineering to maintain technical credibility and unblock teams on complex implementation challenges.

Qualifications

What we look for.

Technical

  • Cloud-Native Data Platform Architecture

    15+ years designing, building, and operating data platform architecture at enterprise scale. Primary expertise in GCP (BigQuery, GCS, Dataflow) with deep operational familiarity in AWS (EKS-based Airflow). Multi-cloud fluency required - ability to architect and operate across multiple cloud providers simultaneously, managing hybrid and multi-cloud data infrastructure decisions.

  • Modern Data Stack Expertise

    Hands-on production experience governing dbt at platform scale, Apache Airflow/Astronomer orchestration, Kafka/Confluent event streaming, Databricks/Apache Spark, Fivetran data integration, and data activation platforms like Hightouch. Must demonstrate mastery in each tool's architectural patterns and governance requirements.

  • Event Streaming and CDC Architecture

    Proven ability designing and operating event streaming pipelines at scale including Schema Registry governance, data contract enforcement, and consumer lag management. Experience with Change Data Capture (CDC) patterns and real-time processing frameworks (Flink preferred) for building reliable, low-latency data pipelines.

  • Data Quality and Observability Frameworks

    Experience owning data quality architecture spanning dbt testing frameworks, anomaly detection systems, schema validation, and data observability tooling. Proven track record establishing organization-wide data quality standards that consumers trust and that enable reliable downstream analytics and ML applications.

  • Data Governance and Healthcare Compliance

    Experience working in regulated environments with HIPAA/PHI handling, data classification systems, access control enforcement, audit logging, and GDPR compliance requirements. Demonstrated ability designing governance frameworks that enable secure data usage while maintaining platform agility and self-service capabilities.

  • Infrastructure-as-Code and DevOps Practices

    Fluency with Infrastructure-as-Code tools (Terraform or OpenTofu preferred) treating infrastructure changes with software engineering discipline. Experience managing Kubernetes environments (EKS), CI/CD pipelines, observability infrastructure, and cloud-native deployment patterns at scale.

  • Python and SQL Production Engineering

    Strong production-grade Python and SQL skills with proven ability writing, reviewing, and raising standards for complex data pipeline code. Comfortable with PySpark/SparkSQL for large-scale batch and streaming workloads, and Go or Python for building Kafka producers/consumers.

  • Incident Response and Operational Management

    Experience leading incident response for critical data platform outages with blameless RCA methodology, systemic root cause identification, and organizational learning from failures. Demonstrated ability implementing operational improvements that prevent recurrence and build platform resilience.

  • Architectural Alignment and Technical Strategy

    Proven ability defining and aligning architectural vision with business goals across entire engineering organizations. Demonstrated track record establishing engineering standards and driving adoption across multiple teams without direct authority, translating business requirements into technical roadmaps.

Education

  • Bachelor's Degree in Computer Science or Related Field

    Foundational education in computer science, software engineering, mathematics, physics, or equivalent discipline. While not strictly required given extensive professional experience, formal education provides structured foundation for complex systems thinking.

Experience

  • Large-Scale Data Platform Ownership

    15+ years professional experience designing, building, and owning data platform architecture at company scale. Leadership of data infrastructure serving millions of users with demonstrated responsibility for architectural decisions spanning ingestion, orchestration, storage, and consumption.

  • Architectural Leadership and Standards Setting

    Demonstrated ability defining and aligning architectural vision with business goals across entire engineering organizations. Proven track record establishing engineering standards, production readiness criteria, and governance frameworks adopted across multiple teams without direct authority.

  • Platform-Scale Event Streaming

    Hands-on experience designing and operating event streaming pipelines at scale with Schema Registry governance, data contracts, and consumer lag management. Understanding of CDC patterns and real-time data processing architectures serving mission-critical applications.

  • Regulated Environment Data Engineering

    Experience operating data platforms in regulated healthcare or financial environments with HIPAA, GDPR, and compliance requirements. Understanding of data classification, access controls, audit logging, and security requirements in healthcare contexts.

  • Data Platform Incident Leadership

    Leadership experience managing data platform outages and P1/P2 incidents with blameless RCA execution, systemic root cause analysis, and implementation of preventive measures. Track record of learning from failures and building organizational resilience.

  • ML and Analytics Platform Enablement

    Experience partnering with ML engineers and data scientists to build feature stores, model training pipelines, and experimentation infrastructure. Understanding of data requirements for machine learning at scale and infrastructure needed to support ML operations.

Skills

Required

  • GCP BigQuery and Cloud Data Warehousing

    Deep expertise in Google Cloud Platform's BigQuery including query optimization, slot reservations, partition strategies, and cost governance. Understanding of BigQuery's strengths for OLAP workloads, its integration with GCP ecosystem, and modern data warehouse architecture patterns.

  • Apache Airflow and Orchestration

    Production expertise with Apache Airflow (Astronomer platform preferred) for building, scheduling, and monitoring complex data pipelines. Understanding of DAG design patterns, error handling, retry strategies, and observability for orchestration systems at scale.

  • Kafka and Event Streaming

    Production experience with Apache Kafka and Confluent platform for event streaming, message queuing, and real-time data pipelines. Knowledge of Schema Registry, topic design, consumer groups, partition strategies, and operational management of streaming infrastructure.

  • dbt and Transformation Frameworks

    Hands-on expertise governing dbt (data build tool) at platform scale including dbt testing, documentation, model organization, and establishing standards. Understanding of dbt's role in modern analytics engineering and patterns for maintaining data quality and governance.

  • Databricks and Apache Spark

    Production experience with Databricks platform and Apache Spark for large-scale batch and streaming transformations. Understanding of Delta Lake format, MLflow integration, Unity Catalog for governance, and Databricks architecture for lakehouse implementations.

  • SQL and Python Production Engineering

    Expert-level SQL writing and optimization for complex analytical queries, CTEs, window functions, and performance tuning. Strong Python programming skills for data engineering including libraries (pandas, PySpark), testing practices, and production code quality standards.

  • Terraform and Infrastructure-as-Code

    Proficiency with Terraform or OpenTofu for defining cloud infrastructure as code. Experience managing GCP resources, EKS clusters, IAM policies, and other cloud infrastructure with version control and CI/CD integration treating infrastructure with software engineering discipline.

  • HIPAA and GDPR Compliance

    Deep understanding of HIPAA regulations for healthcare data handling including PHI protection, access controls, audit requirements, and data retention policies. Knowledge of GDPR compliance requirements for European data subjects and ability translating regulations into technical architecture.

  • Data Quality and Observability

    Experience implementing data quality frameworks including anomaly detection, schema validation, data profiling, and data lineage tracking. Knowledge of observability platforms for monitoring data health and establishing SLOs for data reliability.

  • Technical Documentation and Communication

    Exceptional written communication ability producing clear architecture decision records, RFCs, and design documents that build alignment. Comfort articulating complex technical concepts to both engineering and non-technical audiences with clarity and precision.

Preferred

  • Databricks Unity Catalog and Delta Lake Migration

    Nice to have

    Production experience with Databricks Unity Catalog for centralized governance and Delta Lake at scale. Experience leading or executing BigQuery to Databricks lakehouse migrations or similar cloud data warehouse modernization projects managing complex cutover processes.

  • Change Data Capture (CDC) and Flink

    Nice to have

    Hands-on experience implementing CDC patterns using tools like Debezium and processing with Apache Flink for real-time data pipelines. Understanding of CDC architectures for maintaining consistency between operational systems and analytical platforms.

  • PySpark and SparkSQL for Large-Scale Processing

    Nice to have

    Advanced PySpark and SparkSQL expertise for developing complex batch and streaming transformations processing terabytes of data. Experience optimizing Spark jobs, managing shuffle operations, and tuning Spark configurations for production workloads.

  • Healthcare and Telehealth Industry Experience

    Nice to have

    Direct experience working at direct-to-consumer healthcare, telehealth, or healthcare technology companies. Familiarity with healthcare domain challenges, patient data sensitivity, regulatory requirements, and healthcare data infrastructure patterns.

  • MLOps and Feature Store Architecture

    Nice to have

    Experience building feature stores, ML training pipelines, and experimentation infrastructure. Understanding of data requirements for machine learning, feature engineering at scale, and platforms like Tecton or Databricks Feature Store.

  • Kafka Producer and Consumer Development

    Nice to have

    Experience developing Go or Python services as Kafka producers and consumers for building event-driven data pipelines. Understanding of consumer lag management, exactly-once semantics, and integrating with modern streaming platforms.

  • Reverse ETL and Data Activation

    Nice to have

    Experience implementing reverse ETL solutions like Hightouch for activating data across business systems. Understanding of data activation patterns, CDP implementations, and ensuring data consistency in operational systems.

  • SOX and Audit Compliance

    Nice to have

    Familiarity with SOX (Sarbanes-Oxley) compliance controls in data engineering context including audit logging, change management, and internal controls for financial data integrity. Understanding of public company data governance requirements.

Compensation

Pay and benefits.

Base·USD 220,000 – 260,000

Full posting

Original listing.

Hims & Hers is the leading health and wellness platform, on a mission to help the world feel great through the power of better health. We are redefining healthcare by putting the customer first and delivering access to care that is affordable, accessible, and personal, from diagnosis to treatment to delivery. No two people are the same, so we provide access to personalized care designed for results. By normalizing health & wellness challenges and innovating on their solutions, we’re making better health outcomes easier to achieve. 

Hims & Hers is a public company, traded on the NYSE under the ticker symbol “HIMS.” To learn more about the brand and offerings, you can visit hims.com/about and hims.com/how-it-works . For information on the company’s outstanding benefits, culture, and its talent-first flexible/remote work approach, see below and visit www.hims.com/careers-professionals.

About the Role:

We're looking for a Principal Data Engineer to be the most senior individual contributor on the Data Platform Engineering (DPE) team at Hims & Hers. In this role, you will define and align the architectural vision for our data platform with business goals – working in close partnership with product, engineering, and data science leadership. Your scope is org-wide: you will set technical direction across every surface DPE owns, drive the highest-stakes architectural decisions, and establish the standards that the entire data engineering discipline operates by.


Our platform serves millions of patients across telehealth, prescription, and wellness products. It runs on GCP BigQuery, Airflow on Astronomer/EKS, dbt, Confluent Kafka, Databricks Delta Lake, and Terraform/OpenTofu – and it is in active, consequential evolution: a net-new streaming platform, and a lower environments strategy being built from scratch. You will own those architectural bets.

You Will:

  • Own the long-term technical architecture for DPE across ingestion, orchestration, event streaming, and the platform infrastructure that enables transformation and serving - driving the highest-stakes decisions for the CDC-based streaming platform (Kafka → Flink → BigQuery), orchestration platform evaluation, and lower environment strategy

  • Chair Architecture Review Committee (ARC) decisions; act as the primary technical DRI for cross-team, multi-system, and cost-impacting changes

  • Establish and enforce engineering standards and production readiness criteria across all DPE-owned systems - testing requirements, CI/CD patterns, observability-as-code, logging standards, data contracts, Schema Registry governance, and what 'production-ready' means for emerging streaming and CDC capabilities

  • Own data quality and observability architecture - dbt anomaly detection frameworks, schema validation, data drift alerting, and the platform standards that ensure consumers can trust the data they build on

  • Define the technical strategy for self-service analytics: what platform capabilities enable Analytics Engineering to work independently, what guardrails prevent downstream breakage, and how DPE reduces its bottleneck over time

  • Own evaluation, onboarding, and ongoing governance of DPE-managed tooling - Fivetran, Confluent, and equivalent platforms, including contract management, cost tracking, and deprecation decisions

  • Own data sharing and egress patterns - access provisioning, cross-team data contracts, reverse ETL (Hightouch), and governed consumption paths for internal and external consumers

  • Drive cost governance for platform infrastructure - BigQuery slot reservations, query optimization, partition strategies, orchestration rightsizing, and cloud spend accountability across the full DPE stack

  • Lead incident response for platform-level P1/P2 incidents: act as technical escalation point, facilitate blameless RCAs, and drive systemic fixes that prevent recurrence

  • Produce exemplary technical artifacts - architecture decision records, solution design docs, RFCs - that create alignment and become the team's reference standard

  • Mentor and elevate Staff and Senior Data Engineers; raise the technical ceiling through design reviews, code reviews, and hands-on pairing

  • Partner cross-functionally with ML/Data Science, legal/security/compliance, and DevOps to deliver platform capabilities that are ML-ready, compliant with HIPAA/GDPR, and hardened at the infrastructure layer

  • Contribute hands-on to critical path work

You Have:

  • 15+ years of professional experience designing, building, and owning data platform architecture at company scale

  • Demonstrated ability to define and align architectural vision with business goals across an entire engineering organization

  • Deep expertise in cloud-native data platforms across GCP primary (BigQuery, GCS, Dataflow); AWS operational familiarity required (EKS-based Airflow) - BigQuery strongly preferred; multi-cloud fluency is required, not a plus

  • Hands-on experience with the modern data stack: experience governing dbt at platform scale, Airflow/Astronomer, Kafka/Confluent, Databricks/Spark, Fivetran, and data sharing/activation platforms (Hightouch or equivalent)

  • Experience designing and operating event streaming pipelines at scale - including Schema Registry, data contracts, and consumer lag management

  • Proven track record establishing engineering standards across multiple teams and driving adoption without direct authority

  • Experience owning data quality frameworks - dbt testing, anomaly detection, schema validation, and data observability tooling

  • Experience with data governance and compliance frameworks in a regulated environment - HIPAA/PHI handling, data classification, access controls, audit logging, and GDPR

  • Experience leading incident response for data platform outages - blameless RCA, systemic root cause identification, and operational improvement

  • Infrastructure-as-code fluency - Terraform or equivalent; you treat infrastructure changes like software changes

  • Strong Python and SQL skills; comfortable writing, reviewing, and raising the bar on production-grade pipeline code

  • Clear written communication: you produce design docs and RFCs that build alignment, not confusion. Comfort operating in ambiguity - you define the path, you don't wait for it to be defined

Preferred Qualifications:

  • Databricks, Unity Catalog, and Delta Lake in production at scale

  • Experience with CDC (Change Data Capture) patterns and Flink for real-time data processing

  • PySpark/SparkSQL for large-scale batch and streaming workloads

  • Experience driving a BigQuery → Databricks Lakehouse migration or equivalent cloud data warehouse migration

  • Experience with MLOps - partnering with ML engineers on model training pipelines, feature stores, and experimentation infrastructure

  • Go or Python service development for Kafka producers and consumers

  • Experience at a direct-to-consumer healthcare or telehealth company with HIPAA and GDPR obligations

  • Familiarity with SOX compliance controls in a data engineering context

Our Benefits (there are more but here are some highlights):

  • Competitive salary & equity compensation for full-time roles

  • Unlimited PTO, company holidays, and quarterly mental health days

  • Comprehensive health benefits including medical, dental & vision, and parental leave

  • Employee Stock Purchase Program (ESPP)

  • 401k benefits with employer matching contribution

  • Offsite team retreats

We are committed to building a workforce that reflects diverse perspectives and prioritizes ethics, wellness, and a strong sense of belonging. If you're excited about this role, we encourage you to apply—even if you're not sure if your background or experience is a perfect match.

Hims considers all qualified applicants for employment, including applicants with arrest or conviction records, in accordance with the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance, the California Fair Chance Act, and any similar state or local fair chance laws.

It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.

Hims & Hers is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If you need assistance or an accommodation due to a disability, please contact us at [email protected] and describe the needed accommodation. Your privacy is important to us, and any information you share will only be used for the legitimate purpose of considering your request for accommodation. Hims & Hers gives consideration to all qualified applicants without regard to any protected status, including disability. Please do not send resumes to this email address.

To learn more about how we collect, use, retain, and disclose Personal Information, please visit our Global Candidate Privacy Statement.

Redirects to Hims & Hers's application page.

Other roles

More at Hims & Hers.

View all 13 roles