Principal Data Engineer
Data Engineer · Principal · Full Time · Remote
Opens Hims & Hers's application page
Role
What you'll do.
Principal Data Engineer at Hims & Hers - the leading telehealth platform serving millions of patients. As the most senior individual contributor on the Data Platform Engineering team, you'll architect and own the long-term vision for cloud-native data infrastructure spanning GCP BigQuery, Apache Airflow, Kafka/Confluent, and Databricks. This role requires 15+ years of data platform architecture experience, deep expertise in multi-cloud environments, and demonstrated ability to align technical architecture with business goals across an entire engineering organization while mentoring senior engineers and driving critical cost and compliance initiatives in a regulated healthcare environment.
Responsibilities
- Data Platform Architecture Leadership: Own the long-term technical architecture for Data Platform Engineering across ingestion, orchestration, event streaming, and infrastructure that enables transformation and serving. Drive highest-stakes architectural decisions including CDC-based streaming platform design (Kafka to Flink to BigQuery), orchestration platform evaluation, and lower environment strategy built from scratch to support millions of patients across telehealth, prescription, and wellness products.
- Architecture Review and Technical Governance: Chair Architecture Review Committee (ARC) decisions and act as primary technical DRI for cross-team, multi-system, and cost-impacting changes. Establish and enforce engineering standards and production readiness criteria across all DPE-owned systems including testing requirements, CI/CD patterns, observability-as-code, logging standards, data contracts, and Schema Registry governance.
- Data Quality and Observability Architecture: Own comprehensive data quality frameworks and observability architecture including dbt anomaly detection, schema validation, data drift alerting, and platform standards that ensure downstream consumers can trust data integrity. Establish platform standards for production-ready emerging streaming and CDC capabilities.
- Self-Service Analytics Platform Strategy: Define technical strategy for self-service analytics by determining platform capabilities that enable Analytics Engineering independence, implementing guardrails to prevent downstream breakage, and architecting solutions to reduce DPE's bottleneck over time while maintaining governance and data quality standards.
- Data Platform Tooling Evaluation and Governance: Own evaluation, onboarding, and ongoing governance of DPE-managed tooling including Fivetran, Confluent, and equivalent platforms. Manage contract negotiations, cost tracking, deprecation decisions, and ensure tools align with organizational data strategy and scalability requirements.
- Data Sharing, Access Control, and Reverse ETL: Own data sharing and egress patterns including access provisioning, cross-team data contracts, reverse ETL implementations (Hightouch), and governed consumption paths for internal and external consumers. Ensure compliance with HIPAA and GDPR requirements while enabling efficient data activation.
- Cloud Cost Governance and Optimization: Drive cost governance for platform infrastructure spanning BigQuery slot reservations, query optimization, partition strategies, orchestration rightsizing, and cloud spend accountability across the full DPE stack. Own accountability for optimizing GCP and AWS spend while maintaining performance and reliability.
- Incident Response and Operational Excellence: Lead incident response for platform-level P1/P2 incidents as technical escalation point. Facilitate blameless root cause analyses, drive systemic fixes that prevent recurrence, and establish operational procedures that enhance platform reliability and resilience for millions of patients.
- Technical Documentation and Standards: Produce exemplary technical artifacts including architecture decision records, solution design documents, and RFCs that create alignment and establish reference standards for the team. Document patterns and practices that enable knowledge transfer and consistency across the data engineering discipline.
- Senior Engineer Mentorship and Technical Growth: Mentor and elevate Staff and Senior Data Engineers through hands-on design reviews, code reviews, and pairing sessions. Raise the technical ceiling and establish a culture of architectural excellence, design thinking, and continuous learning across the data engineering organization.
- Cross-Functional Partnership and Compliance: Partner cross-functionally with ML/Data Science, legal/security/compliance, and DevOps teams to deliver platform capabilities that are ML-ready, compliant with HIPAA and GDPR regulations, and hardened at the infrastructure layer for a healthcare environment serving millions of patients.
- Hands-On Execution on Critical Path Work: Contribute directly to highest-impact initiatives including net-new streaming platform development, lower environments strategy, and other critical path projects. Balance architectural leadership with hands-on engineering to maintain technical credibility and unblock teams on complex implementation challenges.
Qualifications
What we look for.
Technical
Cloud-Native Data Platform Architecture
15+ years designing, building, and operating data platform architecture at enterprise scale. Primary expertise in GCP (BigQuery, GCS, Dataflow) with deep operational familiarity in AWS (EKS-based Airflow). Multi-cloud fluency required - ability to architect and operate across multiple cloud providers simultaneously, managing hybrid and multi-cloud data infrastructure decisions.
Modern Data Stack Expertise
Hands-on production experience governing dbt at platform scale, Apache Airflow/Astronomer orchestration, Kafka/Confluent event streaming, Databricks/Apache Spark, Fivetran data integration, and data activation platforms like Hightouch. Must demonstrate mastery in each tool's architectural patterns and governance requirements.
Event Streaming and CDC Architecture
Proven ability designing and operating event streaming pipelines at scale including Schema Registry governance, data contract enforcement, and consumer lag management. Experience with Change Data Capture (CDC) patterns and real-time processing frameworks (Flink preferred) for building reliable, low-latency data pipelines.
Data Quality and Observability Frameworks
Experience owning data quality architecture spanning dbt testing frameworks, anomaly detection systems, schema validation, and data observability tooling. Proven track record establishing organization-wide data quality standards that consumers trust and that enable reliable downstream analytics and ML applications.
Data Governance and Healthcare Compliance
Experience working in regulated environments with HIPAA/PHI handling, data classification systems, access control enforcement, audit logging, and GDPR compliance requirements. Demonstrated ability designing governance frameworks that enable secure data usage while maintaining platform agility and self-service capabilities.
Infrastructure-as-Code and DevOps Practices
Fluency with Infrastructure-as-Code tools (Terraform or OpenTofu preferred) treating infrastructure changes with software engineering discipline. Experience managing Kubernetes environments (EKS), CI/CD pipelines, observability infrastructure, and cloud-native deployment patterns at scale.
Python and SQL Production Engineering
Strong production-grade Python and SQL skills with proven ability writing, reviewing, and raising standards for complex data pipeline code. Comfortable with PySpark/SparkSQL for large-scale batch and streaming workloads, and Go or Python for building Kafka producers/consumers.
Incident Response and Operational Management
Experience leading incident response for critical data platform outages with blameless RCA methodology, systemic root cause identification, and organizational learning from failures. Demonstrated ability implementing operational improvements that prevent recurrence and build platform resilience.
Architectural Alignment and Technical Strategy
Proven ability defining and aligning architectural vision with business goals across entire engineering organizations. Demonstrated track record establishing engineering standards and driving adoption across multiple teams without direct authority, translating business requirements into technical roadmaps.
Education
Bachelor's Degree in Computer Science or Related Field
Foundational education in computer science, software engineering, mathematics, physics, or equivalent discipline. While not strictly required given extensive professional experience, formal education provides structured foundation for complex systems thinking.
Experience
Large-Scale Data Platform Ownership
15+ years professional experience designing, building, and owning data platform architecture at company scale. Leadership of data infrastructure serving millions of users with demonstrated responsibility for architectural decisions spanning ingestion, orchestration, storage, and consumption.
Architectural Leadership and Standards Setting
Demonstrated ability defining and aligning architectural vision with business goals across entire engineering organizations. Proven track record establishing engineering standards, production readiness criteria, and governance frameworks adopted across multiple teams without direct authority.
Platform-Scale Event Streaming
Hands-on experience designing and operating event streaming pipelines at scale with Schema Registry governance, data contracts, and consumer lag management. Understanding of CDC patterns and real-time data processing architectures serving mission-critical applications.
Regulated Environment Data Engineering
Experience operating data platforms in regulated healthcare or financial environments with HIPAA, GDPR, and compliance requirements. Understanding of data classification, access controls, audit logging, and security requirements in healthcare contexts.
Data Platform Incident Leadership
Leadership experience managing data platform outages and P1/P2 incidents with blameless RCA execution, systemic root cause analysis, and implementation of preventive measures. Track record of learning from failures and building organizational resilience.
ML and Analytics Platform Enablement
Experience partnering with ML engineers and data scientists to build feature stores, model training pipelines, and experimentation infrastructure. Understanding of data requirements for machine learning at scale and infrastructure needed to support ML operations.
Skills
Required
GCP BigQuery and Cloud Data Warehousing
Deep expertise in Google Cloud Platform's BigQuery including query optimization, slot reservations, partition strategies, and cost governance. Understanding of BigQuery's strengths for OLAP workloads, its integration with GCP ecosystem, and modern data warehouse architecture patterns.
Apache Airflow and Orchestration
Production expertise with Apache Airflow (Astronomer platform preferred) for building, scheduling, and monitoring complex data pipelines. Understanding of DAG design patterns, error handling, retry strategies, and observability for orchestration systems at scale.
Kafka and Event Streaming
Production experience with Apache Kafka and Confluent platform for event streaming, message queuing, and real-time data pipelines. Knowledge of Schema Registry, topic design, consumer groups, partition strategies, and operational management of streaming infrastructure.
dbt and Transformation Frameworks
Hands-on expertise governing dbt (data build tool) at platform scale including dbt testing, documentation, model organization, and establishing standards. Understanding of dbt's role in modern analytics engineering and patterns for maintaining data quality and governance.
Databricks and Apache Spark
Production experience with Databricks platform and Apache Spark for large-scale batch and streaming transformations. Understanding of Delta Lake format, MLflow integration, Unity Catalog for governance, and Databricks architecture for lakehouse implementations.
SQL and Python Production Engineering
Expert-level SQL writing and optimization for complex analytical queries, CTEs, window functions, and performance tuning. Strong Python programming skills for data engineering including libraries (pandas, PySpark), testing practices, and production code quality standards.
Terraform and Infrastructure-as-Code
Proficiency with Terraform or OpenTofu for defining cloud infrastructure as code. Experience managing GCP resources, EKS clusters, IAM policies, and other cloud infrastructure with version control and CI/CD integration treating infrastructure with software engineering discipline.
HIPAA and GDPR Compliance
Deep understanding of HIPAA regulations for healthcare data handling including PHI protection, access controls, audit requirements, and data retention policies. Knowledge of GDPR compliance requirements for European data subjects and ability translating regulations into technical architecture.
Data Quality and Observability
Experience implementing data quality frameworks including anomaly detection, schema validation, data profiling, and data lineage tracking. Knowledge of observability platforms for monitoring data health and establishing SLOs for data reliability.
Technical Documentation and Communication
Exceptional written communication ability producing clear architecture decision records, RFCs, and design documents that build alignment. Comfort articulating complex technical concepts to both engineering and non-technical audiences with clarity and precision.
Preferred
Databricks Unity Catalog and Delta Lake Migration
Nice to haveProduction experience with Databricks Unity Catalog for centralized governance and Delta Lake at scale. Experience leading or executing BigQuery to Databricks lakehouse migrations or similar cloud data warehouse modernization projects managing complex cutover processes.
Change Data Capture (CDC) and Flink
Nice to haveHands-on experience implementing CDC patterns using tools like Debezium and processing with Apache Flink for real-time data pipelines. Understanding of CDC architectures for maintaining consistency between operational systems and analytical platforms.
PySpark and SparkSQL for Large-Scale Processing
Nice to haveAdvanced PySpark and SparkSQL expertise for developing complex batch and streaming transformations processing terabytes of data. Experience optimizing Spark jobs, managing shuffle operations, and tuning Spark configurations for production workloads.
Healthcare and Telehealth Industry Experience
Nice to haveDirect experience working at direct-to-consumer healthcare, telehealth, or healthcare technology companies. Familiarity with healthcare domain challenges, patient data sensitivity, regulatory requirements, and healthcare data infrastructure patterns.
MLOps and Feature Store Architecture
Nice to haveExperience building feature stores, ML training pipelines, and experimentation infrastructure. Understanding of data requirements for machine learning, feature engineering at scale, and platforms like Tecton or Databricks Feature Store.
Kafka Producer and Consumer Development
Nice to haveExperience developing Go or Python services as Kafka producers and consumers for building event-driven data pipelines. Understanding of consumer lag management, exactly-once semantics, and integrating with modern streaming platforms.
Reverse ETL and Data Activation
Nice to haveExperience implementing reverse ETL solutions like Hightouch for activating data across business systems. Understanding of data activation patterns, CDP implementations, and ensuring data consistency in operational systems.
SOX and Audit Compliance
Nice to haveFamiliarity with SOX (Sarbanes-Oxley) compliance controls in data engineering context including audit logging, change management, and internal controls for financial data integrity. Understanding of public company data governance requirements.
Compensation
Pay and benefits.
Base·USD 220,000 – 260,000
Full posting
Original listing.
Hims & Hers is the leading health and wellness platform, on a mission to help the world feel great through the power of better health. We are redefining healthcare by putting the customer first and delivering access to care that is affordable, accessible, and personal, from diagnosis to treatment to delivery. No two people are the same, so we provide access to personalized care designed for results. By normalizing health & wellness challenges and innovating on their solutions, we’re making better health outcomes easier to achieve.
Hims & Hers is a public company, traded on the NYSE under the ticker symbol “HIMS.” To learn more about the brand and offerings, you can visit hims.com/about and hims.com/how-it-works . For information on the company’s outstanding benefits, culture, and its talent-first flexible/remote work approach, see below and visit www.hims.com/careers-professionals.
About the Role:
We're looking for a Principal Data Engineer to be the most senior individual contributor on the Data Platform Engineering (DPE) team at Hims & Hers. In this role, you will define and align the architectural vision for our data platform with business goals – working in close partnership with product, engineering, and data science leadership. Your scope is org-wide: you will set technical direction across every surface DPE owns, drive the highest-stakes architectural decisions, and establish the standards that the entire data engineering discipline operates by.
Our platform serves millions of patients across telehealth, prescription, and wellness products. It runs on GCP BigQuery, Airflow on Astronomer/EKS, dbt, Confluent Kafka, Databricks Delta Lake, and Terraform/OpenTofu – and it is in active, consequential evolution: a net-new streaming platform, and a lower environments strategy being built from scratch. You will own those architectural bets.
You Will:
Own the long-term technical architecture for DPE across ingestion, orchestration, event streaming, and the platform infrastructure that enables transformation and serving - driving the highest-stakes decisions for the CDC-based streaming platform (Kafka → Flink → BigQuery), orchestration platform evaluation, and lower environment strategy
Chair Architecture Review Committee (ARC) decisions; act as the primary technical DRI for cross-team, multi-system, and cost-impacting changes
Establish and enforce engineering standards and production readiness criteria across all DPE-owned systems - testing requirements, CI/CD patterns, observability-as-code, logging standards, data contracts, Schema Registry governance, and what 'production-ready' means for emerging streaming and CDC capabilities
Own data quality and observability architecture - dbt anomaly detection frameworks, schema validation, data drift alerting, and the platform standards that ensure consumers can trust the data they build on
Define the technical strategy for self-service analytics: what platform capabilities enable Analytics Engineering to work independently, what guardrails prevent downstream breakage, and how DPE reduces its bottleneck over time
Own evaluation, onboarding, and ongoing governance of DPE-managed tooling - Fivetran, Confluent, and equivalent platforms, including contract management, cost tracking, and deprecation decisions
Own data sharing and egress patterns - access provisioning, cross-team data contracts, reverse ETL (Hightouch), and governed consumption paths for internal and external consumers
Drive cost governance for platform infrastructure - BigQuery slot reservations, query optimization, partition strategies, orchestration rightsizing, and cloud spend accountability across the full DPE stack
Lead incident response for platform-level P1/P2 incidents: act as technical escalation point, facilitate blameless RCAs, and drive systemic fixes that prevent recurrence
Produce exemplary technical artifacts - architecture decision records, solution design docs, RFCs - that create alignment and become the team's reference standard
Mentor and elevate Staff and Senior Data Engineers; raise the technical ceiling through design reviews, code reviews, and hands-on pairing
Partner cross-functionally with ML/Data Science, legal/security/compliance, and DevOps to deliver platform capabilities that are ML-ready, compliant with HIPAA/GDPR, and hardened at the infrastructure layer
Contribute hands-on to critical path work
You Have:
15+ years of professional experience designing, building, and owning data platform architecture at company scale
Demonstrated ability to define and align architectural vision with business goals across an entire engineering organization
Deep expertise in cloud-native data platforms across GCP primary (BigQuery, GCS, Dataflow); AWS operational familiarity required (EKS-based Airflow) - BigQuery strongly preferred; multi-cloud fluency is required, not a plus
Hands-on experience with the modern data stack: experience governing dbt at platform scale, Airflow/Astronomer, Kafka/Confluent, Databricks/Spark, Fivetran, and data sharing/activation platforms (Hightouch or equivalent)
Experience designing and operating event streaming pipelines at scale - including Schema Registry, data contracts, and consumer lag management
Proven track record establishing engineering standards across multiple teams and driving adoption without direct authority
Experience owning data quality frameworks - dbt testing, anomaly detection, schema validation, and data observability tooling
Experience with data governance and compliance frameworks in a regulated environment - HIPAA/PHI handling, data classification, access controls, audit logging, and GDPR
Experience leading incident response for data platform outages - blameless RCA, systemic root cause identification, and operational improvement
Infrastructure-as-code fluency - Terraform or equivalent; you treat infrastructure changes like software changes
Strong Python and SQL skills; comfortable writing, reviewing, and raising the bar on production-grade pipeline code
Clear written communication: you produce design docs and RFCs that build alignment, not confusion. Comfort operating in ambiguity - you define the path, you don't wait for it to be defined
Preferred Qualifications:
Databricks, Unity Catalog, and Delta Lake in production at scale
Experience with CDC (Change Data Capture) patterns and Flink for real-time data processing
PySpark/SparkSQL for large-scale batch and streaming workloads
Experience driving a BigQuery → Databricks Lakehouse migration or equivalent cloud data warehouse migration
Experience with MLOps - partnering with ML engineers on model training pipelines, feature stores, and experimentation infrastructure
Go or Python service development for Kafka producers and consumers
Experience at a direct-to-consumer healthcare or telehealth company with HIPAA and GDPR obligations
Familiarity with SOX compliance controls in a data engineering context
Our Benefits (there are more but here are some highlights):
Competitive salary & equity compensation for full-time roles
Unlimited PTO, company holidays, and quarterly mental health days
Comprehensive health benefits including medical, dental & vision, and parental leave
Employee Stock Purchase Program (ESPP)
401k benefits with employer matching contribution
Offsite team retreats
We are committed to building a workforce that reflects diverse perspectives and prioritizes ethics, wellness, and a strong sense of belonging. If you're excited about this role, we encourage you to apply—even if you're not sure if your background or experience is a perfect match.
Hims considers all qualified applicants for employment, including applicants with arrest or conviction records, in accordance with the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance, the California Fair Chance Act, and any similar state or local fair chance laws.
It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.
Hims & Hers is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If you need assistance or an accommodation due to a disability, please contact us at [email protected] and describe the needed accommodation. Your privacy is important to us, and any information you share will only be used for the legitimate purpose of considering your request for accommodation. Hims & Hers gives consideration to all qualified applicants without regard to any protected status, including disability. Please do not send resumes to this email address.
To learn more about how we collect, use, retain, and disclose Personal Information, please visit our Global Candidate Privacy Statement.
Redirects to Hims & Hers's application page.
Other roles
More at Hims & Hers.
Staff Data Engineer
Staff
Sr. Software Engineer, Patient Platform (Backend)
Senior
Software Engineer II
Mid
Staff Forward Deployed Engineer
Staff
Principal Engineer (Fullstack/Backend)
Principal