Sr. Staff Machine Learning Systems Engineer

ML Engineer · Staff · Full Time · Remote

US Remote · RemoteUSD 240k – 265k3d ago
Apply for this role

Opens Hims & Hers's application page

Role

What you'll do.

Sr. Staff Machine Learning Systems Engineer at Hims & Hers, a NYSE-traded digital health platform, seeking a leadership-level engineer to own AI/ML evaluation systems and data pipelines in a regulated healthcare environment. This role requires 10+ years of ML infrastructure and data engineering expertise with hands-on depth in evaluation methodology, statistical testing, and large-scale data systems, plus demonstrated ability to lead cross-functional initiatives and mentor senior engineers while building trusted, production-ready AI systems for patient care.

Responsibilities

  • Own Evaluation System Architecture: Set technical direction for comprehensive evaluation systems including metric design, LLM judge and scorer calibration, statistical regression methodology, and infrastructure for tracking and categorizing model failures over time in production healthcare environments.
  • Design and Scale Data Pipelines: Build and maintain robust data pipelines handling ingestion, transformation, dataset versioning, labeling workflows, and calibration systems that support both AI evaluation and downstream data science operations serving millions of patients.
  • Lead Multi-Team Initiatives: Drive cross-functional projects spanning engineering, product, and clinical teams to establish evaluation frameworks for new AI services, replace manual review processes with statistically sound automated gates, and implement adversarial red-team testing programs that reduce safety risks before production deployment.
  • Establish Statistical Testing Standards: Design and defend production-grade statistical testing frameworks including paired significance testing, multiple comparison corrections, and human label agreement metrics that inform critical decisions about AI model reliability and patient safety.
  • Develop Reusable Evaluation Methodology: Transform ambiguous, cross-team challenges into standardized approaches and platforms that other teams leverage, taking vague requirements from conception through fully-specified shipped systems without handoffs.
  • Lead Platform Modernization: Execute major system re-architecture initiatives removing brittle logic and modernizing core evaluation infrastructure with org-wide impact, ensuring scalability and maintainability for growing AI/ML capabilities.
  • Build Cross-Functional Relationships: Establish trusted relationships across ML engineering, data science, platform engineering, clinical affairs, legal, and product teams to ensure early involvement in decisions and facilitate knowledge sharing about evaluation best practices.
  • Mentor Senior Engineers and Raise Technical Bar: Mentor experienced engineers and ML specialists, share evaluation methodology expertise through internal talks and technical write-ups, and elevate the statistical rigor and systems thinking capabilities of teams tackling related problems.

Qualifications

What we look for.

Technical

  • ML Infrastructure and Evaluation Systems Expertise

    Deep hands-on experience designing, building, and scaling ML evaluation platforms with proven ability to implement LLM judge calibration, regression testing methodology, and failure tracking systems that inform production deployment decisions.

  • Statistical Testing and Validation Mastery

    Expert-level knowledge of statistical testing frameworks including paired significance testing, multiple comparison corrections, human label agreement metrics, and the ability to design methodologies that real production decisions depend on.

  • Data Pipeline Architecture

    Advanced proficiency building high-throughput data engineering systems including dataset versioning, feature engineering pipelines, benchmark datasets, labeling and calibration workflows, and transformation systems that handle scale and complexity.

  • Adversarial and Red-Team Evaluation

    Demonstrated experience designing comprehensive adversarial testing suites, failure taxonomies, and red-team evaluation approaches that proactively identify safety issues and edge cases before production deployment in sensitive environments.

  • Python Development

    Strong hands-on Python programming skills for building evaluation infrastructure, data pipelines, and analysis tools at production scale.

Education

  • Bachelor's Degree in Computer Science, Engineering, or Related Field

    Foundational education in computer science, software engineering, mathematics, statistics, or equivalent field providing strong algorithmic and systems thinking background.

  • Advanced Statistical and Mathematical Proficiency

    Deep understanding of statistical testing, experimental design, probability theory, and computational mathematics sufficient to design and defend production evaluation frameworks.

Experience

  • 10+ Years ML Infrastructure and Data Engineering

    Extensive background in ML infrastructure, data engineering, or evaluation/testing systems with track record of impact reaching beyond individual projects, demonstrating ability to influence how organizations approach ML systems problems.

  • Building Organizational Standards

    Proven history of creating approaches and methodologies that became standard practices for broader teams, not just solving individual problems but changing how groups tackle entire categories of problems.

  • Multi-Team Project Leadership

    Demonstrated experience leading complex, multi-team projects to completion over extended timelines, including navigating and resolving genuine technical disagreements to drive alignment across stakeholders.

  • Mentoring and Team Development

    Track record of mentoring experienced engineers and raising technical standards, with visible impact on team capabilities and decision-making quality across organizations.

  • Regulated Industry Experience

    Prior experience working in regulated environments such as healthcare, fintech, or life sciences, understanding compliance requirements, audit standards, and risk management in high-stakes domains.

Skills

Required

  • Python

    Production-grade Python development for building ML evaluation systems, data pipelines, and statistical analysis tools.

  • Statistical Testing Methodology

    Expertise in paired significance testing, multiple comparison corrections, statistical power analysis, and hypothesis testing frameworks.

  • Data Pipeline Engineering

    Design and implementation of scalable ETL systems, dataset versioning, feature engineering workflows, and data validation frameworks.

  • ML Evaluation Systems

    Designing metrics, judges, scorers, and evaluation infrastructure that measures model reliability and guides deployment decisions.

  • Distributed Systems Design

    Understanding of scalable system architecture, data processing frameworks, and infrastructure patterns for high-throughput data operations.

  • Cross-Functional Leadership

    Ability to navigate complex stakeholder environments, drive consensus among technical leaders, and communicate technical concepts to non-technical audiences.

  • SQL

    Advanced SQL for data analysis, pipeline design, and querying large-scale datasets.

Preferred

  • Databricks Platform Experience

    Nice to have

    Hands-on experience with Databricks, Unity Catalog, and lakehouse architecture for ML and data workflows at scale.

  • MLflow and Model Management

    Nice to have

    Experience with MLflow, model versioning systems, and experiment tracking platforms for production ML operations.

  • LLM and Foundation Model Evaluation

    Nice to have

    Specific experience designing evaluation frameworks and judges for large language models and modern foundation models in production systems.

  • Reporting and Analytics Tools

    Nice to have

    Building automated reporting systems, dashboards, and communication tools for non-technical stakeholders (Slack bots, spreadsheet automation, BI tools).

  • Healthcare or Regulated Industry Experience

    Nice to have

    Prior work in healthcare, fintech, or other regulated industries understanding compliance, audit requirements, and risk management processes.

  • Public Speaking and Technical Communication

    Nice to have

    Track record of company-wide talks, technical write-ups, or published work demonstrating ability to influence how broader organizations approach problems.

  • Kubernetes and Container Orchestration

    Nice to have

    Experience deploying and managing ML infrastructure using Kubernetes, Docker, or container-based systems at scale.

Tech stack

Languages

PythonSQL

Frameworks

MLflowApache SparkPandas

Databases

DatabricksPostgres or Similar Relational Databases

Tools

Jupyter NotebooksGit and Version ControlCI/CD SystemsMonitoring and Observability Tools

Other

Statistical Analysis and Hypothesis TestingData Versioning and LineageAutomated Reporting and BI

Compensation

Pay and benefits.

Base·USD 240,000 – 265,000

Equity·Stock options

Benefits

  • Competitive Salary and Equity Compensation

    Market-competitive base salary with meaningful equity stake as a public company employee, providing long-term financial alignment with Hims & Hers mission and performance.

  • Unlimited Paid Time Off

    Flexible vacation policy with unlimited PTO, company holidays, and quarterly mental health days supporting work-life balance and employee wellness.

  • Comprehensive Health Benefits

    Full medical, dental, and vision coverage with extensive parental leave policies supporting employees and their families.

  • Employee Stock Purchase Program (ESPP)

    Opportunity to purchase Hims & Hers stock at a discount, building personal wealth while strengthening investment in company success.

  • 401(k) with Employer Matching

    Retirement savings plan with employer matching contributions supporting long-term financial security and planning.

  • Team Offsite Retreats

    Regular team gatherings and company-wide offsites building relationships, fostering collaboration, and celebrating achievements across the organization.

  • Remote and Flexible Work Culture

    Talent-first flexible and remote work approach enabling engineers to work from anywhere, supporting the company's commitment to employee autonomy and work-life integration.

Process

Interview steps.

  1. 01

    Initial Screening Call

    Recruiter conversation exploring background in ML infrastructure, data engineering, and leadership experience to assess cultural fit and baseline qualifications for the Sr. Staff role.

  2. 02

    Technical Deep Dive Interview

    Detailed technical discussion with ML engineering leadership covering evaluation system design, statistical methodology, data pipeline architecture, and specific experience with LLM judges, regression testing, and production ML systems.

  3. 03

    Systems Design Round

    Collaborative problem-solving session designing an evaluation framework or data pipeline architecture for a healthcare AI use case, assessing cross-functional thinking and ability to navigate ambiguity.

  4. 04

    Leadership and Impact Discussion

    Conversation with senior leadership exploring track record of building organizational standards, mentoring engineers, leading multi-team initiatives, and influence on how organizations approach ML systems problems.

  5. 05

    Stakeholder Alignment Discussion

    Meeting with product, clinical, and data science leaders to discuss cross-functional collaboration style, ability to build trust across disciplines, and understanding of regulated healthcare requirements.

  6. 06

    Final Round with Hiring Manager

    Comprehensive conversation with the Sr. Staff role's manager covering role expectations, team dynamics, growth opportunities, and alignment on technical vision for evaluation systems at Hims & Hers.

Full posting

Original listing.

Hims & Hers is the leading health and wellness platform, on a mission to help the world feel great through the power of better health. We are redefining healthcare by putting the customer first and delivering access to care that is affordable, accessible, and personal, from diagnosis to treatment to delivery. No two people are the same, so we provide access to personalized care designed for results. By normalizing health & wellness challenges and innovating on their solutions, we’re making better health outcomes easier to achieve. 

Hims & Hers is a public company, traded on the NYSE under the ticker symbol “HIMS.” To learn more about the brand and offerings, you can visit hims.com/about and hims.com/how-it-works . For information on the company’s outstanding benefits, culture, and its talent-first flexible/remote work approach, see below and visit www.hims.com/careers-professionals.

About the Role:

How do we make advanced AI/ML not just powerful but trustworthy enough to run in a regulated healthcare environment? We're looking for a Senior Staff engineer who can own that question end to end: the data pipelines that feed our models and evaluations, and the evaluation infrastructure — judges, scorers, statistical regression gates, red-team testing — that decides whether AI models are effective and safe to ship to patients.

This is a leadership role for someone who works comfortably across disciplines. You'll set technical direction for how we build, version, and trust the data and judgments that our AI products are evaluated against. Your scope will expand from raw data ingestion and feature/dataset pipelines, through evaluation methodology and statistical rigor, to the reporting surfaces that let clinical and product teams act on what we learn.

You'll spend most of your time on problems that don't have an existing playbook: ambiguous, cross-team, and genuinely hard to reason about. The job is to bring clarity to that ambiguity, chart a path the rest of the team and organization can follow, and see it through from idea to production, building the relationships and buy-in along the way to make it stick.

 

You Will:

Own the evaluation as a whole, not just a slice of it

  • Set the technical direction for our evaluation systems; metric, judge and scorer design, the statistical methodology behind regression decisions, and the infrastructure that tracks and categorizes failures over time.

  • Design and scale the data pipelines — ingestion, transformation, dataset versioning, labeling and calibration workflows — that both evaluation and downstream data science work depend on.

  • Proactively address challenges in scaling and complexity AI evaluation.

Lead projects that span teams and quarters

  • Define how we evaluate any new AI service from scratch. Drive multi-team initiatives like replacing manual, inconsistent review processes with statistically sound, automated gates.

  • Own our approach to adversarial and red-team evaluation as a risk-reduction program, designing the test suites and failure taxonomies that catch safety and edge-case issues before they reach patients.

  • Work through complex, cross-team technical disagreements and drive alignment across engineering, product and AI leaders.

Turn hard, ambiguous problems into solutions other teams can build on

  • Originate new approaches and methodology that becomes a reusable standard rather than a one-off fix.

  • Take vague, cross-team pain points ("we do this manually and it's inconsistent") all the way from a rough idea to a fully-specified, shipped system, without needing to hand off any part of the journey.

  • Lead major platform improvements; re-architecting core systems, removing brittle logic, modernizing how things run with impact that's felt org-wide, not just on your own team.

Grow the people and the network around you

  • Build real working relationships across ML engineering, data science, platform engineering, clinical, legal, and product, the kind of trust that gets you looped in early, before decisions are locked in.

  • Become a go-to voice on evaluation methodology and data pipeline design: share what you've learned in internal talks, write things up so other teams can use them, and expect your ideas to shape how others approach similar problems.

  • Mentor other engineers, including experienced ones, and help raise the technical and statistical bar of the teams you work with.

You Have:

  • 10+ years of experience in ML infrastructure, data engineering, or evaluation/testing systems, with a track record of impact that reaches beyond a single team or project.

  • Hands-on depth in evaluation systems: designing and calibrating LLM judges/scorers, building statistically sound regression-testing methodology (e.g., paired significance testing with proper correction for multiple comparisons), measuring agreement against human labels, and designing adversarial/red-team evaluation approaches.

  • Hands-on depth in data pipeline engineering: dataset versioning, feature and benchmark pipelines, labeling and calibration workflows, and high-throughput ingestion and transformation systems.

  • A history of building things that became the standard approach for others — not just solving your own problem, but changing how a broader group of people tackle a category of problem.

  • Experience leading multi-team projects to completion, including navigating and resolving genuine technical disagreement along the way.

  • A track record of mentoring other engineers, including senior ones, and visibly raising the bar for the teams around you.

  • Excellent communication — comfortable adapting the same idea for different audiences, and confident building support for it well before launch.

  • Strong Python, and enough statistical fluency to design and defend a testing framework that real production decisions ride on.

Nice to Have:

  • Experience with Databricks, MLflow, Unity Catalog, or similar data/eval platforms.

  • Experience building reporting tools for people without direct engineering access (e.g., automated Slack digests, spreadsheet reports for non-technical teams).

  • Prior experience in a regulated industry (healthcare, fintech, life sciences).

  • A track record of company-wide talks or write-ups that changed how other teams approached a problem.

Our Benefits (there are more but here are some highlights):

  • Competitive salary & equity compensation for full-time roles

  • Unlimited PTO, company holidays, and quarterly mental health days

  • Comprehensive health benefits including medical, dental & vision, and parental leave

  • Employee Stock Purchase Program (ESPP)

  • 401k benefits with employer matching contribution

  • Offsite team retreats

We are committed to building a workforce that reflects diverse perspectives and prioritizes ethics, wellness, and a strong sense of belonging. If you're excited about this role, we encourage you to apply—even if you're not sure if your background or experience is a perfect match.

Hims considers all qualified applicants for employment, including applicants with arrest or conviction records, in accordance with the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance, the California Fair Chance Act, and any similar state or local fair chance laws.

It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.

Hims & Hers is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If you need assistance or an accommodation due to a disability, please contact us at [email protected] and describe the needed accommodation. Your privacy is important to us, and any information you share will only be used for the legitimate purpose of considering your request for accommodation. Hims & Hers gives consideration to all qualified applicants without regard to any protected status, including disability. Please do not send resumes to this email address.

To learn more about how we collect, use, retain, and disclose Personal Information, please visit our Global Candidate Privacy Statement.

Redirects to Hims & Hers's application page.

Other roles

More at Hims & Hers.

View all 11 roles