ML Engineer

Senior · Full Time · Remote

Palo Alto, CA · RemoteUSD 139k – 226k1mo ago
Apply for this role

Opens Docker's application page

Role

What you'll do.

Docker is seeking a senior-level Machine Learning Engineer to join their Intelligence team, focusing on building advanced AI-driven security and governance capabilities for container environments. The role involves developing cutting-edge ML systems that enhance Docker's platform security, with a particular emphasis on creating intelligent, trustworthy autonomous workflows.

Responsibilities

  • ML System Development: Design, train, evaluate, and deploy ML systems for governance and security capabilities, including prompt injection detection, behavioral anomaly detection, and trust scoring
  • Infrastructure Engineering: Build and maintain supporting ML infrastructure including data pipelines, feature stores, model serving, and evaluation harnesses
  • Technical Leadership: Set technical direction for ML work, own system architecture, evaluation methodology, and model lifecycle management
  • Strategic Decision Making: Make pragmatic build-vs-buy decisions, leveraging frontier models and managed services while identifying opportunities for custom system development
  • Team Development: Participate in recruiting, mentoring, and team growth for the Intelligence organization
  • Operational Responsibility: Potentially participate in 24/7 on-call rotation for the Agentic Platform, maintaining service reliability and performance

Qualifications

What we look for.

Technical

  • Machine Learning Expertise

    Deep applied ML/AI expertise with proven track record of shipping production systems, especially in fraud, abuse, safety, security, or trust domains

  • Software Engineering Skills

    Minimum 4+ years of professional backend, infrastructure, or platform engineering experience

  • ML Infrastructure

    Experience building and owning complete ML system ecosystems, including data pipelines, model serving, evaluation, and monitoring

  • AI Technology Proficiency

    Advanced experience with LLM-based systems, including evaluation, prompt engineering, fine-tuning, retrieval, and agent frameworks

Education

  • Academic Background

    Bachelor's degree in Computer Science, Engineering, or related field; equivalent practical experience acceptable

Experience

  • Professional Experience

    5+ years of applied ML/AI expertise with production system deployment

  • Domain Expertise

    Experience in handling adversarial dynamics, imbalanced data, and high-stakes decision-making environments

Skills

Required

  • Machine Learning

    Advanced ML model development, training, and deployment skills

  • Programming

    Proficient in Python, with strong software engineering capabilities

  • AI Technologies

    Expertise in Large Language Models, prompt engineering, and AI safety techniques

Preferred

  • Cloud Platforms

    Nice to have

    Experience with AWS, GCP, or Azure ML services

  • Container Technologies

    Nice to have

    Deep understanding of containerization and Docker ecosystem

Tech stack

Languages

Python

Frameworks

TensorFlowPyTorch

Databases

PostgreSQL

Tools

DockerKubernetes

Other

MLflow

Compensation

Pay and benefits.

Base·USD 138,500 – 225,500

Equity·Stock options

Benefits

  • Remote Work

    Fully remote-first culture with flexible work arrangements

  • Parental Leave

    16 weeks of paid parental leave after 6 months of employment

  • Technology Stipend

    $100 monthly technology allowance for home office setup

  • Professional Development

    Training stipend for conferences, courses, and classes

  • Equity

    Stock options to share in company's growth and success

Process

Interview steps.

  1. 01

    Initial Screening

    HR review of application and initial qualifications

  2. 02

    Technical Phone Screen

    Detailed discussion of ML engineering experience and technical capabilities

  3. 03

    Technical Interviews

    Multiple rounds of in-depth technical interviews focusing on ML systems design, coding, and problem-solving

  4. 04

    Hiring Manager Interview

    Strategic discussion about team fit, technical vision, and role expectations

  5. 05

    Final Interview

    Comprehensive evaluation of technical skills, team compatibility, and overall potential contribution

Full posting

Original listing.

Docker has been one of the most loved brands in developer tooling, trusted by more than 20 million monthly users and over 20 billion container image pulls. From solo founders to the world's largest companies, developers rely on Docker to build, share, and run their applications across our suite of products including Docker Desktop, Docker Hub, and Docker Scout.

We are a globally distributed, remote-first team building the tools that define how software gets built and delivered. As AI agents redefine software development, Docker is at the center of that shift, providing the sandboxed environments, verified images, and secure infrastructure that make autonomous workflows trustworthy by default.

Docker's long-term vision is to become the runtime for trusted autonomy. As agents become more capable and autonomous, governance, policy, identity, and audit become foundational.

The Intelligence team builds intelligence-driven product capabilities that make software and agent execution on Docker safer, more effective, more trustworthy, and more efficient. Because Docker sits at the intersection of models, tools, software, identities, credentials, networks, and execution, we have visibility into behavior and context few other platforms can see, and we think that visibility is the foundation for a new layer of value across the platform.

About the role

We're hiring a ML Engineer as one of the founding engineers on Intelligence Org. You'll work directly with the team's first engineers and manager to figure out what to build, how to build it, and how it fits into the broader Docker platform. This is a hands-on builder role with staff-level scope: you'll shape technical direction, ship the first versions of intelligence capabilities into customer hands, and grow the foundations (data, evaluation, infrastructure) the team will rely on as it scales.

Responsibilities

  • Design, train, evaluate, and ship ML systems that power governance and security capabilities, starting with problems like prompt injection detection, behavioral anomaly detection, trust scoring, and policy recommendations.

  • Build the supporting infrastructure: data pipelines, feature stores, model serving, evaluation harnesses, and the feedback loops that make iteration fast.

  • Make pragmatic build-vs-buy calls. Use frontier models, off-the-shelf tooling, and managed services to move quickly; invest in custom systems where they create durable advantage.

  • Set technical direction for the team's ML work. Own the architecture, evaluation methodology, model lifecycle, and the bar for shipping.

  • Help recruit, mentor, and shape the team as it grows.

  • This role may require participation in a 24/7 on-call rotation for the Agentic Platform; carry genuine pager responsibility for the services you build and operate

Qualifications

  • 5+ years of deep applied ML/AI expertise with a track record of shipping production systems. Experience in fraud, abuse, safety, security, or trust domains, where adversarial dynamics, imbalanced data, and high-stakes decisions is valuable.

  • 4+ years of professional, hands-on, full-time software engineering experience in backend, infrastructure, or platform engineering.

  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience

  • You've built and owned the systems around ML models, i.e. data pipelines, serving, evaluation, monitoring etc. and have shipped customer-facing products end to end.

  • You use modern AI tools fluently in your day-to-day work and have a sharp instinct for when frontier models can replace traditional ML, when they can't, and when to combine the two.

  • Experience with LLM-based systems in production - evaluation, prompt engineering, fine-tuning, retrieval, guardrails, agent frameworks.

  • Familiarity with the agent / MCP ecosystem.

  • You're energized by an early-stage effort where the roadmap is being written as the work happens, and you make crisp decisions with incomplete information.

  • Collaborative and low-ego. You work well across teams, write clearly, and bring others along.

Docker considers visa sponsorship on a case-by-case basis based on business needs.

Perks

  • Freedom & flexibility; fit your work around your life

  • Designated quarterly Whaleness Days plus end of year Whaleness break

  • Home office setup; we want you comfortable while you work

  • 16 weeks of paid Parental leave (after 6 months of employment)

  • Technology stipend equivalent to $100 USD net/month

  • PTO plan that encourages you to take time to do the things you enjoy

  • Training stipend for conferences, courses and classes

  • Equity; we are a growing start-up and want all employees to have a share in the success of the company

  • Docker Swag

  • Medical benefits, retirement and holidays vary by country

  • Remote-first culture, with offices in Seattle and Paris

Docker embraces diversity and equal opportunity. We are committed to building a team that represents a variety of backgrounds, perspectives, and skills. The more inclusive we are, the better our company will be.

#LI-REMOTE

Redirects to Docker's application page.

Other roles

More at Docker.

View all 24 roles