Software Engineer, Infrastructure Platform

Infrastructure Platform Engineer · Senior · Full Time · Remote

Canada · RemoteCAD 130k – 180k4mo ago
Apply for this role

Opens Docker's application page

Role

What you'll do.

Docker is seeking a skilled Software Engineer for its Infrastructure Platform team to build and operate cloud-native platform services. The ideal candidate will focus on creating self-service infrastructure solutions, reducing operational toil through automation, and improving platform reliability using cutting-edge technologies like Kubernetes, Go, and AI-assisted workflows.

Responsibilities

  • Self-Service Platform Services: Build and operate internal platform services and APIs in Go, creating golden paths for self-serve onboarding, deployment, and operational workflows with clear documentation and measurable outcomes.
  • Infrastructure as Code and Reliability: Implement infrastructure automation using Terraform and GitOps practices, define SLOs, improve alerting, and contribute to safe delivery patterns with testing gates and rollback mechanisms.
  • Kubernetes and Networking Foundations: Operate and scale multi-tenant EKS clusters, manage traffic and ingress systems, and continuously evaluate and adopt improvements with incremental rollout strategies.
  • AI and Agentic Workflows: Develop and iterate on AI-powered operational workflows that reduce toil, including automated triage, context gathering, and safe runbook execution with strong observability and auditability.
  • Incident Response and On-Call: Participate in on-call rotations, engage in incident response, and contribute to a culture of sustainable reliability through blameless postmortems and preventative measures.

Qualifications

What we look for.

Technical

  • Backend Development

    4+ years of backend software engineering experience in cloud or distributed systems

  • Programming Languages

    Strong software development skills in Go, with expertise in design, testing, debugging, and code review

  • Cloud Infrastructure

    Experience shipping and operating cloud services in production with a solid foundation in Linux, networking, and cloud security

Education

  • Computer Science

    Bachelor's degree in Computer Science, Software Engineering, or related technical field preferred

Experience

  • Production Systems

    Minimum 3 years of experience operating cloud services with a focus on impact and skill

  • Operational Automation

    Proven track record of building operational automation with emphasis on safety, guardrails, and auditability

Skills

Required

  • Go Programming

    Proficient in Go language development for backend and infrastructure services

  • Cloud Native Technologies

    Strong understanding of Kubernetes, containerization, and cloud infrastructure

  • Infrastructure as Code

    Experience with Terraform and GitOps practices for infrastructure automation

Preferred

  • Kubernetes Expertise

    Nice to have

    Advanced knowledge of EKS, ingress, CNI, service mesh, and load balancing

  • Observability Tools

    Nice to have

    Familiarity with OpenTelemetry, Prometheus, Grafana, and SLO practices

  • CI/CD

    Nice to have

    Experience with GitHub Actions, Argo CD, canary deployments, and automated rollbacks

Tech stack

Languages

Go

Frameworks

Terraform

Databases

Not Specified

Tools

KubernetesEKSGitHub Actions

Other

OpenTelemetryPrometheusGrafana

Compensation

Pay and benefits.

Base·CAD 130,000 – 180,000

Equity·Stock options

Benefits

  • Remote Work

    Flexible, remote-first work arrangement with global team collaboration

  • Home Office Setup

    Stipend for comfortable home office equipment

  • Technology Allowance

    $100 monthly net technology stipend

  • Parental Leave

    16 weeks of paid parental leave

  • Training Support

    Stipend for conferences, courses, and professional development

  • Equity

    Stock options to share in company's growth and success

  • Paid Time Off

    Flexible PTO plan encouraging work-life balance

  • Whaleness Days

    Quarterly designated days off and end-of-year break

Process

Interview steps.

  1. 01

    Initial Screening

    Phone or video call with recruiting team to discuss background and role fit

  2. 02

    Technical Interview

    In-depth technical discussion focusing on infrastructure, cloud-native technologies, and system design

  3. 03

    Coding Assessment

    Practical coding challenge to evaluate Go programming skills and infrastructure automation capabilities

  4. 04

    Team Match Interview

    Interviews with potential team members to assess cultural and technical alignment

  5. 05

    Final Interview

    Meeting with hiring manager to discuss role expectations, team dynamics, and career growth opportunities

Full posting

Original listing.

At Docker, we make app development easier so developers can focus on what matters. Our remote-first team spans the globe, united by a passion for innovation and great developer experiences. With over 20 million monthly users and 20 billion image pulls, Docker is the #1 tool for building, sharing, and running apps—trusted by startups and Fortune 100s alike. We’re growing fast and just getting started. Come join us for a whale of a ride!

Our Infrastructure Engineering team builds and operates the cloud-native platform that powers Docker’s suite of products. We design resilient services, automate where it helps most, and measure what matters so hundreds of engineers can ship safely to millions of users every day.

A core focus is self-service. We build paved-road platform capabilities that let internal teams provision, deploy, observe, and operate services with minimal friction and strong guardrails. We treat the platform as a product with clear contracts, well-defined defaults, and great documentation. Success is measured by adoption and fewer support requests.

How We Work

  • Write it down, ship it, iterate: RFCs and design docs, code review, and small safe releases.

  • Sustainable reliability: we prioritize root-cause fixes, good alerts, and automation over heroics.

  • Cross-functional by default: we partner closely with product and security teams.

  • AI-accelerated execution: we build agentic workflows to reduce toil and improve incident response, with guardrails, auditability, and human review.

What You’ll Work On

  • Reducing toil through automation, including AI-assisted and agentic operational workflows.

  • Building self-service onboarding and deployment workflows that reduce tickets and speed delivery.

  • Scaling Kubernetes foundations and evolving our traffic and ingress stack.

Responsibilities

1) Self-Service Platform Services

  • Build and operate internal platform services and APIs in Go, including provisioning, quotas and policies, cost insights, and platform workflows.

  • Deliver golden paths for self-serve onboarding and day-2 operations, including access, deployment setup, observability defaults, and governance guardrails.

  • Partner with teams to drive adoption through clear docs, examples, and measurable outcomes.

2) Infrastructure as Code and Reliability

  • Codify infrastructure with Terraform and GitOps practices, and contribute to platform tooling in Go.

  • Define and improve SLOs, alerting, and operational readiness. Participate in incident response and preventive follow-ups.

  • Help standardize safe delivery patterns, including testing gates, canaries, and rollback triggers, so deployments are routine and low-risk.

3) Kubernetes and Networking Foundations

  • Operate and scale multi-tenant EKS clusters and traffic and ingress systems to deliver secure, reliable routing.

  • Evaluate and adopt improvements with a bias toward incremental rollout and measurable impact.

4) AI and Agentic Workflows for Reliability

  • Build and iterate on agentic workflows that reduce operational toil, including triage support, context gathering, safe runbook execution, and remediation suggestions.

  • Integrate automation into delivery and operations in a way that is safe, observable, and auditable.

5) On-Call and Incident Response

Operational ownership is part of this role.

  • You’ll join an on-call rotation after onboarding and shadowing, and participate in incident response during your shifts.

  • We aim for sustainable on-call through good alerting, automation, and blameless postmortems focused on prevention.

Qualifications

Core Engineering Skills (must-have)

  • 4+ years of backend software engineering experience building large-scale cloud or distributed systems

  • Strong software development skills in Go or a similar language, including design, testing, debugging, and code review.

  • Experience shipping and operating cloud services in production, often 3+ years. We hire for skill and impact, not years alone.

  • Solid foundation in Linux, networking fundamentals, and cloud security.

  • Experience building operational automation, including AI-assisted or agentic workflows, with an emphasis on safety, guardrails, and auditability.

  • Clear written and verbal communication in a remote environment, including RFCs, incident writeups, and async collaboration.

Nice-to-have

  • Kubernetes and EKS experience, plus ingress, CNI, service mesh, and familiarity with L4 and L7 load balancing.

  • Observability tooling such as OpenTelemetry, Prometheus, and Grafana, plus alerting and SLO practice.

  • CI/CD and progressive delivery, including GitHub Actions or Argo CD, canaries, and automated rollback.

  • Cost optimization at scale, including FinOps and capacity modeling.

  • Distributed systems, containers, and Go-based platform tooling.

We value depth in one area and curiosity across others, and we will help you grow in the rest.

What to Expect

First 30 Days

  • Ship your first change to a Terraform module or internal service and learn how we operate.

  • Shadow on-call and build context on our platform and reliability priorities.

First 90 Days

  • Own a component and deliver an improvement from design to production with measurable impact.

  • Join the on-call rotation and contribute effectively during your shifts.

First Year

  • Lead or co-lead a meaningful platform initiative, with scope that scales by level, and help reduce toil through automation.

  • Become a trusted contributor in one or more areas such as platform services, Kubernetes and networking foundations, or reliability automation.

Docker considers sponsorship on a case-by-case basis based on business needs.

We use Covey as part of our hiring and / or promotional process for jobs in NYC and certain features may qualify it as an AEDT. As part of the evaluation process we provide Covey with job requirements and candidate submitted applications. We began using Covey Scout for Inbound on April 13, 2024.

Please see the independent bias audit report covering our use of Covey here.

Perks

  • Freedom & flexibility; fit your work around your life

  • Designated quarterly Whaleness Days plus end of year Whaleness break

  • Home office setup; we want you comfortable while you work

  • 16 weeks of paid Parental leave

  • Technology stipend equivalent to $100 net/month

  • PTO plan that encourages you to take time to do the things you enjoy

  • Training stipend for conferences, courses and classes

  • Equity; we are a growing start-up and want all employees to have a share in the success of the company

  • Docker Swag

  • Medical benefits, retirement and holidays vary by country

  • Remote-first culture, with offices in Seattle and Paris

Docker embraces diversity and equal opportunity. We are committed to building a team that represents a variety of backgrounds, perspectives, and skills. The more inclusive we are, the better our company will be.

#LI-REMOTE

Redirects to Docker's application page.

Other roles

More at Docker.

View all 21 roles