Staff Software Engineer, Cloud Sandboxes (West Coast)

Staff Engineer · Staff · Full Time · Remote

Seattle, WA · RemoteUSD 170k – 276k1mo ago
Apply for this role

Opens Docker's application page

Role

What you'll do.

Join Docker's Cloud Sandboxes team as a Staff Software Engineer to architect and operate core distributed systems powering Docker's cloud-native agentic platform. This role focuses on designing scalable microVM orchestration, multi-tenant workload scheduling, and high-performance control plane systems that enable developers to deploy autonomous workflows securely and reliably. You'll partner with product and security teams to advance container infrastructure while solving complex distributed systems challenges at scale.

Responsibilities

  • Design and Implement Core Platform Services: Architect, build, and deploy mission-critical services that form the foundation of Docker's Cloud Sandboxes platform, ensuring they meet enterprise-grade performance and reliability standards for processing billions of container operations.
  • Build Scalable MicroVM Orchestration Systems: Develop distributed systems for microVM orchestration, workload scheduling, and lifecycle management that efficiently handle multi-tenant environments while maintaining security isolation and resource optimization across cloud infrastructure.
  • Develop High-Performance Control Plane APIs: Build low-latency APIs and control plane components that manage multi-tenant workloads, enabling secure and efficient deployment of agentic workloads across Docker's cloud platform with industry-leading performance characteristics.
  • Ensure System Reliability and Observability: Implement comprehensive monitoring, logging, and alerting strategies to maintain high availability and visibility across Docker's Cloud Sandbox infrastructure, meeting SLAs for critical platform services.
  • Cross-Functional Collaboration: Partner with product management, platform engineering, and security teams to translate customer requirements into scalable technical capabilities and drive architectural decisions that balance developer experience with infrastructure security.
  • Contribute to Architecture and Code Quality: Lead technical discussions, conduct thorough code reviews, author design documents, and establish best practices that elevate engineering standards across the Cloud Sandboxes team.
  • Advance CI/CD Infrastructure: Drive automation initiatives and improvements to deployment pipelines, reducing time-to-market for platform features while ensuring build reliability and deployment safety across development and production environments.
  • Debug Production Issues in Distributed Systems: Diagnose and resolve complex issues in production cloud environments using deep observability, systematic debugging techniques, and incident analysis to prevent future occurrences.
  • On-Call Response and Incident Management: Participate in on-call rotation for the Cloud Sandboxes team, respond to critical incidents, execute incident response procedures, and drive continuous improvement of system resilience and mean-time-to-resolution.

Qualifications

What we look for.

Technical

  • Go or Java Proficiency

    Advanced proficiency in Go and/or Java for building distributed backend systems, with demonstrated experience shipping production systems written in either or both languages.

  • Container Orchestration Expertise

    Deep understanding of container orchestration platforms, particularly Kubernetes, including cluster management, workload scheduling, resource allocation, and operational patterns in production environments.

  • Distributed Systems Architecture

    Strong expertise in microservices architecture patterns, service communication protocols, distributed consensus mechanisms, and designing systems that operate reliably across multiple nodes and availability zones.

  • Cloud Infrastructure Mastery

    Hands-on experience designing and operating on AWS, Azure, or GCP, including compute services, networking, storage, and understanding of cloud-native scalability patterns and cost optimization.

  • Infrastructure Automation

    Proficiency with infrastructure-as-code tools, CI/CD pipeline design and implementation, containerization best practices, and modern deployment automation frameworks.

  • High-Availability System Design

    Experience designing, implementing, and operating production systems with stringent uptime requirements, including redundancy patterns, failover mechanisms, and disaster recovery strategies.

  • Observability and Monitoring

    Expertise in implementing comprehensive monitoring, logging, and observability solutions for distributed systems, including metrics collection, tracing, alerting, and performance analysis.

  • Distributed Debugging and Troubleshooting

    Advanced ability to diagnose issues in complex distributed environments using logs, metrics, traces, and systematic debugging methodologies specific to cloud-scale systems.

Education

  • Bachelor's Degree in Computer Science or Engineering

    Bachelor's degree in Computer Science, Engineering, or related technical field from an accredited institution.

  • Equivalent Practical Experience

    Equivalent professional experience demonstrating mastery of computer science fundamentals and software engineering principles through substantial production engineering work.

Experience

  • Large-Scale Backend System Development

    Minimum 10+ years of professional backend software engineering experience building, scaling, and operating large-scale distributed systems that handle significant throughput and complex operational requirements.

  • Cloud and Distributed Systems Production Experience

    Demonstrated track record of designing, implementing, and operating cloud-native or distributed systems in production environments at scale, with measurable impact on system performance or reliability.

  • Security and Multi-Tenancy

    Practical experience implementing security controls in production systems, including multi-tenant isolation patterns, authentication, authorization, and compliance with enterprise security requirements.

  • Incident Response and On-Call Management

    Experience participating in on-call rotations, investigating production incidents, implementing root cause analyses, and driving improvements in system reliability and incident response processes.

Skills

Required

  • Go Programming

    Production-grade proficiency in Go for building backend services, microservices, and distributed systems with emphasis on performance, concurrency, and operational excellence.

  • Java Programming

    Production-grade proficiency in Java for large-scale systems, with experience in modern frameworks, JVM optimization, and building highly concurrent applications.

  • Kubernetes Administration

    Deep hands-on expertise with Kubernetes including cluster configuration, workload deployment, service discovery, storage management, and troubleshooting.

  • Microservices Architecture

    Strong understanding of microservices design patterns, API gateway patterns, service-to-service communication, and distributed transaction handling.

  • Cloud Platform Expertise

    Production experience with AWS (EC2, ECS, EKS, networking, storage), GCP (GKE, Compute Engine, Cloud Run), or Azure (AKS, Container Instances) for building scalable infrastructure.

  • System Design and Scalability

    Ability to design large-scale systems that handle high throughput and complexity, including database sharding, caching strategies, load balancing, and performance optimization.

  • Infrastructure as Code

    Practical experience with infrastructure automation using Terraform, CloudFormation, Helm, or similar declarative infrastructure tools for repeatable deployments.

  • Monitoring and Observability

    Hands-on experience implementing observability across distributed systems using metrics, logs, and traces for comprehensive system visibility and troubleshooting.

  • Distributed Systems Debugging

    Sophisticated troubleshooting methodology for complex multi-node systems including log analysis, metrics interpretation, and systematic problem isolation in production environments.

  • Technical Communication

    Ability to clearly articulate technical architecture decisions, document complex systems, and collaborate effectively across remote, distributed teams using written and verbal communication.

Preferred

  • Cloud Platform Infrastructure Products

    Nice to have

    Prior experience contributing to cloud-scale compute platforms, container orchestration systems, or managed infrastructure services at companies like Google Cloud, AWS, Azure, or similar organizations.

  • Service Mesh Architecture

    Nice to have

    Hands-on experience with service mesh technologies (Istio, Linkerd, Consul) for managing service-to-service communication, security policies, and traffic management at scale.

  • Advanced Networking

    Nice to have

    Deep understanding of container networking, overlay networks, DNS resolution in distributed systems, network policies, and troubleshooting network-level issues in cloud environments.

  • Policy Enforcement Systems

    Nice to have

    Experience implementing policy enforcement, authorization frameworks, and admission control mechanisms in cloud-native environments for security and compliance.

  • Observability Stack Implementation

    Nice to have

    Expertise with observability platforms including Prometheus for metrics, OpenTelemetry for instrumentation, Grafana for visualization, and ELK or similar stacks for logging.

  • Multi-Tenant Security

    Nice to have

    Advanced knowledge of security best practices specific to multi-tenant cloud systems including workload isolation, data protection, secure credential management, and audit logging.

  • Hyperscale Infrastructure Experience

    Nice to have

    Background working at hyperscale companies or in developer infrastructure teams where you've operated systems handling millions of concurrent users or billions of transactions.

  • Container Ecosystem Knowledge

    Nice to have

    Deep familiarity with container technologies, image registries, container runtime security, and the broader Docker/container ecosystem.

Tech stack

Languages

GoJava

Frameworks

KubernetesMicroservices ArchitectureDocker Ecosystem

Databases

Distributed DatabasesCloud-Native Data Stores

Tools

CI/CD Pipeline ToolsInfrastructure as Code ToolsPrometheusGrafanaContainer Runtime Technologies

Other

OpenTelemetryService Mesh TechnologiesCloud Security and ComplianceDistributed Consensus Mechanisms

Compensation

Pay and benefits.

Base·USD 170,350 – 275,550

Equity·Stock options

Benefits

  • Flexible Work Arrangement

    Remote-first culture with full flexibility to structure your work around your life, supporting work-life balance and personal priorities.

  • Quarterly Whaleness Days Plus Extended Year-End Break

    Designated quarterly wellness days coupled with an extended Whaleness break at year-end to recharge and disconnect from work.

  • Home Office Setup Support

    Company investment in your home office environment to ensure you have a comfortable, productive workspace.

  • Paid Parental Leave

    Comprehensive 16 weeks of paid parental leave available after 6 months of employment to support family growth.

  • Technology Stipend

    Monthly technology stipend of $100 USD (net) to support home office upgrades, software subscriptions, or development tools.

  • Generous PTO Policy

    Flexible time-off plan designed to encourage taking time for rest, personal pursuits, and experiences outside of work.

  • Professional Development Stipend

    Training budget for conferences, courses, certifications, and classes to support continuous learning and career growth.

  • Equity Ownership

    Stock options and equity grants aligned with company performance, enabling employees to participate in Docker's growth as a growing technology company.

  • Comprehensive Health Benefits

    Medical benefits, retirement plans, and holiday schedules tailored to your country of residence, ensuring local compliance and support.

  • Global Office Access

    Access to remote-first culture with physical offices in Seattle and Paris for occasional collaboration, team gatherings, or local hub work.

  • Docker Branded Merchandise

    Exclusive Docker swag and company merchandise as part of the broader community experience.

Full posting

Original listing.

Docker has been one of the most loved brands in developer tooling, trusted by more than 20 million monthly users and over 20 billion container image pulls. From solo founders to the world's largest companies, developers rely on Docker to build, share, and run their applications across our suite of products including Docker Desktop, Docker Hub, and Docker Scout.

We are a globally distributed, remote-first team building the tools that define how software gets built and delivered. As AI agents redefine software development, Docker is at the center of that shift, providing the sandboxed environments, verified images, and secure infrastructure that make autonomous workflows trustworthy by default.

We are a globally distributed, remote-first team building the tools that define how software gets built and delivered. As AI agents redefine software development, Docker is at the center of that shift, providing the sandboxed environments, verified images, and secure infrastructure that make autonomous workflows trustworthy by default.

As a Staff Software Engineer on the Cloud Sandboxes team, you’ll design and build the core systems that power Docker’s cloud agentic platform. Your work will focus on creating scalable, reliable, and secure infrastructure that enables developers to deploy and manage agentic workloads efficiently and with confidence.

If you thrive on solving distributed systems challenges, enjoy working at the intersection of developer experience and cloud infrastructure, and want to help shape the future of Docker’s platform, we’d love to hear from you.

Responsibilities

  • Design, implement, and operate core services that power Docker’s Cloud Sandboxes platform

  • Build scalable systems for microVM orchestration, workload scheduling, and lifecycle management

  • Develop high-performance APIs and control plane components for managing multi-tenant workloads

  • Ensure system reliability, observability, and performance across Docker’s Cloud Sandbox infrastructure

  • Collaborate with product, platform, and security teams to deliver customer-focused capabilities

  • Participate in architectural discussions, code reviews, and design documents

  • Contribute to automation and CI/CD improvements across the deployment pipeline

  • Debug and resolve production issues across distributed systems in cloud environments

  • Take part in on-call rotation for your team; respond to incidents, debug production issues, and drive continuous improvement of system reliability

Qualifications

Required:

  • 10+ years of backend software engineering experience building large-scale cloud or distributed systems

  • Strong proficiency in Go and/or Java

  • Deep understanding of container orchestration, Kubernetes, and microservices architecture

  • Experience designing and operating highly available, secure, and observable production systems

  • Strong understanding of cloud infrastructure (AWS, Azure, or GCP) and related scalability patterns

  • Familiarity with CI/CD pipelines, monitoring, and infrastructure-as-code tooling

  • Excellent problem-solving and debugging skills in distributed environments

  • Strong communication skills and ability to collaborate across remote, cross-functional teams

  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.

Preferred:

  • Experience contributing to cloud-scale compute platforms or container infrastructure products

  • Knowledge of service mesh, networking, or policy enforcement systems

  • Experience with observability stacks (Prometheus, OpenTelemetry, Grafana, etc.)

  • Familiarity with security best practices for multi-tenant cloud systems

  • Prior experience in developer infrastructure, cloud platforms, or hyperscale environments

Docker considers sponsorship on a case-by-case basis based on business needs.

Perks

  • Freedom & flexibility; fit your work around your life

  • Designated quarterly Whaleness Days plus end of year Whaleness break

  • Home office setup; we want you comfortable while you work

  • 16 weeks of paid Parental leave (after 6 months of employment)

  • Technology stipend equivalent to $100 USD net/month

  • PTO plan that encourages you to take time to do the things you enjoy

  • Training stipend for conferences, courses and classes

  • Equity; we are a growing start-up and want all employees to have a share in the success of the company

  • Docker Swag

  • Medical benefits, retirement and holidays vary by country

  • Remote-first culture, with offices in Seattle and Paris

Docker embraces diversity and equal opportunity. We are committed to building a team that represents a variety of backgrounds, perspectives, and skills. The more inclusive we are, the better our company will be.

#LI-REMOTE

Redirects to Docker's application page.

Other roles

More at Docker.

View all 21 roles