Staff Software Engineer, Cloud Sandboxes (West Coast)
Staff Engineer · Staff · Full Time · Remote
Opens Docker's application page
Role
What you'll do.
Join Docker's Cloud Sandboxes team as a Staff Software Engineer to architect and operate core distributed systems powering Docker's cloud-native agentic platform. This role focuses on designing scalable microVM orchestration, multi-tenant workload scheduling, and high-performance control plane systems that enable developers to deploy autonomous workflows securely and reliably. You'll partner with product and security teams to advance container infrastructure while solving complex distributed systems challenges at scale.
Responsibilities
- Design and Implement Core Platform Services: Architect, build, and deploy mission-critical services that form the foundation of Docker's Cloud Sandboxes platform, ensuring they meet enterprise-grade performance and reliability standards for processing billions of container operations.
- Build Scalable MicroVM Orchestration Systems: Develop distributed systems for microVM orchestration, workload scheduling, and lifecycle management that efficiently handle multi-tenant environments while maintaining security isolation and resource optimization across cloud infrastructure.
- Develop High-Performance Control Plane APIs: Build low-latency APIs and control plane components that manage multi-tenant workloads, enabling secure and efficient deployment of agentic workloads across Docker's cloud platform with industry-leading performance characteristics.
- Ensure System Reliability and Observability: Implement comprehensive monitoring, logging, and alerting strategies to maintain high availability and visibility across Docker's Cloud Sandbox infrastructure, meeting SLAs for critical platform services.
- Cross-Functional Collaboration: Partner with product management, platform engineering, and security teams to translate customer requirements into scalable technical capabilities and drive architectural decisions that balance developer experience with infrastructure security.
- Contribute to Architecture and Code Quality: Lead technical discussions, conduct thorough code reviews, author design documents, and establish best practices that elevate engineering standards across the Cloud Sandboxes team.
- Advance CI/CD Infrastructure: Drive automation initiatives and improvements to deployment pipelines, reducing time-to-market for platform features while ensuring build reliability and deployment safety across development and production environments.
- Debug Production Issues in Distributed Systems: Diagnose and resolve complex issues in production cloud environments using deep observability, systematic debugging techniques, and incident analysis to prevent future occurrences.
- On-Call Response and Incident Management: Participate in on-call rotation for the Cloud Sandboxes team, respond to critical incidents, execute incident response procedures, and drive continuous improvement of system resilience and mean-time-to-resolution.
Qualifications
What we look for.
Technical
Go or Java Proficiency
Advanced proficiency in Go and/or Java for building distributed backend systems, with demonstrated experience shipping production systems written in either or both languages.
Container Orchestration Expertise
Deep understanding of container orchestration platforms, particularly Kubernetes, including cluster management, workload scheduling, resource allocation, and operational patterns in production environments.
Distributed Systems Architecture
Strong expertise in microservices architecture patterns, service communication protocols, distributed consensus mechanisms, and designing systems that operate reliably across multiple nodes and availability zones.
Cloud Infrastructure Mastery
Hands-on experience designing and operating on AWS, Azure, or GCP, including compute services, networking, storage, and understanding of cloud-native scalability patterns and cost optimization.
Infrastructure Automation
Proficiency with infrastructure-as-code tools, CI/CD pipeline design and implementation, containerization best practices, and modern deployment automation frameworks.
High-Availability System Design
Experience designing, implementing, and operating production systems with stringent uptime requirements, including redundancy patterns, failover mechanisms, and disaster recovery strategies.
Observability and Monitoring
Expertise in implementing comprehensive monitoring, logging, and observability solutions for distributed systems, including metrics collection, tracing, alerting, and performance analysis.
Distributed Debugging and Troubleshooting
Advanced ability to diagnose issues in complex distributed environments using logs, metrics, traces, and systematic debugging methodologies specific to cloud-scale systems.
Education
Bachelor's Degree in Computer Science or Engineering
Bachelor's degree in Computer Science, Engineering, or related technical field from an accredited institution.
Equivalent Practical Experience
Equivalent professional experience demonstrating mastery of computer science fundamentals and software engineering principles through substantial production engineering work.
Experience
Large-Scale Backend System Development
Minimum 10+ years of professional backend software engineering experience building, scaling, and operating large-scale distributed systems that handle significant throughput and complex operational requirements.
Cloud and Distributed Systems Production Experience
Demonstrated track record of designing, implementing, and operating cloud-native or distributed systems in production environments at scale, with measurable impact on system performance or reliability.
Security and Multi-Tenancy
Practical experience implementing security controls in production systems, including multi-tenant isolation patterns, authentication, authorization, and compliance with enterprise security requirements.
Incident Response and On-Call Management
Experience participating in on-call rotations, investigating production incidents, implementing root cause analyses, and driving improvements in system reliability and incident response processes.
Skills
Required
Go Programming
Production-grade proficiency in Go for building backend services, microservices, and distributed systems with emphasis on performance, concurrency, and operational excellence.
Java Programming
Production-grade proficiency in Java for large-scale systems, with experience in modern frameworks, JVM optimization, and building highly concurrent applications.
Kubernetes Administration
Deep hands-on expertise with Kubernetes including cluster configuration, workload deployment, service discovery, storage management, and troubleshooting.
Microservices Architecture
Strong understanding of microservices design patterns, API gateway patterns, service-to-service communication, and distributed transaction handling.
Cloud Platform Expertise
Production experience with AWS (EC2, ECS, EKS, networking, storage), GCP (GKE, Compute Engine, Cloud Run), or Azure (AKS, Container Instances) for building scalable infrastructure.
System Design and Scalability
Ability to design large-scale systems that handle high throughput and complexity, including database sharding, caching strategies, load balancing, and performance optimization.
Infrastructure as Code
Practical experience with infrastructure automation using Terraform, CloudFormation, Helm, or similar declarative infrastructure tools for repeatable deployments.
Monitoring and Observability
Hands-on experience implementing observability across distributed systems using metrics, logs, and traces for comprehensive system visibility and troubleshooting.
Distributed Systems Debugging
Sophisticated troubleshooting methodology for complex multi-node systems including log analysis, metrics interpretation, and systematic problem isolation in production environments.
Technical Communication
Ability to clearly articulate technical architecture decisions, document complex systems, and collaborate effectively across remote, distributed teams using written and verbal communication.
Preferred
Cloud Platform Infrastructure Products
Nice to havePrior experience contributing to cloud-scale compute platforms, container orchestration systems, or managed infrastructure services at companies like Google Cloud, AWS, Azure, or similar organizations.
Service Mesh Architecture
Nice to haveHands-on experience with service mesh technologies (Istio, Linkerd, Consul) for managing service-to-service communication, security policies, and traffic management at scale.
Advanced Networking
Nice to haveDeep understanding of container networking, overlay networks, DNS resolution in distributed systems, network policies, and troubleshooting network-level issues in cloud environments.
Policy Enforcement Systems
Nice to haveExperience implementing policy enforcement, authorization frameworks, and admission control mechanisms in cloud-native environments for security and compliance.
Observability Stack Implementation
Nice to haveExpertise with observability platforms including Prometheus for metrics, OpenTelemetry for instrumentation, Grafana for visualization, and ELK or similar stacks for logging.
Multi-Tenant Security
Nice to haveAdvanced knowledge of security best practices specific to multi-tenant cloud systems including workload isolation, data protection, secure credential management, and audit logging.
Hyperscale Infrastructure Experience
Nice to haveBackground working at hyperscale companies or in developer infrastructure teams where you've operated systems handling millions of concurrent users or billions of transactions.
Container Ecosystem Knowledge
Nice to haveDeep familiarity with container technologies, image registries, container runtime security, and the broader Docker/container ecosystem.
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 170,350 – 275,550
Equity·Stock options
Benefits
Flexible Work Arrangement
Remote-first culture with full flexibility to structure your work around your life, supporting work-life balance and personal priorities.
Quarterly Whaleness Days Plus Extended Year-End Break
Designated quarterly wellness days coupled with an extended Whaleness break at year-end to recharge and disconnect from work.
Home Office Setup Support
Company investment in your home office environment to ensure you have a comfortable, productive workspace.
Paid Parental Leave
Comprehensive 16 weeks of paid parental leave available after 6 months of employment to support family growth.
Technology Stipend
Monthly technology stipend of $100 USD (net) to support home office upgrades, software subscriptions, or development tools.
Generous PTO Policy
Flexible time-off plan designed to encourage taking time for rest, personal pursuits, and experiences outside of work.
Professional Development Stipend
Training budget for conferences, courses, certifications, and classes to support continuous learning and career growth.
Equity Ownership
Stock options and equity grants aligned with company performance, enabling employees to participate in Docker's growth as a growing technology company.
Comprehensive Health Benefits
Medical benefits, retirement plans, and holiday schedules tailored to your country of residence, ensuring local compliance and support.
Global Office Access
Access to remote-first culture with physical offices in Seattle and Paris for occasional collaboration, team gatherings, or local hub work.
Docker Branded Merchandise
Exclusive Docker swag and company merchandise as part of the broader community experience.
Full posting
Original listing.
Docker has been one of the most loved brands in developer tooling, trusted by more than 20 million monthly users and over 20 billion container image pulls. From solo founders to the world's largest companies, developers rely on Docker to build, share, and run their applications across our suite of products including Docker Desktop, Docker Hub, and Docker Scout.
We are a globally distributed, remote-first team building the tools that define how software gets built and delivered. As AI agents redefine software development, Docker is at the center of that shift, providing the sandboxed environments, verified images, and secure infrastructure that make autonomous workflows trustworthy by default.
We are a globally distributed, remote-first team building the tools that define how software gets built and delivered. As AI agents redefine software development, Docker is at the center of that shift, providing the sandboxed environments, verified images, and secure infrastructure that make autonomous workflows trustworthy by default.
As a Staff Software Engineer on the Cloud Sandboxes team, you’ll design and build the core systems that power Docker’s cloud agentic platform. Your work will focus on creating scalable, reliable, and secure infrastructure that enables developers to deploy and manage agentic workloads efficiently and with confidence.
If you thrive on solving distributed systems challenges, enjoy working at the intersection of developer experience and cloud infrastructure, and want to help shape the future of Docker’s platform, we’d love to hear from you.
Responsibilities
Design, implement, and operate core services that power Docker’s Cloud Sandboxes platform
Build scalable systems for microVM orchestration, workload scheduling, and lifecycle management
Develop high-performance APIs and control plane components for managing multi-tenant workloads
Ensure system reliability, observability, and performance across Docker’s Cloud Sandbox infrastructure
Collaborate with product, platform, and security teams to deliver customer-focused capabilities
Participate in architectural discussions, code reviews, and design documents
Contribute to automation and CI/CD improvements across the deployment pipeline
Debug and resolve production issues across distributed systems in cloud environments
Take part in on-call rotation for your team; respond to incidents, debug production issues, and drive continuous improvement of system reliability
Qualifications
Required:
10+ years of backend software engineering experience building large-scale cloud or distributed systems
Strong proficiency in Go and/or Java
Deep understanding of container orchestration, Kubernetes, and microservices architecture
Experience designing and operating highly available, secure, and observable production systems
Strong understanding of cloud infrastructure (AWS, Azure, or GCP) and related scalability patterns
Familiarity with CI/CD pipelines, monitoring, and infrastructure-as-code tooling
Excellent problem-solving and debugging skills in distributed environments
Strong communication skills and ability to collaborate across remote, cross-functional teams
Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
Preferred:
Experience contributing to cloud-scale compute platforms or container infrastructure products
Knowledge of service mesh, networking, or policy enforcement systems
Experience with observability stacks (Prometheus, OpenTelemetry, Grafana, etc.)
Familiarity with security best practices for multi-tenant cloud systems
Prior experience in developer infrastructure, cloud platforms, or hyperscale environments
Docker considers sponsorship on a case-by-case basis based on business needs.
Perks
Freedom & flexibility; fit your work around your life
Designated quarterly Whaleness Days plus end of year Whaleness break
Home office setup; we want you comfortable while you work
16 weeks of paid Parental leave (after 6 months of employment)
Technology stipend equivalent to $100 USD net/month
PTO plan that encourages you to take time to do the things you enjoy
Training stipend for conferences, courses and classes
Equity; we are a growing start-up and want all employees to have a share in the success of the company
Docker Swag
Medical benefits, retirement and holidays vary by country
Remote-first culture, with offices in Seattle and Paris
Docker embraces diversity and equal opportunity. We are committed to building a team that represents a variety of backgrounds, perspectives, and skills. The more inclusive we are, the better our company will be.
#LI-REMOTE
Redirects to Docker's application page.
Other roles
More at Docker.
Senior Software Engineer, Secure Build
Senior
Staff Software Engineer, Developer Experience
Staff
Senior Software Engineer, Sandboxes (EU or East Coast Preferred)
Senior
Staff Software Engineer, Networking (Seattle or SF Bay Area)
Staff
Principal Software Engineer, Networking (Seattle or SF Bay Area)
Principal