Software Engineer, Core Infrastructure (Mid-Senior level)

Infrastructure Engineer · Mid · Full Time

TorontoCAD 160k – 220k1mo ago
Apply for this role

Opens Zip's application page

Role

What you'll do.

Join Zip's Core Infrastructure team as a Software Engineer to lead the design, build, and operation of global multi-region, highly scalable infrastructure systems. This mid-senior level role offers end-to-end ownership of critical infrastructure components including Kubernetes platforms, AI infrastructure, observability, and deployment pipelines while collaborating with world-class engineers from Apple, Airbnb, and Meta to support enterprise procurement innovation at scale.

Responsibilities

  • Multi-Region Architecture Development: Design, build, and operate global multi-cell architectures that enable Zip's enterprise platform to scale across regions while maintaining performance, reliability, and data sovereignty requirements for Fortune 500 clients.
  • Infrastructure Component Ownership: Assume end-to-end ownership of one or more core infrastructure systems including Kubernetes platform management, AI infrastructure provisioning, observability stacks, deployment pipelines, and cost optimization initiatives that directly impact platform reliability and operational efficiency.
  • Scalable System Design and Implementation: Champion technical excellence by leading architectural design, implementation, and operational excellence of highly scalable systems that support Zip's $500+ billion annual spend processing across thousands of enterprise customers.
  • Cross-Functional Engineering Collaboration: Partner strategically with product engineering, platform, and data teams across multiple time zones to design infrastructure solutions that accelerate Zip's product roadmap and enable rapid feature deployment without compromising system stability.
  • Platform Reliability and Operations: Establish and maintain SLOs, incident response procedures, and operational runbooks for critical infrastructure components; conduct post-incident reviews and implement preventative measures to continuously improve platform availability and performance.
  • Infrastructure Modernization and Optimization: Evaluate, adopt, and optimize emerging infrastructure technologies and patterns; lead initiatives to improve deployment velocity, reduce operational overhead, and optimize cloud infrastructure costs across the organization.

Qualifications

What we look for.

Technical

  • Large-Scale Distributed Systems

    Proven experience designing, building, and operating distributed systems at scale, with deep understanding of consistency models, fault tolerance, load balancing, and multi-region deployment patterns.

  • Container Orchestration and Kubernetes

    Advanced proficiency in Kubernetes administration, including cluster provisioning, networking, persistent storage, security policies, and operational management at production scale.

  • Cloud Infrastructure Platforms

    Hands-on expertise with major cloud providers (AWS preferred based on company tech stack) including infrastructure-as-code, networking, compute optimization, and cost management strategies.

  • Infrastructure Automation and DevOps

    Strong experience with infrastructure-as-code tools (Terraform, CloudFormation), CI/CD pipeline design, deployment automation, and configuration management in production environments.

  • Observability and Monitoring

    Deep expertise with observability platforms (Datadog preferred), including metrics collection, log aggregation, distributed tracing, alerting strategies, and performance optimization.

  • Database Systems and Data Infrastructure

    Strong understanding of relational and non-relational databases, data pipelines, and query optimization; experience with migration strategies and multi-region data replication patterns.

Education

  • Bachelor's Degree in Computer Science or Related Field

    Bachelor's degree in Computer Science, Software Engineering, Physics, Mathematics, or equivalent technical discipline; advanced degree beneficial but not required.

Experience

  • 4+ Years Software Engineering Experience

    Minimum 4 years of professional software engineering experience with demonstrated progression in scope and complexity, preferably including infrastructure, backend systems, or platform engineering focus.

  • Large-Scale Platform Operations

    Proven track record of independently building, deploying, and operating large-scale platforms or infrastructure systems in production environments supporting millions of users or processing significant data volumes.

  • Multi-Team Collaboration and Communication

    Demonstrated ability to communicate technical concepts effectively across diverse stakeholder groups, including non-technical audiences; proven success collaborating with distributed teams across multiple time zones.

  • Architectural Decision Making

    Experience leading architectural decisions, evaluating technology trade-offs, and making sound technical judgments that balance performance, scalability, maintainability, and business requirements.

Skills

Required

  • Kubernetes and Container Orchestration

    Production-level expertise in Kubernetes cluster design, management, and optimization including networking, storage, security, and RBAC policies.

  • AWS Cloud Services

    Strong proficiency with AWS services including EC2, RDS, S3, VPC, IAM, CloudFormation, and other core infrastructure services used in enterprise deployments.

  • Infrastructure-as-Code (IaC)

    Advanced experience with Terraform, CloudFormation, or similar IaC tools for automating infrastructure provisioning, versioning, and management.

  • Distributed Systems Design

    Strong grasp of distributed system principles including consensus algorithms, replication strategies, failure modes, and patterns for building resilient infrastructure.

  • Systems Programming and Performance Optimization

    Proficiency with systems-level programming, performance profiling, bottleneck identification, and optimization techniques for high-throughput systems.

  • Backend Programming Languages

    Strong proficiency in Python, Go, Java, or similar languages commonly used in infrastructure and backend systems development.

Preferred

  • Datadog Observability Platform

    Nice to have

    Hands-on experience implementing and optimizing Datadog for metrics collection, log aggregation, distributed tracing, and complex alerting in production environments.

  • Message Queue and Stream Processing

    Nice to have

    Experience with distributed message systems like Celery, Apache Kafka, RabbitMQ, or Redis; understanding of event-driven architecture patterns and async processing.

  • AI/ML Infrastructure

    Nice to have

    Exposure to AI infrastructure requirements including GPU cluster management, model serving frameworks, training pipeline orchestration, and cost optimization for ML workloads.

  • Multi-Region and High-Availability Architecture

    Nice to have

    Demonstrated success designing and operating multi-region deployments, cross-region failover mechanisms, and disaster recovery strategies for mission-critical systems.

  • Open Source Contribution

    Nice to have

    Active involvement in open source projects, particularly infrastructure or DevOps tools, demonstrating commitment to community and technical depth.

  • Financial or Enterprise SaaS Systems

    Nice to have

    Prior experience building or operating infrastructure for financial technology, enterprise software, or high-compliance environments where security and reliability are paramount.

  • Cost Optimization and FinOps

    Nice to have

    Track record of optimizing cloud infrastructure costs, implementing resource scheduling, and driving FinOps practices across engineering organizations.

Tech stack

Languages

PythonGoSQLBash/Shell

Frameworks

KubernetesCeleryDBOS

Databases

PostgreSQLRedisElasticsearch

Tools

TerraformDatadogDockerJenkins or GitHub ActionsArgoCDPrometheus

Other

AWS Cloud PlatformDoclingGit Version ControlHelm Package ManagerLinux System Administration

Compensation

Pay and benefits.

Base·CAD 160,000 – 220,000

Equity·Stock options

Benefits

  • Equity Compensation

    Competitive startup equity package providing long-term value participation as Zip scales its enterprise platform globally.

  • Comprehensive Health Coverage

    100% coverage options for health, vision, and dental insurance with multiple plan choices and family coverage options.

  • On-Campus Meals

    Catered breakfast, lunch, and dinner daily at Zip's offices to support employee wellness and team collaboration.

  • Flexible PTO Policy

    Unlimited flexible paid time off allowing employees to balance work and personal commitments without rigid accrual limits.

  • Wellness Benefits

    ClassPass membership providing access to thousands of fitness, wellness, and mental health activities and studios.

  • Commuter and Transportation Benefits

    Monthly commuter benefits supporting sustainable and convenient transportation options to and from the office.

  • Team Culture and Events

    Regular team building events, happy hours, and social gatherings fostering collaboration and company culture.

  • Remote Work Support

    Home office stipend for equipment and setup, plus phone and internet reimbursement to support distributed work quality.

  • Hybrid Work Flexibility

    Hybrid work model with 5 flexible remote days per quarter, allowing balanced in-office collaboration and remote productivity.

  • Family Support Benefits

    Paid parental leave and fertility benefits supporting employees at all life stages and family planning needs.

  • Employee Assistance Program (EAP)

    Comprehensive employee assistance program providing confidential counseling, mental health support, and personal resources.

  • AI Tool Access

    Unlimited AI token usage providing access to cutting-edge AI capabilities and tools for productivity and innovation.

Process

Interview steps.

  1. 01

    Initial Screening

    Brief conversation with Zip's recruiting team to discuss your background, infrastructure engineering experience, and interest in scaling enterprise procurement systems.

  2. 02

    Technical Interview - Systems Design

    Deep dive into distributed systems architecture and infrastructure design challenges. Expect discussions on multi-region deployment patterns, Kubernetes optimization, and scalability trade-offs relevant to processing $500+ billion in enterprise procurement spend.

  3. 03

    Technical Interview - Infrastructure Implementation

    Practical assessment of your infrastructure-as-code, automation, and hands-on experience with AWS, Kubernetes, or related technologies. May involve reviewing past projects or whiteboarding infrastructure solutions.

  4. 04

    Infrastructure Leadership and Collaboration

    Conversation with Core Infrastructure team members and potential collaborators focused on your approach to technical decision-making, cross-functional communication, and building resilient systems in distributed teams.

  5. 05

    Team and Cultural Fit Discussion

    Final conversation with engineering leadership exploring your alignment with Zip's values of ownership, open communication, underdog mindset, and commitment to driving innovation in enterprise software.

Full posting

Original listing.

About Zip

Zip is the AI platform for enterprise procurement — built for humans and agents working together. By orchestrating procurement across teams, tools, and suppliers with the help of AI agents, companies can secure the resources they need to innovate faster than ever before.

The world’s most influential enterprises trust Zip, including T-Mobile, OpenAI, AMD, Mars, Dollar Tree, and more. Together they’ve saved over $8 billion and processed over $500 billion in spend. Zip’s team includes product leaders from Apple, Airbnb, and Meta, as well as former procurement leaders from United Health, Sanofi, MGM Resorts, Discover, and NASA.

Backed by Adams Street, Alkeon, BOND, CRV, DST, Tiger Global, and Y Combinator, Zip has raised $371 million, most recently at a $2.2 billion valuation and has been recognized by Forbes Fintech 50, Fast Company's Most Innovative Companies, Inc. Best in Business, and LinkedIn Top Startups.

Your Role

We are looking for an experienced software engineer to join the Core Infrastructure team in Toronto, Canada. This role will become a key contributor to Zip’s multi-region expansion, and take ownership in one or more of the Core Infra components, such as Kubernetes platform, AI infrastructure and more.

You Will

  • Play a pivotal role in multi-region growth by building and operating global multi-cell architectures.

  • Assume end-to-end ownership of one or more infrastructure components, including Kubernetes platforms, AI Infra, observability, deployment pipeline and cost optimization.

  • Champion technical excellence by leading the design, implementation, and operational excellence of highly scalable systems.

  • Partner with engineering colleagues to accelerate and empower Zip’s product roadmap.

Qualifications

  • 4 years of experience in software engineering.

  • Strong track record of building and operating large scale platforms or infrastructure independently.

  • Excellent communication skills, both written and verbal.

  • Bachelor's degree or higher in Computer Science or a related technical field (e.g., Physics, Mathematics).

Nice to have

  • Hands-on expertise in building and operating large-scale systems, utilizing technologies like AWS, Datadog, Kubernetes, and open-source solutions such as Celery, Redis, Docling and DBOS.

  • Demonstrated effective collaboration with internal stakeholders across multiple time zones.

Perks & Benefits

At Zip, we’re committed to providing our employees with everything they need to do their best work.

  • 📈 Start-up equity

  • 🦷 100% health, vision & dental coverage options

  • 🍽️ Catered breakfast, lunch, & dinner

  • 🌴 Flexible PTO

  • 🏋️‍♀️ ClassPass membership

  • 🚍 Monthly commuter benefit

  • 🚠 Team building events & happy hours

  • 💻 Home office stipend

  • 🛜 Phone/internet reimbursement

  • 🏠 Hybrid model + 5 flexible remote days per quarter

  • 🍼 Paid parental leave

  • 🐣 Fertility benefits

  • 🧑‍🧑‍🧒‍🧒 Employee Assistance Program (EAP)

  • 🤖 Unlimited AI token usage

We're looking to hire Zipsters and that means hiring people who take ownership, communicate openly, have an underdog mindset, and are excited to increase the pace of innovation for every business in the world. We encourage all candidates to apply even if your experience doesn't exactly match up to our job description. We are committed to building a diverse and inclusive workspace where everyone (regardless of age, religion, ethnicity, gender, sexual orientation, and more) feels like they belong. We look forward to hearing from you!

Redirects to Zip's application page.

Other roles

More at Zip.

View all 9 roles