Lead Infrastructure Engineer

Senior · Full Time

MontrealUSD 180k – 280k2w ago
Apply for this role

Opens AtoB's application page

Role

What you'll do.

Lead Infrastructure Engineer at AtoB is a Senior/Staff-level role focused on designing and operating production Kubernetes infrastructure, modernizing CI/CD pipelines, and bridging infrastructure and data engineering disciplines. You'll take ownership of mission-critical systems supporting AtoB's fintech platform for transportation, with emphasis on zero-downtime architectures, multi-tenant infrastructure, and lakehouse data platforms while collaborating across technical and business teams.

Responsibilities

  • Scope and Lead Complex Infrastructure Projects: Identify, scope, and lead large-scale, often ambiguous technical initiatives that establish foundational infrastructure for early-stage products, enabling iterative evolution and scaling with clear success metrics and phased delivery
  • Kubernetes Infrastructure Operations: Own full operational lifecycle of production Kubernetes (EKS) infrastructure managed through Porter, including cluster provisioning, node management, workload scheduling, and performance optimization
  • Enforce Kubernetes Best Practices: Apply industry best practices across resource management, health checks and probes, horizontal/vertical autoscaling, pod disruption budgets, network policies, and security controls to ensure platform reliability
  • Design Zero-Downtime Release Architecture: Tune deployments and cluster configurations to achieve zero-downtime releases through deployment strategies such as rolling updates, canary deployments, blue-green deployments, and graceful shutdown handling
  • Build and Improve CI/CD Pipelines: Design, implement, and continuously improve continuous integration and continuous deployment pipelines that enable fast, reliable, and safe software delivery with comprehensive testing, security scanning, and automated rollback capabilities
  • Develop Observability and Incident Response: Contribute to comprehensive observability infrastructure including metrics collection, distributed tracing, centralized logging, and alerting systems; establish incident response procedures and post-incident review processes
  • Bridge Infrastructure and Data Engineering: Work at the intersection of platform engineering and data engineering, designing infrastructure to support modern data platforms, lakehouse architectures, and analytics workloads while maintaining operational excellence
  • Cross-functional Stakeholder Collaboration: Partner directly with BizOps, finance, business teams, and engineering stakeholders to translate business requirements into technical solutions, provide cost analysis, and drive data-informed infrastructure decisions
  • Database and Query Optimization: Write and optimize SQL queries for performance and cost efficiency, analyze query execution plans, establish indexing strategies, and mentor team members on SQL best practices
  • Implement Modern Data Engineering Patterns: Research, evaluate, and implement contemporary data engineering approaches including lakehouse architectures, Apache Iceberg, Change Data Capture patterns, and transformation frameworks to modernize data infrastructure

Qualifications

What we look for.

Technical

  • Kubernetes (EKS) Production Expertise

    Minimum 8-12 years of hands-on experience operating Kubernetes in production, with deep knowledge of Amazon EKS, resource management, autoscaling policies, pod disruption budgets, and failure recovery patterns

  • CI/CD Pipeline Design and Implementation

    Proven ability to architect and maintain enterprise-scale CI/CD systems supporting fast, safe deployments with automated testing, security scanning, and deployment verification

  • Advanced SQL and Query Optimization

    Strong SQL expertise including query optimization, execution plan analysis, indexing strategies, and cost management in cloud data warehouses

  • Infrastructure-as-Code (IaC)

    Proficiency with IaC tools such as Terraform or CloudFormation for reproducible, version-controlled infrastructure management

  • AWS Cloud Platform

    Deep knowledge of AWS services including EC2, ECS/EKS, S3, RDS, VPC networking, IAM, and multi-region architectures

  • Modern Data Engineering Concepts

    Understanding of contemporary data engineering including lakehouse architectures, Apache Iceberg, Change Data Capture (CDC), data transformation frameworks (dbt), and data quality patterns

  • Observability and Monitoring

    Experience designing comprehensive observability solutions including metrics, logging, distributed tracing, and alerting for production systems

  • Distributed Systems Concepts

    Understanding of distributed systems principles including consistency models, replication strategies, failover mechanisms, and multi-region deployment patterns

Education

  • Bachelor's Degree in Computer Science or Related Field

    Formal education in computer science, software engineering, or equivalent field providing foundational knowledge of algorithms, data structures, and system design principles

Experience

  • Senior Infrastructure Engineering (8-12 years)

    Substantial experience in infrastructure engineering roles, with demonstrated progression to senior levels, leading technical initiatives and owning production systems at scale

  • Production Kubernetes Operations

    Deep hands-on experience operating Kubernetes clusters in production environments, managing hundreds or thousands of pods, handling cluster upgrades, and optimizing performance

  • Large-Scale Platform Engineering

    Track record of designing and building infrastructure platforms that serve dozens or hundreds of engineers, with focus on developer experience and operational efficiency

  • Mission-Critical System Ownership

    Demonstrated ability to take ownership of complex systems from design through production operation, including on-call responsibilities and incident response leadership

  • Cross-functional Technical Leadership

    Experience leading technical initiatives involving coordination with multiple teams including data engineering, platform teams, and non-technical stakeholders

Skills

Required

  • Kubernetes Production Operations

    8-12 years of deep expertise operating Kubernetes in production environments, including EKS, with mastery of resource management, health checks, autoscaling, and pod disruption budgets

  • Zero-Downtime Architecture Design

    Proven ability to design, tune, and operate systems achieving zero-downtime releases through graceful rollout strategies, canary deployments, and advanced Kubernetes patterns

  • CI/CD Pipeline Architecture

    Strong experience designing, building, and maintaining large-scale CI/CD pipelines ensuring fast, reliable, and safe software delivery with automated testing and deployment strategies

  • SQL Performance Optimization

    Advanced SQL skills with demonstrated ability to write optimized queries, analyze execution plans, and reduce query costs in production data environments

  • Modern Data Engineering Fundamentals

    Demonstrated hunger to learn and apply contemporary data engineering patterns including lakehouse architectures, Apache Iceberg, Change Data Capture (CDC), and transformation frameworks like dbt

  • Cross-functional Leadership

    Excellent collaboration skills with ability to communicate technical concepts to non-engineering stakeholders including BizOps, finance, and business teams while translating business requirements into technical solutions

  • End-to-End Ownership Mindset

    Demonstrated track record of owning systems from architectural design through implementation, deployment, operation, and optimization with accountability for reliability and performance

Preferred

  • Multi-region Architecture

    Nice to have

    Experience designing and implementing cross-region failover, disaster recovery strategies, and data replication patterns for high-availability systems

  • Multi-tenant Infrastructure

    Nice to have

    Expertise in designing and operating multi-tenant systems with tenant isolation, fair resource sharing, and per-tenant observability and scaling

  • Apache Iceberg and Lakehouse Catalogs

    Nice to have

    Hands-on experience with Apache Iceberg, AWS Glue, or other modern lakehouse catalog systems for building scalable data platforms

  • EKS and Porter Experience

    Nice to have

    Prior experience managing EKS clusters or using Porter orchestrator for Kubernetes management at scale

  • Modern Data Stack

    Nice to have

    Familiarity with contemporary data engineering tools including dbt, Snowflake, Debezium, Estuary, and similar components in the modern data stack

  • Fintech and Regulated Environment Compliance

    Nice to have

    Experience building and operating systems in regulated industries with knowledge of compliance requirements like PCI DSS, SOC 2, or other financial services standards

  • Platform Engineering Leadership

    Nice to have

    Track record of building internal developer platforms, developer experience tools, or infrastructure abstractions that enable engineering teams

Tech stack

Languages

SQLPythonYAMLBash/Shell

Frameworks

Kubernetes (EKS)dbt (data build tool)Apache Iceberg

Databases

SnowflakeApache Iceberg / LakehousePostgreSQL

Tools

PorterAWS GlueCI/CD Platforms (GitHub Actions, GitLab CI, or similar)Debezium / EstuaryObservability Stack (Datadog, New Relic, Prometheus, or similar)

Other

AWS (EC2, S3, RDS, networking)Infrastructure-as-Code (IaC)Container Technology (Docker)Multi-region and Disaster Recovery

Compensation

Pay and benefits.

Base·USD 180,000 – 280,000

Equity·Stock options

Benefits

  • Equity Compensation

    Meaningful equity stake in venture-backed fintech startup positioned as next billion-dollar company, with exposure to significant upside and alignment with company success

  • Comprehensive Health Coverage

    Medical, dental, and vision insurance with company contributions, designed to support employee and family health needs

  • Professional Development

    Learning budget and opportunities to develop expertise in cutting-edge technologies including modern data engineering, advanced Kubernetes patterns, and fintech domain knowledge

  • Remote-Friendly Work Environment

    Flexibility to work across geographies with headquarters in San Francisco; ability to collaborate asynchronously with distributed team

  • High-Impact Role

    Opportunity to shape greenfield data platform on modern lakehouse foundations with small team structure enabling significant influence on technical direction and product strategy

  • Industry Expertise Access

    Work alongside engineers from leading companies (Google, Uber, Meta, Shopify, Stripe, Chime) providing mentorship and access to world-class infrastructure expertise

  • Meaningful Mission

    Contribute to modernizing payments infrastructure for transportation and logistics industry, directly supporting hard-working trucking and logistics operators

Full posting

Original listing.

Our mission

The trucking and logistics industry provides the backbone of the economy. But the payments infrastructure on which it runs is broken. For the hard-working men and women of this sector, the existing suite of payment tools is outdated, difficult to use, prone to fraud, and saddled with shady fee structures. The incumbent players in this space often overlook the economic and practical needs of this user base.


We're changing that. AtoB is building Stripe for Transportation — modernizing the payments infrastructure for trucking and logistics. Supply chains rely on the timely movement of capital to function efficiently. Our end game is a world in which that capital movement occurs fairly, smoothly, and without delay. As we pursue that end game, we aim to center our customers in every way — offering them world-class customer experience and building products that work with and around the unique constraints of their daily lives. We build for fleet managers in the office and drivers on the road. We strive for products that are efficient, satisfying, and useful. Our customers enable our modern economy — they deserve it.


Our history and background

Our founding team has backgrounds in payments, working on autonomous vehicles at Cruise Automation, leading ops and growth for Uber, and building apps that were featured on the Apple app store. We have staff and senior engineers from Google, Uber, Meta, Shopify, Stripe, Chime, and other leading technology companies. 


We have raised $125 million+ from investors such as General Catalyst, Elad Gil, Bloomberg Beta, Y Combinator, XYZ; founders and CEOs of companies such as Google (Eric Schmidt), Salesforce (Marc Benioff), Coinbase (Brian Armstrong), DoorDash (Tony Xu), Instacart, Gusto; strategic investors like Mastercard, Flexport and Samsara.


We were named to Forbes annual Next Billion-Dollar Startup List, and have just recently been selected to join the World Economic Forum as a Global Innovator.

What You'll Do

Infrastructure & Kubernetes

  • Scope and lead large, often ambiguous technical projects, laying the groundwork for early-stage products to iteratively evolve and scale

  • Own and operate our Kubernetes (EKS) infrastructure, managed via the Porter orchestrator

  • Apply and enforce Kubernetes best practices: resource management, health checks, autoscaling, pod disruption budgets, and graceful rollout strategies

  • Tune deployments and cluster configuration for zero-downtime releases and high availability

  • Design and improve CI/CD pipelines for fast, reliable, and safe delivery

  • Contribute to observability, alerting, and incident response practices across the platform

Must-Have

  • 8-12 years Deep expertise in Kubernetes — production operations, best practices, and performance tuning for zero-downtime architectures

  • Strong CI/CD experience — designing, building, and maintaining pipelines at scale

  • Demonstrated hunger to learn modern data engineering — lakehouse architectures, Iceberg, CDC, transformation frameworks (e.g., dbt or coalesce.io or others)

  • Strong SQL skills and the ability to optimize queries for performance and cost

  • Excellent cross-functional collaboration skills — comfortable working directly with BizOps and other non-engineering stakeholders

  • Ownership mindset — able to take systems from design through operation end to end

Nice-to-Have

  • Cross-region architecture experience (multi-region failover, DR, data replication)

  • Cross-tenant / multi-tenant architecture expertise

  • Experience with Apache Iceberg, AWS Glue, or other lakehouse catalogs

  • Experience with EKS, Porter, or similar Kubernetes management platforms

  • Familiarity with modern data stack tooling: dbt, Snowflake, CDC tools (Debezium, Estuary, etc.)

  • Experience in fintech or other regulated environments (PCI DSS, SOC 2)

Why You Should Join

  • Shape a greenfield data platform built on modern lakehouse foundations

  • Real ownership across infrastructure and data — small team, high impact

  • Work at the intersection of platform engineering and data engineering, two of the fastest-growing disciplines in the industry

Redirects to AtoB's application page.

Other roles

More at AtoB.