Lead Infrastructure Engineer
Senior · Full Time
Opens AtoB's application page
Role
What you'll do.
Lead Infrastructure Engineer at AtoB is a Senior/Staff-level role focused on designing and operating production Kubernetes infrastructure, modernizing CI/CD pipelines, and bridging infrastructure and data engineering disciplines. You'll take ownership of mission-critical systems supporting AtoB's fintech platform for transportation, with emphasis on zero-downtime architectures, multi-tenant infrastructure, and lakehouse data platforms while collaborating across technical and business teams.
Responsibilities
- Scope and Lead Complex Infrastructure Projects: Identify, scope, and lead large-scale, often ambiguous technical initiatives that establish foundational infrastructure for early-stage products, enabling iterative evolution and scaling with clear success metrics and phased delivery
- Kubernetes Infrastructure Operations: Own full operational lifecycle of production Kubernetes (EKS) infrastructure managed through Porter, including cluster provisioning, node management, workload scheduling, and performance optimization
- Enforce Kubernetes Best Practices: Apply industry best practices across resource management, health checks and probes, horizontal/vertical autoscaling, pod disruption budgets, network policies, and security controls to ensure platform reliability
- Design Zero-Downtime Release Architecture: Tune deployments and cluster configurations to achieve zero-downtime releases through deployment strategies such as rolling updates, canary deployments, blue-green deployments, and graceful shutdown handling
- Build and Improve CI/CD Pipelines: Design, implement, and continuously improve continuous integration and continuous deployment pipelines that enable fast, reliable, and safe software delivery with comprehensive testing, security scanning, and automated rollback capabilities
- Develop Observability and Incident Response: Contribute to comprehensive observability infrastructure including metrics collection, distributed tracing, centralized logging, and alerting systems; establish incident response procedures and post-incident review processes
- Bridge Infrastructure and Data Engineering: Work at the intersection of platform engineering and data engineering, designing infrastructure to support modern data platforms, lakehouse architectures, and analytics workloads while maintaining operational excellence
- Cross-functional Stakeholder Collaboration: Partner directly with BizOps, finance, business teams, and engineering stakeholders to translate business requirements into technical solutions, provide cost analysis, and drive data-informed infrastructure decisions
- Database and Query Optimization: Write and optimize SQL queries for performance and cost efficiency, analyze query execution plans, establish indexing strategies, and mentor team members on SQL best practices
- Implement Modern Data Engineering Patterns: Research, evaluate, and implement contemporary data engineering approaches including lakehouse architectures, Apache Iceberg, Change Data Capture patterns, and transformation frameworks to modernize data infrastructure
Qualifications
What we look for.
Technical
Kubernetes (EKS) Production Expertise
Minimum 8-12 years of hands-on experience operating Kubernetes in production, with deep knowledge of Amazon EKS, resource management, autoscaling policies, pod disruption budgets, and failure recovery patterns
CI/CD Pipeline Design and Implementation
Proven ability to architect and maintain enterprise-scale CI/CD systems supporting fast, safe deployments with automated testing, security scanning, and deployment verification
Advanced SQL and Query Optimization
Strong SQL expertise including query optimization, execution plan analysis, indexing strategies, and cost management in cloud data warehouses
Infrastructure-as-Code (IaC)
Proficiency with IaC tools such as Terraform or CloudFormation for reproducible, version-controlled infrastructure management
AWS Cloud Platform
Deep knowledge of AWS services including EC2, ECS/EKS, S3, RDS, VPC networking, IAM, and multi-region architectures
Modern Data Engineering Concepts
Understanding of contemporary data engineering including lakehouse architectures, Apache Iceberg, Change Data Capture (CDC), data transformation frameworks (dbt), and data quality patterns
Observability and Monitoring
Experience designing comprehensive observability solutions including metrics, logging, distributed tracing, and alerting for production systems
Distributed Systems Concepts
Understanding of distributed systems principles including consistency models, replication strategies, failover mechanisms, and multi-region deployment patterns
Education
Bachelor's Degree in Computer Science or Related Field
Formal education in computer science, software engineering, or equivalent field providing foundational knowledge of algorithms, data structures, and system design principles
Experience
Senior Infrastructure Engineering (8-12 years)
Substantial experience in infrastructure engineering roles, with demonstrated progression to senior levels, leading technical initiatives and owning production systems at scale
Production Kubernetes Operations
Deep hands-on experience operating Kubernetes clusters in production environments, managing hundreds or thousands of pods, handling cluster upgrades, and optimizing performance
Large-Scale Platform Engineering
Track record of designing and building infrastructure platforms that serve dozens or hundreds of engineers, with focus on developer experience and operational efficiency
Mission-Critical System Ownership
Demonstrated ability to take ownership of complex systems from design through production operation, including on-call responsibilities and incident response leadership
Cross-functional Technical Leadership
Experience leading technical initiatives involving coordination with multiple teams including data engineering, platform teams, and non-technical stakeholders
Skills
Required
Kubernetes Production Operations
8-12 years of deep expertise operating Kubernetes in production environments, including EKS, with mastery of resource management, health checks, autoscaling, and pod disruption budgets
Zero-Downtime Architecture Design
Proven ability to design, tune, and operate systems achieving zero-downtime releases through graceful rollout strategies, canary deployments, and advanced Kubernetes patterns
CI/CD Pipeline Architecture
Strong experience designing, building, and maintaining large-scale CI/CD pipelines ensuring fast, reliable, and safe software delivery with automated testing and deployment strategies
SQL Performance Optimization
Advanced SQL skills with demonstrated ability to write optimized queries, analyze execution plans, and reduce query costs in production data environments
Modern Data Engineering Fundamentals
Demonstrated hunger to learn and apply contemporary data engineering patterns including lakehouse architectures, Apache Iceberg, Change Data Capture (CDC), and transformation frameworks like dbt
Cross-functional Leadership
Excellent collaboration skills with ability to communicate technical concepts to non-engineering stakeholders including BizOps, finance, and business teams while translating business requirements into technical solutions
End-to-End Ownership Mindset
Demonstrated track record of owning systems from architectural design through implementation, deployment, operation, and optimization with accountability for reliability and performance
Preferred
Multi-region Architecture
Nice to haveExperience designing and implementing cross-region failover, disaster recovery strategies, and data replication patterns for high-availability systems
Multi-tenant Infrastructure
Nice to haveExpertise in designing and operating multi-tenant systems with tenant isolation, fair resource sharing, and per-tenant observability and scaling
Apache Iceberg and Lakehouse Catalogs
Nice to haveHands-on experience with Apache Iceberg, AWS Glue, or other modern lakehouse catalog systems for building scalable data platforms
EKS and Porter Experience
Nice to havePrior experience managing EKS clusters or using Porter orchestrator for Kubernetes management at scale
Modern Data Stack
Nice to haveFamiliarity with contemporary data engineering tools including dbt, Snowflake, Debezium, Estuary, and similar components in the modern data stack
Fintech and Regulated Environment Compliance
Nice to haveExperience building and operating systems in regulated industries with knowledge of compliance requirements like PCI DSS, SOC 2, or other financial services standards
Platform Engineering Leadership
Nice to haveTrack record of building internal developer platforms, developer experience tools, or infrastructure abstractions that enable engineering teams
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 180,000 – 280,000
Equity·Stock options
Benefits
Equity Compensation
Meaningful equity stake in venture-backed fintech startup positioned as next billion-dollar company, with exposure to significant upside and alignment with company success
Comprehensive Health Coverage
Medical, dental, and vision insurance with company contributions, designed to support employee and family health needs
Professional Development
Learning budget and opportunities to develop expertise in cutting-edge technologies including modern data engineering, advanced Kubernetes patterns, and fintech domain knowledge
Remote-Friendly Work Environment
Flexibility to work across geographies with headquarters in San Francisco; ability to collaborate asynchronously with distributed team
High-Impact Role
Opportunity to shape greenfield data platform on modern lakehouse foundations with small team structure enabling significant influence on technical direction and product strategy
Industry Expertise Access
Work alongside engineers from leading companies (Google, Uber, Meta, Shopify, Stripe, Chime) providing mentorship and access to world-class infrastructure expertise
Meaningful Mission
Contribute to modernizing payments infrastructure for transportation and logistics industry, directly supporting hard-working trucking and logistics operators
Full posting
Original listing.
Our mission
The trucking and logistics industry provides the backbone of the economy. But the payments infrastructure on which it runs is broken. For the hard-working men and women of this sector, the existing suite of payment tools is outdated, difficult to use, prone to fraud, and saddled with shady fee structures. The incumbent players in this space often overlook the economic and practical needs of this user base.
We're changing that. AtoB is building Stripe for Transportation — modernizing the payments infrastructure for trucking and logistics. Supply chains rely on the timely movement of capital to function efficiently. Our end game is a world in which that capital movement occurs fairly, smoothly, and without delay. As we pursue that end game, we aim to center our customers in every way — offering them world-class customer experience and building products that work with and around the unique constraints of their daily lives. We build for fleet managers in the office and drivers on the road. We strive for products that are efficient, satisfying, and useful. Our customers enable our modern economy — they deserve it.
Our history and background
Our founding team has backgrounds in payments, working on autonomous vehicles at Cruise Automation, leading ops and growth for Uber, and building apps that were featured on the Apple app store. We have staff and senior engineers from Google, Uber, Meta, Shopify, Stripe, Chime, and other leading technology companies.
We have raised $125 million+ from investors such as General Catalyst, Elad Gil, Bloomberg Beta, Y Combinator, XYZ; founders and CEOs of companies such as Google (Eric Schmidt), Salesforce (Marc Benioff), Coinbase (Brian Armstrong), DoorDash (Tony Xu), Instacart, Gusto; strategic investors like Mastercard, Flexport and Samsara.
We were named to Forbes annual Next Billion-Dollar Startup List, and have just recently been selected to join the World Economic Forum as a Global Innovator.
What You'll Do
Infrastructure & Kubernetes
Scope and lead large, often ambiguous technical projects, laying the groundwork for early-stage products to iteratively evolve and scale
Own and operate our Kubernetes (EKS) infrastructure, managed via the Porter orchestrator
Apply and enforce Kubernetes best practices: resource management, health checks, autoscaling, pod disruption budgets, and graceful rollout strategies
Tune deployments and cluster configuration for zero-downtime releases and high availability
Design and improve CI/CD pipelines for fast, reliable, and safe delivery
Contribute to observability, alerting, and incident response practices across the platform
Must-Have
8-12 years Deep expertise in Kubernetes — production operations, best practices, and performance tuning for zero-downtime architectures
Strong CI/CD experience — designing, building, and maintaining pipelines at scale
Demonstrated hunger to learn modern data engineering — lakehouse architectures, Iceberg, CDC, transformation frameworks (e.g., dbt or coalesce.io or others)
Strong SQL skills and the ability to optimize queries for performance and cost
Excellent cross-functional collaboration skills — comfortable working directly with BizOps and other non-engineering stakeholders
Ownership mindset — able to take systems from design through operation end to end
Nice-to-Have
Cross-region architecture experience (multi-region failover, DR, data replication)
Cross-tenant / multi-tenant architecture expertise
Experience with Apache Iceberg, AWS Glue, or other lakehouse catalogs
Experience with EKS, Porter, or similar Kubernetes management platforms
Familiarity with modern data stack tooling: dbt, Snowflake, CDC tools (Debezium, Estuary, etc.)
Experience in fintech or other regulated environments (PCI DSS, SOC 2)
Why You Should Join
Shape a greenfield data platform built on modern lakehouse foundations
Real ownership across infrastructure and data — small team, high impact
Work at the intersection of platform engineering and data engineering, two of the fastest-growing disciplines in the industry
Redirects to AtoB's application page.
Other roles