Zip

Engineering Manager, Online Storage - SF

Zip1 weeks ago
Location

San Francisco

Type

Full Time

Salary

USD 250,000 – 300,000

Level

Manager

Role

Engineering Manager

Posted

Jul 16, 2026

Full TimeManager

The role

Summary

Lead Zip's foundational Online Storage infrastructure team as an Engineering Manager, overseeing critical systems including Aurora RDS, DynamoDB, OpenSearch, and permissions frameworks. This role combines hands-on technical leadership with team growth, scaling the team from 6+ to 12+ engineers while driving multi-cell scalability initiatives and ensuring reliable storage infrastructure serving enterprise procurement at scale.

What you'll do

Team Leadership and Growth: Build and scale the Online Storage engineering team from current size to 12+ engineers within 6 months. Drive hiring strategy, conduct technical interviews, and establish team culture. Provide career coaching, performance management, and professional development opportunities for direct reports. Establish hiring criteria aligned with infrastructure engineering excellence and promote diversity in technical hiring.
Multi-Cell Scalability Architecture: Lead the strategic multi-cell scalability program enabling data isolation, regional redistribution, and intelligent cell provisioning. Collaborate with Core Infrastructure team to design and implement systems supporting enterprise-scale deployment across multiple geographic regions. Define technical roadmap for achieving data residency requirements and reducing single points of failure.
Reliability and Disaster Recovery Operations: Own comprehensive reliability engineering for all core storage infrastructure including Aurora RDS, ElastiCache, MemoryDB, DynamoDB, S3, and OpenSearch. Establish and enforce Recovery Point Objective (RPO) and Recovery Time Objective (RTO) targets. Design and maintain disaster recovery procedures, backup strategies, and failover mechanisms. Conduct regular disaster recovery drills and post-mortem analysis.
Unified Search Platform Technical Direction: Set architectural vision and technical direction for Zip's unified search platform consolidation. Define search infrastructure standards, evaluate technology choices (OpenSearch, Elasticsearch, alternatives), and establish performance benchmarks. Coordinate with product teams on search capabilities while maintaining platform stability and operational efficiency.
Permissions System Consolidation: Lead the consolidation and modernization of Zip's permissions framework ensuring every user sees exactly the data they're authorized to access. Define technical approach for unified permissions system, establish auditing and compliance standards, and ensure scalability across multi-tenant enterprise deployments.
Internal Platform Partnership and Guardrails: Operate as strategic internal platform partner for product teams. Triage and prioritize infrastructure requests balancing product velocity against platform stability and reliability. Establish guardrails, best practices, and SLOs for infrastructure consumption. Facilitate communication between product teams and infrastructure engineering to align on capacity planning and feature roadmaps.
Operational Excellence and On-Call Management: Drive operational excellence across the data layer including on-call health, incident response processes, and observability infrastructure. Establish on-call rotations with clear escalation procedures and on-call compensation. Implement comprehensive monitoring, alerting, and logging for all storage systems. Lead blameless post-incident reviews and establish continuous improvement cycles.

What we look for

Technical

Production Data Infrastructure at ScaleDemonstrated hands-on experience building, deploying, or operating production data infrastructure serving enterprise workloads. This includes relational database systems, distributed caching layers, object storage, or search infrastructure operating at significant scale (millions of queries per day, terabytes of data).
Relational Database SystemsStrong working knowledge of relational database operations, optimization, and maintenance. Experience with Amazon Aurora RDS, PostgreSQL, or MySQL in production environments, including index design, query optimization, replication, backup strategies, and disaster recovery procedures.
Distributed Storage and CachingHands-on experience with distributed storage solutions and caching layers. Proficiency with technologies such as DynamoDB, Redis, ElastiCache, MemoryDB, or similar distributed data systems. Understanding of consistency models, partitioning strategies, and trade-offs between different storage architectures.
Search InfrastructureExperience building, deploying, or maintaining search infrastructure at scale. Familiarity with OpenSearch, Elasticsearch, or similar distributed search platforms. Understanding of indexing strategies, query performance tuning, cluster management, and search relevance optimization.
Infrastructure as Code and DevOpsProficiency with Infrastructure as Code tools (Terraform, CloudFormation, or equivalent) and containerization technologies (Docker, Kubernetes). Experience with AWS services ecosystem including RDS, DynamoDB, S3, ElastiCache, and monitoring/observability tools.

Education

Computer Science Degree or Related FieldBachelor's degree or higher in Computer Science, Computer Engineering, Software Engineering, or related technical discipline. Equivalent professional experience may be considered for candidates with exceptional track records in infrastructure engineering roles.

Experience

Software EngineeringMinimum 6+ years of professional software engineering experience delivering production systems. Background should include significant work on backend systems, infrastructure, or platform engineering with exposure to distributed systems concepts and large-scale system design.
Engineering ManagementMinimum 2+ years of direct experience managing engineering teams of 6 or more people. Track record of successful hiring, team development, performance management, and scaling teams while maintaining quality and culture. Demonstrated ability to balance technical contribution with management responsibilities.
Infrastructure Engineering LeadershipProven experience leading infrastructure or platform engineering initiatives. Background should include mentoring junior engineers, establishing best practices, and driving technical standardization. Experience with on-call management, incident response, and operational excellence initiatives.

Skills

Required skills

Cloud Infrastructure ArchitectureDesign and implementation of scalable cloud infrastructure on AWS, GCP, or Azure. Experience architecting multi-region deployments, implementing disaster recovery, and optimizing cloud costs. Strong understanding of cloud-native patterns and best practices for enterprise applications.
Distributed Systems DesignDeep understanding of distributed systems principles including consistency, availability, partition tolerance, CAP theorem, and eventual consistency models. Ability to design systems that scale horizontally and handle failures gracefully.
Data Modeling and Database DesignExpert-level ability to design efficient data models for relational and non-relational databases. Experience with schema design, indexing strategies, query optimization, and understanding trade-offs between different storage paradigms.
Team Leadership and MentorshipProven ability to lead and develop high-performing engineering teams. Skills include hiring, onboarding, performance management, career development, and creating psychological safety. Experience establishing team processes and scaling teams while maintaining culture.
Incident Response and Post-Mortem LeadershipExperience managing critical infrastructure incidents and driving blameless post-incident review processes. Ability to triage issues, communicate during crises, and implement systemic improvements. Understanding of SLOs, error budgets, and reliability targets.
Observability and MonitoringExpertise in implementing comprehensive observability (logs, metrics, traces) for complex infrastructure systems. Experience with monitoring platforms like Datadog, New Relic, Prometheus, or ELK stack. Ability to establish meaningful alerts and drive visibility into system behavior.

Nice to have

SaaS Platform OperationsExperience operating SaaS platforms at scale with uptime requirements and compliance considerations. Knowledge of multi-tenancy architecture, data isolation, and security considerations specific to SaaS businesses.
FinTech Infrastructure ExperienceBackground working on infrastructure for financial technology companies. Understanding of compliance requirements (SOC 2, PCI-DSS), audit trails, regulatory considerations, and high-reliability expectations in fintech environments.
Zero-to-One Project LeadershipTrack record of bootstrapping and shipping large-scale infrastructure projects from conceptualization through production deployment. Experience defining requirements, building minimum viable infrastructure, and iterating to scale.
Startup Engineering CultureExperience in early-stage or high-growth startup environments with diverse technical responsibilities. Comfort with ambiguity, ability to wear multiple hats, and success in fast-moving organizations with limited resources.
Hiring at ScaleDemonstrated success in building engineering recruitment processes, sourcing talent, and scaling teams rapidly. Experience identifying and attracting senior infrastructure engineers and building diverse teams.
Permissions and Authorization SystemsExperience implementing or maintaining complex authorization and permissions systems. Familiarity with role-based access control (RBAC), attribute-based access control (ABAC), or similar frameworks. Understanding of audit trails and compliance logging.

Compensation & benefits

Salary

USD 250,000 – 300,000 (annual)

Stock options

Available


Apply for this position

You'll be redirected to the company's application page


Zip

Zip

View all jobs

Zip provides an intake-to-pay platform designed to streamline procurement processes, automate approvals, and improve visibility and control for organizations.

San Francisco, CA, USAFounded 2019ziphq.com

Tech Stack

Languages
PythonTypeScript/JavaScriptGoSQL
Frameworks
Spring Boot / Spring FrameworkNode.jsFastAPI / Django
Databases
Amazon Aurora RDSDynamoDBAmazon S3OpenSearchElastiCache/MemoryDB
Tools
TerraformKubernetesDataDog or New RelicPagerDutyGitHubAWS CloudFormation
Other
AWS EcosystemMulti-Region ArchitectureDisaster Recovery and Business ContinuityFinTech ComplianceAgile and Scrum Methodologies
Apply Now