TLM, Production Engineering

Engineering Manager · Lead · Full Time

New York, NY (HQ)USD 168k – 325k1d ago
Apply for this role

Opens Ramp's application page

Role

What you'll do.

Team Lead Manager of Production Engineering at Ramp, a Series C fintech company automating spend management for 70,000+ companies. Lead Ramp's infrastructure ownership layer ensuring systems are faster, more reliable, and scalable across compute, storage, messaging, and observability. This role requires shipping high-quality architectures for critical financial systems, leading cross-team initiatives, and driving infrastructure modernization including cellular architecture migrations and AI-native platform capabilities.

Responsibilities

  • Build and Operate Critical Infrastructure: Architect, build, and operate mission-critical infrastructure across Ramp's compute, storage, messaging, and observability stack handling real financial transactions at scale. Own end-to-end systems reliability, performance optimization, and capacity planning for distributed systems processing over $200B in annualized spend.
  • Drive Architectural Innovation and Technical Strategy: Lead architectural transformations across the organization, not just flag problems. Propose infrastructure solutions to reliability and scalability issues, coordinate with cross-functional engineering teams, and maintain accountability through completion. Establish standards and patterns that become the foundation for company-wide infrastructure practices.
  • Lead Cellular Architecture Migration: Take hands-on ownership of Ramp's most consequential infrastructure shift: transitioning to cellular architecture. Enable global scalability, international market expansion, highly regulated environments (FedRAMP compliance), and enterprise-grade SLAs. Design and implement systems supporting multi-tenant isolation and regulatory compliance.
  • Partner with Product Teams on AI Infrastructure: Proactively collaborate with product engineering teams to enable AI-powered products. Define golden paths and reusable infrastructure patterns for AI workloads, anticipate emerging AI infrastructure challenges before they become production blockers, and establish best practices for AI-native platform capabilities.
  • Build Developer Tooling and Self-Service Platforms: Design and implement internal developer platforms that empower engineering teams to self-serve infrastructure questions regarding cost attribution, performance monitoring, and reliability metrics. Reduce friction in deployment pipelines, observability systems, and infrastructure provisioning.
  • Lead On-Call Operations and Incident Response: Participate in on-call rotation and drive incident response leadership across critical financial systems. Transform every incident into a learning opportunity by identifying root causes and implementing systemic improvements to eliminate entire classes of problems, not just symptoms.
  • Champion Company-Wide Reliability and Scalability: Proactively initiate cross-team architectural reviews and lead company-wide reliability and scalability initiatives. Take full ownership from problem identification through solution proposal, stakeholder alignment, and delivery, ensuring Production Engineering drives technical outcomes beyond team boundaries.
  • Partner at Design Phase with Product Engineering: Embed infrastructure expertise early in product development by reviewing architectures, establishing golden paths, and making it easy for teams to build correctly the first time. Prevent technical debt accumulation through proactive collaboration and standards enforcement.

Qualifications

What we look for.

Technical

  • Distributed Systems Architecture

    Deep expertise in designing and implementing distributed systems patterns. Proficiency in consensus algorithms, data replication, partitioning strategies, and handling failure scenarios across multiple availability zones and regions.

  • Container Orchestration and Deployment

    Production-level experience with Kubernetes or equivalent container orchestration platforms. Expertise in container design, networking policies, resource management, and deployment automation pipelines.

  • Database and Storage Systems

    Hands-on experience with both SQL and NoSQL databases at scale. Understanding of database performance tuning, replication strategies, backup and recovery, and operational considerations for production databases.

  • Infrastructure as Code

    Proficiency with Terraform, CloudFormation, or equivalent IaC tools. Experience designing repeatable, testable, and maintainable infrastructure code following software engineering best practices.

  • Observability and Monitoring

    Expert-level knowledge of observability platforms and practices. Proficiency with tools like Prometheus, Grafana, DataDog, or equivalent. Experience designing SLOs, implementing alerting strategies, and building dashboards for system health.

  • Event-Driven Architecture

    Experience designing and implementing event-driven systems using message queues and stream processing. Familiarity with event sourcing patterns and workflow orchestration principles.

  • Software Engineering Fundamentals

    Strong software engineering practices including writing clean, well-tested, production-ready code. Proficiency in multiple programming languages. Deep understanding of testing strategies, code review practices, and refactoring techniques.

Education

  • Computer Science or Related Field

    Bachelor's degree in Computer Science, Software Engineering, or equivalent practical experience. Advanced degrees or specialized certifications in distributed systems, cloud architecture, or site reliability engineering are advantageous but not required.

Experience

  • Infrastructure Leadership at Scale

    3+ years of software engineering management shipping high-quality architectures for critical systems. Demonstrated track record leading technical projects end-to-end with cross-team coordination and architectural ownership. Experience building systems handling millions of transactions and serving at scale.

  • Distributed Systems Production Experience

    Hands-on experience architecting, deploying, and operating distributed systems at production scale. Deep understanding of distributed system challenges including eventual consistency, fault tolerance, consensus algorithms, and horizontal scaling patterns.

  • Cloud Infrastructure and Deployment

    Extensive hands-on experience with major cloud providers, preferably AWS. Proficiency in container orchestration, networking, load balancing, auto-scaling, and infrastructure-as-code tooling. Experience designing for high availability, disaster recovery, and multi-region deployments.

  • Observability and Reliability Engineering

    Strong familiarity with observability practices including SLO definition, error budget management, alerting strategies, and dashboard design. Experience establishing reliability standards, defining metrics, and building observability infrastructure for distributed systems.

  • Mission-Critical Systems Ownership

    Proven experience owning infrastructure for mission-critical or highly regulated systems. Comfortable with incident response, post-mortems, and driving systemic improvements. Experience with incident command systems and escalation procedures.

Skills

Required

  • AWS Cloud Services

    Expert-level experience with AWS services including EC2, ECS/EKS, RDS, DynamoDB, S3, Lambda, CloudFormation, and VPC networking. Expertise in designing scalable, secure, and cost-efficient cloud architectures on AWS.

  • Kubernetes Orchestration

    Production experience deploying, managing, and scaling Kubernetes clusters. Deep knowledge of pods, services, deployments, stateful sets, ingress controllers, network policies, and RBAC security models.

  • Infrastructure as Code (Terraform)

    Advanced proficiency with Terraform for declaring and managing cloud infrastructure. Experience with state management, module design, testing infrastructure code, and GitOps workflows.

  • Distributed Systems Design

    Deep expertise in distributed system patterns, trade-offs, and failure modes. Proficiency with consistency models, consensus protocols, idempotency, and designing for fault tolerance.

  • Database Administration and Optimization

    Hands-on expertise in SQL and NoSQL database operations. Experience with query optimization, indexing strategies, replication configuration, backup strategies, and capacity planning.

  • Observability and SRE Practices

    Expert implementation of SLOs, error budgets, and observability frameworks. Proficiency with monitoring tools, logging aggregation, tracing systems, and metrics collection for production systems.

  • Cross-Functional Technical Leadership

    Demonstrated ability to lead technical initiatives across team boundaries, coordinate with product engineering, and drive consensus on architectural decisions affecting the entire organization.

  • Modern Programming Languages

    Strong proficiency in at least two programming languages commonly used in infrastructure or systems engineering such as Go, Python, Java, Rust, or TypeScript. Clean code practices and deep language-specific ecosystem knowledge required.

  • AI Tooling and Coding Agents

    Comfortable and proficient using AI tooling and coding agents like GitHub Copilot, ChatGPT, and Claude as part of everyday engineering workflow. Ability to leverage AI to accelerate infrastructure development while maintaining code quality and security.

Preferred

  • Cellular and Multi-Tenant Architecture

    Nice to have

    Experience designing or operating cellular architectures or multi-tenant SaaS platforms. Understanding of tenant isolation, resource quotas, and patterns for scaling to thousands of independent deployments.

  • Workflow Orchestration (Temporal)

    Nice to have

    Hands-on experience with Temporal or similar workflow orchestration frameworks. Understanding of durable execution, failure recovery, and building resilient business processes.

  • Developer Platform and Internal Tooling

    Nice to have

    Track record building or improving developer platforms, internal tools, or developer experience. Experience measuring and optimizing developer velocity and infrastructure self-service capabilities.

  • Fintech and Regulated Industries

    Nice to have

    Prior experience in fintech, payments processing, or highly regulated industries. Familiarity with compliance requirements like FedRAMP, SOC 2, PCI-DSS, and designing systems for regulatory constraints.

  • Multi-Region and Edge Infrastructure

    Nice to have

    Experience designing and operating multi-region deployments and edge computing infrastructure. Understanding of geo-distribution, latency optimization, and regional compliance requirements.

  • API Gateway and Service Mesh

    Nice to have

    Production experience with API gateways, service meshes (Istio, Consul, Linkerd), or similar infrastructure patterns. Understanding of load balancing, traffic management, and service-to-service communication.

  • Cost Optimization and FinOps

    Nice to have

    Track record of implementing cost attribution, analyzing infrastructure spending, and optimizing cloud costs. Experience with commitment discount strategies and cost-aware architecture design.

  • Security and Compliance Infrastructure

    Nice to have

    Hands-on experience with infrastructure security including secrets management, certificate rotation, network security, encryption, and compliance automation. Understanding of threat models and security best practices.

Compensation

Pay and benefits.

Base·USD 168,000 – 324,500

Equity·Stock options

Full posting

Original listing.

About Ramp

Ramp is building the smart infrastructure for finance teams, embedded in the transaction flow of every dollar a business spends. We automate how over $200B in annualized spend flows in and out of 70,000+ companies: authorizing payments, flagging risk, categorizing spend, and closing books.

The problems are high-stakes, data-dense, and unforgiving.

We hire people with high agency and high urgency. We look for slope over intercept. We care less about where you trained and more about what you’ve built. At Ramp, everyone is a builder who owns problems end to end and makes consequential decisions that shape the outcome.

The median Ramp customer saves 5% and grows revenue 16% in their first year – far in excess of businesses operating without Ramp. We believe every ambitious company deserves the same.

If you want to build systems that directly shape how companies move and manage billions, Ramp is the place to do it.

About Production Engineering

Production Engineering is Ramp's infrastructure ownership layer. We exist to make Ramp faster, more reliable, and more scalable — and we do that by being embedded in the problems, not adjacent to them.

A few things that define how we operate:

  • One team, one company, one objective. There is no "infra team" and "product team" — there is Ramp. We share the company's goals as our own. When a product team struggles with reliability or scalability, that is our struggle.

  • If reliability or scalability is at risk, we own it. We don't wait to be invited, and we don't ask whose code it is. If a system is slow, if it breaks, if it won't scale — that's ours to lead, regardless of where it lives in the stack.

  • We go first, and we go fast. When the path isn't obvious, we don't wait for someone else to find it. We move with urgency, propose the solution, align the stakeholders, and stay in until it's done — not until our ticket is closed.

  • We lead the way. We find the next problem before it finds us. And when we solve it, we don't just fix it for ourselves — the patterns and standards we establish become the foundation the rest of Ramp builds on. That's not a side effect of the job; it's the job.

  • We stay calibrated. Speed means nothing if we're moving in the wrong direction. We regularly stop and ask honestly whether what we're working on is still the highest-leverage thing we could be doing — the discipline that makes sure our effort compounds toward what actually matters.

You cannot build the future of finance on shaky infrastructure.

Our Teams

Production Engineering is organized into teams — but we think about teams differently. Teams are mutable. We build them around what needs to be done, not around what we need to do given the teams we already have. The structure below reflects our current priorities; it will evolve as Ramp does.

Right now, we operate across four areas:

  • Compute — the foundation everything runs on: container orchestration, networking, load balancing, edge infrastructure, and the deployment systems that get code from engineers' laptops to production reliably and at scale

  • Storage — databases, caching, object storage, and the data infrastructure that underpins everything

  • Workflows & Messaging — the systems that power Ramp's financial workflows and event-driven architecture

  • Internal Infrastructure — the platform that makes every builder at Ramp faster and more autonomous: observability, cost attribution, and CI/CD systems that give teams visibility into what they build and what it costs

What You'll Do

Production Engineers at Ramp are full software engineers who happen to specialize in infrastructure. You write production code, you own systems end-to-end, and you drive technical outcomes across the organization — not just within your team.

Day to day, you will:

  • Build and operate critical infrastructure across Ramp's compute, storage, messaging, and observability stack — owning the systems that handle real financial transactions at scale.

  • Drive architectural change — not just flag problems. When you surface a reliability or scalability issue, you own the path forward: you propose the solution, find the owners across engineering, and stay in until it's resolved.

  • Partner with product teams at the design phase — reviewing architectures, embedding golden paths, and making it easy to build correctly the first time.

  • Build Ramp's next level of scale — you'll be a hands-on contributor to the most consequential infrastructure shift happening right now: our move to a cellular architecture, enabling Ramp to scale, reach international markets, operate in highly regulated and constrained environments (e.g. FedRAMP), and deliver on enterprise-grade SLAs.

  • Enable AI-native engineering — as Ramp builds increasingly AI-powered products, PE is the team that makes sure the platform can support them. You'll proactively partner with product teams on AI infrastructure patterns, define the golden paths that turn one-off solutions into reusable foundations, and stay ahead of emerging challenges before they become blockers.

  • Build developer tooling and self-service infrastructure — so that other teams can answer their own questions (cost, performance, reliability) without involving PE.

  • Participate in on-call rotation — and more importantly, use every incident as a signal to eliminate the root cause, not just resolve the symptom.

  • Lead across the company — PE doesn't just review designs or show up when called. We proactively initiate cross-team architectural reviews, and we take full ownership of company-wide reliability and scalability initiatives: identifying the problem, proposing the solution, aligning the stakeholders, and staying in until it's done.

What We Look For

We don't hire for a checklist. We hire for the instincts that make someone exceptional on this team.

The mindset we're looking for:

  • You can't walk past something broken without wanting to fix it — and you don't stop at the workaround. You find the root cause and eliminate the class of problem.

  • You think in systems, not tasks. You understand that "fixing the bug" is step one; making it so the bug can't happen again is the job.

  • You move with urgency but without chaos. You know when to go fast and when to slow down and do it right.

  • You treat product teams as your users. You build for them, communicate with them proactively, and measure success by how much faster they can ship. But you're also an owner — when something is broken or blocking, you don't wait to be asked. We win when our customers win.

  • You leave things better than you found them. Every system you touch, every process you encounter — you move it forward. Not always dramatically, but always directionally. You picked it up in a certain state; you leave it in a better one.

  • You are equally comfortable reading a database query plan, reviewing a system design doc, writing a Terraform module, and jumping into an incident at 2am.

Experience profile:

  • 3+ years of software engineering management shipping high-quality architectures for critical systems

  • Strong software engineering fundamentals — you write clean, well-tested, production-ready code

  • Hands-on experience with distributed systems at production scale

  • Experience with at least one major cloud provider (AWS preferred)

  • Familiarity with observability practices (SLOs, error budgets, alerting, dashboards)

  • Track record of leading technical projects end-to-end, including cross-team coordination

  • Comfortable using AI tooling and coding agents as part of your everyday engineering workflow — we expect our engineers to leverage these tools to move faster and think bigger

Bonus (not required):

  • Experience with cellular or multi-tenant architecture patterns

  • Prior work on workflow orchestration systems (Temporal)

  • Contributions to developer experience or internal platform tooling

  • Experience in fintech, payments, or regulated industries (FedRAMP, SOC 2)

Benefits available to all full-time Ramp employees (Global)

  • Flexible PTO

  • Centralized home-office equipment ordering

  • Health and wellness stipend

  • Budget for intra-office travel

  • Weekly coffee stipend

United States

  • 100% medical, dental & vision insurance coverage for you, with partial coverage for dependents

  • One Medical annual membership

  • 401(k), including employer match on contributions made while employed by Ramp

  • Fertility HRA (up to $10,000 per year)

  • Parental leave: up to 16 weeks (birthing + bonding) or 8 weeks (bonding only) at 100% pay

  • Pet insurance

  • In-office perks: lunch, snacks, drinks, and more

  • Relocation expense coverage to NYC or SF (if needed)

Canada

  • Group medical, dental, and vision coverage through Sun Life

  • Life, AD&D, and disability coverage

  • Fertility drug coverage (up to $4,000 lifetime)

  • Group Retirement Plan with employer match (RRSP + DPSP)

  • Parental leave: up to 16 weeks (birthing + bonding) or 8 weeks (bonding only) at 100% pay, with additional time available at reduced pay

  • Employee Assistance Program and virtual care through Lumino Health

United Kingdom

  • Private medical insurance through Freedom Elite

  • Virtual GP and at-home care via eMed x Livi

  • Workplace pension through Penfold, with salary sacrifice option

  • Parental leave: up to 16 weeks (birthing + bonding) or 8 weeks (bonding only) at 100% pay with additional time available at reduced pay

Referral Instructions

If you are being referred for the role, please contact that person to apply on your behalf.

 

Other notices

Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

 

Beware of recruiting scams: Ramp will only contact you through official @Ramp.com email addresses and will never ask for payment or sensitive personal information during the hiring process.

 

Ramp Applicant Privacy Notice

Redirects to Ramp's application page.

Other roles

More at Ramp.

View all 25 roles