Staff Site Reliability Engineer, Release Engineering

Site Reliability Engineer · Staff · Full Time

New York City OfficeUSD 208k – 274k2mo ago
Apply for this role

Opens Plaid's application page

Role

What you'll do.

Staff Site Reliability Engineer at Plaid's Infrastructure team, focusing on Release Engineering. This hands-on technical leadership role involves architecting SLO and error-budget frameworks, driving progressive delivery adoption, and ensuring production readiness across product engineering. You'll design and scale reliability practices, build self-service platform tooling, and lead incident response while preparing systems for an AI-driven development landscape. Requires 8+ years in backend systems, SRE, or platform engineering with proven expertise in reliability program design and canary deployment systems.

Responsibilities

  • Architect and Manage SLO and Error-Budget Framework: Design and implement comprehensive Service Level Objectives (SLO) and error-budget programs that empower engineering teams to leverage reliability data for informed product and release decisions. Establish metrics-driven governance structures that balance innovation velocity with production stability across all product teams.
  • Lead Reliability Standards Expansion: Define and scale Plaid's reliability practices across all product engineering organizations. Convert foundational infrastructure investments into lasting operational habits, documentation, and cultural shifts that embed production-readiness thinking into team workflows and decision-making processes.
  • Drive Progressive Delivery and Automated Safety Gates Adoption: Promote widespread adoption of progressive delivery patterns, including canary rollouts, metric-gated analysis, and automated rollback mechanisms. Design intuitive tooling and self-service platforms that enable teams to maintain high development velocity without compromising production safety or user experience.
  • Guide Product Teams Toward Production Readiness: Partner with emerging product teams to ensure production readiness through hands-on expertise in observability, incident response procedures, and scalable deployment health practices. Provide technical mentorship and establish maturity assessment frameworks that guide teams through reliability milestones.
  • Lead Critical Incident Response and Platform Improvements: Direct response efforts during critical production incidents, ensuring minimal customer impact and rapid resolution. Facilitate comprehensive post-mortem analysis processes and translate findings into permanent platform improvements, preventive tooling, and architectural enhancements.
  • Cross-Functional Platform Collaboration: Collaborate with SRE, Platform Engineering, and Infrastructure teams to translate complex production requirements into intuitive, user-friendly platform features and tooling. Serve as a bridge between operational needs and platform capabilities, ensuring solutions address real-world deployment and reliability challenges.
  • Prepare Systems for AI-Driven Development Velocity: Architect and scale safety nets, deployment systems, and reliability infrastructure to handle increased volume and frequency of code changes driven by AI-assisted development tools. Establish guardrails and automated checks that maintain production stability as development velocity accelerates.

Qualifications

What we look for.

Technical

  • Canary Rollout and Progressive Delivery Systems

    Direct hands-on experience building or operating canary deployment systems, metric-gated analysis pipelines, or automated rollback infrastructure in production environments at scale.

  • Service Level Objectives and Reliability Frameworks

    Proven expertise designing and implementing reliability programs such as service maturity models, SLI frameworks, or error-budget systems that achieved measurable cross-team adoption and cultural impact.

  • Systems Programming and Backend Development

    Strong technical proficiency in backend systems development, with demonstrated expertise in Go or similar systems languages. Ability to author and review production-grade infrastructure code.

  • Kubernetes and Container Orchestration

    Solid experience with Kubernetes architecture, container deployment patterns, and orchestration best practices in cloud-native environments.

  • Observability and Monitoring Stack

    Deep expertise with Prometheus, time-series databases, and observability platforms. Ability to design comprehensive monitoring, alerting, and tracing strategies for complex distributed systems.

  • Service Mesh and Advanced Networking

    Prior exposure to service mesh technologies and their role in enabling progressive delivery, traffic management, and reliability patterns in microservice architectures.

  • Infrastructure as Code and GitOps

    Hands-on experience with ArgoCD, Terraform, or similar infrastructure-as-code and GitOps tools for managing declarative, version-controlled infrastructure and deployments.

Education

  • Bachelor's Degree in Computer Science or Related Field

    Formal education in Computer Science, Software Engineering, Mathematics, or related technical discipline, or equivalent professional experience demonstrating advanced systems thinking.

Experience

  • Senior Backend or Platform Engineering

    Minimum 8 years of professional experience in backend systems development, Site Reliability Engineering (SRE), or platform engineering roles, with progressive responsibility and impact on infrastructure scale.

  • Production Reliability and Incident Management

    Substantial experience designing and operating production systems with proven expertise in incident response, root cause analysis, postmortem processes, and translating incidents into platform improvements.

  • Organizational Influence and Change Leadership

    Demonstrated ability to drive organizational change and influence engineering culture without formal authority. Track record of getting buy-in from multiple engineering teams for reliability initiatives and practices.

Skills

Required

  • Go Programming

    Advanced proficiency in Go for systems programming, with ability to design and implement high-performance, concurrent infrastructure components.

  • Kubernetes Architecture

    Deep understanding of Kubernetes deployment models, networking, storage, and operational patterns at enterprise scale.

  • Prometheus and Time-Series Monitoring

    Expert-level knowledge of Prometheus architecture, metric design, and time-series analysis for building effective observability solutions.

  • Canary Deployment Systems

    Hands-on experience with canary rollout patterns, progressive delivery orchestration, and metric-gated deployment safety mechanisms.

  • SLO and SLI Design

    Expertise in defining Service Level Objectives, Service Level Indicators, and error budgets that align with business objectives and technical capabilities.

  • Incident Command and Post-Mortem Facilitation

    Mastery of incident management processes, blameless postmortem facilitation, and translating incident learnings into actionable improvements.

  • Technical Leadership and Influence

    Ability to influence engineering culture, drive adoption of reliability practices, and lead technical initiatives across multiple teams without formal authority.

Preferred

  • Service Mesh Technologies (Istio, Envoy)

    Nice to have

    Experience with service mesh platforms for traffic management, observability integration, and enabling progressive delivery patterns in microservice environments.

  • ArgoCD and GitOps Workflows

    Nice to have

    Hands-on experience implementing GitOps deployment patterns and continuous delivery pipelines using declarative infrastructure approaches.

  • Distributed Systems Design

    Nice to have

    Understanding of distributed systems principles, consensus algorithms, and patterns relevant to building resilient infrastructure platforms.

  • Policy as Code and OPA/Conftest

    Nice to have

    Experience with policy-as-code frameworks for automating compliance checks, deployment guardrails, and infrastructure validation across organizations.

  • Terraform and Infrastructure as Code

    Nice to have

    Proficiency with Terraform for managing complex multi-cloud or multi-region infrastructure deployments in a reproducible, version-controlled manner.

  • Cost Optimization and Resource Management

    Nice to have

    Track record of designing systems that optimize cloud infrastructure costs while maintaining performance and reliability standards.

  • FinTech or Payment Systems

    Nice to have

    Prior experience working in financial technology, payment systems, or regulated industries where reliability and audit requirements are critical.

Tech stack

Languages

GoPython

Frameworks

KubernetesIstio/Service MeshArgoCD

Databases

PrometheusPostgreSQL

Tools

PrometheusGrafanaTerraformGit/GitHubPagerDutySplunk or ELK Stack

Other

CI/CD PipelinesObservability ArchitectureCanary Analysis and ExperimentationProduction Readiness Frameworks

Compensation

Pay and benefits.

Base·USD 207,600 – 273,600

Equity·Stock options

Benefits

  • Comprehensive Health Coverage

    Medical, dental, and vision insurance plans covering employee and family members with competitive deductibles and out-of-pocket maximums.

  • 401(k) Retirement Plan

    Company-sponsored 401(k) retirement savings plan with employer matching contributions to support long-term financial planning.

  • Equity and Stock Options

    Participation in company equity programs and stock options aligned with company performance, providing wealth-building opportunities for eligible employees.

  • Flexible Work Arrangements

    Flexible remote work options and location flexibility for engineering teams, enabling work-life balance and access to distributed talent.

  • Professional Development

    Learning budgets, conference attendance support, and professional development opportunities to expand technical skills and industry knowledge.

  • Paid Time Off

    Generous paid time off policies including vacation days, sick leave, and company holidays to support employee wellbeing and work-life balance.

  • Diversity and Inclusion Programs

    Commitment to building a diverse workforce with employee resource groups, mentorship programs, and inclusive hiring practices.

  • Financial Wellness Programs

    Employee assistance programs and financial planning resources, including expertise in fintech benefits given Plaid's industry focus.

Process

Interview steps.

  1. 01

    Initial Screening and Experience Review

    Recruiter conducts initial phone screening to assess background in SRE, platform engineering, or backend systems. Focus on verifying professional experience level, familiarity with release engineering practices, and alignment with Staff-level expectations.

  2. 02

    Technical Architecture Discussion

    First-round conversation with current SRE or Infrastructure team members exploring specific projects, reliability program design decisions, and technical approach to solving deployment challenges. Discussion centers on canary systems, SLO design, and production incident examples.

  3. 03

    Systems Design Interview

    Detailed technical interview focusing on designing a production-grade progressive delivery system or SLO framework. Candidates work through trade-offs in monitoring, rollback strategies, and handling high-velocity deployments while maintaining safety.

  4. 04

    Behavioral and Leadership Interview

    Conversation with Engineering Manager or Infrastructure leadership exploring organizational influence, change management approach, incident response philosophy, and mentorship style. Emphasis on driving adoption across skeptical teams and handling high-pressure incidents.

  5. 05

    Fintech Domain and Culture Fit Discussion

    Optional conversation exploring fintech experience, understanding of Plaid's mission around financial inclusion, and cultural alignment with Plaid's principles of inventing tomorrow and embracing openness.

  6. 06

    Offer and Compensation Discussion

    HR and hiring manager finalize offer including base salary, equity package, and benefits. Discussion covers relocation assistance if needed, start date, and any accommodations required for the onboarding process.

Full posting

Original listing.

We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam.

Plaid's Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely. Release Engineering owns the path from merge to production, including Plaid's zero-touch deployment system, progressive rollouts, metric-gated analysis, and automatic rollback. Our goal is to make safe shipping the default for every product team.

As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale Plaid's reliability practices across product engineering. You'll architect our SLO and error-budget programs, drive the adoption of progressive delivery, and ensure new products are production-ready. By partnering across product and platform teams, you'll translate complex production needs into intuitive, self-service tooling. This is a hands-on technical leadership role where you'll shape the future of our deployment systems—ensuring they remain fast and safe even as AI-assisted development increases code velocity.

What excites you

  • Lead the expansion of reliability standards across product engineering, converting foundational infrastructure into lasting operational habits and tooling.

  • Architect and manage the SLO and error-budget framework, empowering teams to utilize reliability data for strategic product and release choices.

  • Promote widespread use of progressive delivery and automated safety gates, ensuring high velocity without compromising production stability.

  • Guide emerging product teams toward production readiness through expertise in observability, incident response, and scalable deployment health.

  • Collaborate with SRE, Platform, and Infrastructure teams to transform complex production requirements into intuitive, self-service platform features.

  • Direct the response to critical incidents and ensure the resulting post-mortem actions yield permanent improvements to the platform.

  • Prepare for an AI-driven development landscape by scaling our safety nets to handle an increased volume and frequency of code changes.

What excites us

  • Over 8 years of professional experience in backend systems, SRE, or platform engineering roles.

  • Proven track record of designing reliability programs—such as service maturity models or SLI frameworks—that achieved cross-team adoption.

  • Direct experience building or operating canary rollout systems, metric-gated analysis, or automated rollback infrastructure.

  • Technical proficiency in software development, with a preference for Go or similar systems languages.

  • Ability to drive organizational change and influence engineering culture without formal authority.

  • Sound technical judgment in high-stakes production scenarios, balancing user impact with developer velocity.

  • Prior exposure to Kubernetes, service mesh technologies, Prometheus, or ArgoCD is considered a strong asset.

Our culture is rooted in impact and collective growth. We seek technical leaders who resonate with our principles of inventing tomorrow and embracing openness.

Our mission at Plaid is to unlock financial freedom for everyone. To support that mission, we seek to build a diverse team of driven individuals who care deeply about making the financial ecosystem more equitable. We recognize that strong qualifications can come from both prior work experiences and lived experiences. We encourage you to apply to a role even if your experience doesn't fully match the job description. We are always looking for team members that will bring something unique to Plaid!

Plaid is proud to be an equal opportunity employer and values diversity at our company. We do not discriminate based on race, color, national origin, ethnicity, religion or religious belief, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, transgender status, sexual stereotypes, age, military or veteran status, disability, or other applicable legally protected characteristics. We also consider qualified applicants with criminal histories, consistent with applicable federal, state, and local laws. Plaid is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance with your application or interviews due to a disability, please let us know at [email protected].

Please review our Candidate Privacy Notice here.

Additional compensation in the form(s) of equity and/or commission are dependent on the position offered. Plaid provides a comprehensive benefit plan, including medical, dental, vision, and 401(k). Pay is based on factors such as (but not limited to) scope and responsibilities of the position, candidate's work experience and skillset, and location. Pay and benefits are subject to change at any time, consistent with the terms of any applicable compensation or benefit plans.

Redirects to Plaid's application page.

Other roles

More at Plaid.

View all 23 roles