Staff Site Reliability Engineer, Release Engineering
Site Reliability Engineer · Staff · Full Time
Opens Plaid's application page
Role
What you'll do.
Staff Site Reliability Engineer at Plaid's Infrastructure team, focusing on Release Engineering. This hands-on technical leadership role involves architecting SLO and error-budget frameworks, driving progressive delivery adoption, and ensuring production readiness across product engineering. You'll design and scale reliability practices, build self-service platform tooling, and lead incident response while preparing systems for an AI-driven development landscape. Requires 8+ years in backend systems, SRE, or platform engineering with proven expertise in reliability program design and canary deployment systems.
Responsibilities
- Architect and Manage SLO and Error-Budget Framework: Design and implement comprehensive Service Level Objectives (SLO) and error-budget programs that empower engineering teams to leverage reliability data for informed product and release decisions. Establish metrics-driven governance structures that balance innovation velocity with production stability across all product teams.
- Lead Reliability Standards Expansion: Define and scale Plaid's reliability practices across all product engineering organizations. Convert foundational infrastructure investments into lasting operational habits, documentation, and cultural shifts that embed production-readiness thinking into team workflows and decision-making processes.
- Drive Progressive Delivery and Automated Safety Gates Adoption: Promote widespread adoption of progressive delivery patterns, including canary rollouts, metric-gated analysis, and automated rollback mechanisms. Design intuitive tooling and self-service platforms that enable teams to maintain high development velocity without compromising production safety or user experience.
- Guide Product Teams Toward Production Readiness: Partner with emerging product teams to ensure production readiness through hands-on expertise in observability, incident response procedures, and scalable deployment health practices. Provide technical mentorship and establish maturity assessment frameworks that guide teams through reliability milestones.
- Lead Critical Incident Response and Platform Improvements: Direct response efforts during critical production incidents, ensuring minimal customer impact and rapid resolution. Facilitate comprehensive post-mortem analysis processes and translate findings into permanent platform improvements, preventive tooling, and architectural enhancements.
- Cross-Functional Platform Collaboration: Collaborate with SRE, Platform Engineering, and Infrastructure teams to translate complex production requirements into intuitive, user-friendly platform features and tooling. Serve as a bridge between operational needs and platform capabilities, ensuring solutions address real-world deployment and reliability challenges.
- Prepare Systems for AI-Driven Development Velocity: Architect and scale safety nets, deployment systems, and reliability infrastructure to handle increased volume and frequency of code changes driven by AI-assisted development tools. Establish guardrails and automated checks that maintain production stability as development velocity accelerates.
Qualifications
What we look for.
Technical
Canary Rollout and Progressive Delivery Systems
Direct hands-on experience building or operating canary deployment systems, metric-gated analysis pipelines, or automated rollback infrastructure in production environments at scale.
Service Level Objectives and Reliability Frameworks
Proven expertise designing and implementing reliability programs such as service maturity models, SLI frameworks, or error-budget systems that achieved measurable cross-team adoption and cultural impact.
Systems Programming and Backend Development
Strong technical proficiency in backend systems development, with demonstrated expertise in Go or similar systems languages. Ability to author and review production-grade infrastructure code.
Kubernetes and Container Orchestration
Solid experience with Kubernetes architecture, container deployment patterns, and orchestration best practices in cloud-native environments.
Observability and Monitoring Stack
Deep expertise with Prometheus, time-series databases, and observability platforms. Ability to design comprehensive monitoring, alerting, and tracing strategies for complex distributed systems.
Service Mesh and Advanced Networking
Prior exposure to service mesh technologies and their role in enabling progressive delivery, traffic management, and reliability patterns in microservice architectures.
Infrastructure as Code and GitOps
Hands-on experience with ArgoCD, Terraform, or similar infrastructure-as-code and GitOps tools for managing declarative, version-controlled infrastructure and deployments.
Education
Bachelor's Degree in Computer Science or Related Field
Formal education in Computer Science, Software Engineering, Mathematics, or related technical discipline, or equivalent professional experience demonstrating advanced systems thinking.
Experience
Senior Backend or Platform Engineering
Minimum 8 years of professional experience in backend systems development, Site Reliability Engineering (SRE), or platform engineering roles, with progressive responsibility and impact on infrastructure scale.
Production Reliability and Incident Management
Substantial experience designing and operating production systems with proven expertise in incident response, root cause analysis, postmortem processes, and translating incidents into platform improvements.
Organizational Influence and Change Leadership
Demonstrated ability to drive organizational change and influence engineering culture without formal authority. Track record of getting buy-in from multiple engineering teams for reliability initiatives and practices.
Skills
Required
Go Programming
Advanced proficiency in Go for systems programming, with ability to design and implement high-performance, concurrent infrastructure components.
Kubernetes Architecture
Deep understanding of Kubernetes deployment models, networking, storage, and operational patterns at enterprise scale.
Prometheus and Time-Series Monitoring
Expert-level knowledge of Prometheus architecture, metric design, and time-series analysis for building effective observability solutions.
Canary Deployment Systems
Hands-on experience with canary rollout patterns, progressive delivery orchestration, and metric-gated deployment safety mechanisms.
SLO and SLI Design
Expertise in defining Service Level Objectives, Service Level Indicators, and error budgets that align with business objectives and technical capabilities.
Incident Command and Post-Mortem Facilitation
Mastery of incident management processes, blameless postmortem facilitation, and translating incident learnings into actionable improvements.
Technical Leadership and Influence
Ability to influence engineering culture, drive adoption of reliability practices, and lead technical initiatives across multiple teams without formal authority.
Preferred
Service Mesh Technologies (Istio, Envoy)
Nice to haveExperience with service mesh platforms for traffic management, observability integration, and enabling progressive delivery patterns in microservice environments.
ArgoCD and GitOps Workflows
Nice to haveHands-on experience implementing GitOps deployment patterns and continuous delivery pipelines using declarative infrastructure approaches.
Distributed Systems Design
Nice to haveUnderstanding of distributed systems principles, consensus algorithms, and patterns relevant to building resilient infrastructure platforms.
Policy as Code and OPA/Conftest
Nice to haveExperience with policy-as-code frameworks for automating compliance checks, deployment guardrails, and infrastructure validation across organizations.
Terraform and Infrastructure as Code
Nice to haveProficiency with Terraform for managing complex multi-cloud or multi-region infrastructure deployments in a reproducible, version-controlled manner.
Cost Optimization and Resource Management
Nice to haveTrack record of designing systems that optimize cloud infrastructure costs while maintaining performance and reliability standards.
FinTech or Payment Systems
Nice to havePrior experience working in financial technology, payment systems, or regulated industries where reliability and audit requirements are critical.
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 207,600 – 273,600
Equity·Stock options
Benefits
Comprehensive Health Coverage
Medical, dental, and vision insurance plans covering employee and family members with competitive deductibles and out-of-pocket maximums.
401(k) Retirement Plan
Company-sponsored 401(k) retirement savings plan with employer matching contributions to support long-term financial planning.
Equity and Stock Options
Participation in company equity programs and stock options aligned with company performance, providing wealth-building opportunities for eligible employees.
Flexible Work Arrangements
Flexible remote work options and location flexibility for engineering teams, enabling work-life balance and access to distributed talent.
Professional Development
Learning budgets, conference attendance support, and professional development opportunities to expand technical skills and industry knowledge.
Paid Time Off
Generous paid time off policies including vacation days, sick leave, and company holidays to support employee wellbeing and work-life balance.
Diversity and Inclusion Programs
Commitment to building a diverse workforce with employee resource groups, mentorship programs, and inclusive hiring practices.
Financial Wellness Programs
Employee assistance programs and financial planning resources, including expertise in fintech benefits given Plaid's industry focus.
Process
Interview steps.
- 01
Initial Screening and Experience Review
Recruiter conducts initial phone screening to assess background in SRE, platform engineering, or backend systems. Focus on verifying professional experience level, familiarity with release engineering practices, and alignment with Staff-level expectations.
- 02
Technical Architecture Discussion
First-round conversation with current SRE or Infrastructure team members exploring specific projects, reliability program design decisions, and technical approach to solving deployment challenges. Discussion centers on canary systems, SLO design, and production incident examples.
- 03
Systems Design Interview
Detailed technical interview focusing on designing a production-grade progressive delivery system or SLO framework. Candidates work through trade-offs in monitoring, rollback strategies, and handling high-velocity deployments while maintaining safety.
- 04
Behavioral and Leadership Interview
Conversation with Engineering Manager or Infrastructure leadership exploring organizational influence, change management approach, incident response philosophy, and mentorship style. Emphasis on driving adoption across skeptical teams and handling high-pressure incidents.
- 05
Fintech Domain and Culture Fit Discussion
Optional conversation exploring fintech experience, understanding of Plaid's mission around financial inclusion, and cultural alignment with Plaid's principles of inventing tomorrow and embracing openness.
- 06
Offer and Compensation Discussion
HR and hiring manager finalize offer including base salary, equity package, and benefits. Discussion covers relocation assistance if needed, start date, and any accommodations required for the onboarding process.
Full posting
Original listing.
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Washington D.C., London and Amsterdam.
Plaid's Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely. Release Engineering owns the path from merge to production, including Plaid's zero-touch deployment system, progressive rollouts, metric-gated analysis, and automatic rollback. Our goal is to make safe shipping the default for every product team.
As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale Plaid's reliability practices across product engineering. You'll architect our SLO and error-budget programs, drive the adoption of progressive delivery, and ensure new products are production-ready. By partnering across product and platform teams, you'll translate complex production needs into intuitive, self-service tooling. This is a hands-on technical leadership role where you'll shape the future of our deployment systems—ensuring they remain fast and safe even as AI-assisted development increases code velocity.
What excites you
Lead the expansion of reliability standards across product engineering, converting foundational infrastructure into lasting operational habits and tooling.
Architect and manage the SLO and error-budget framework, empowering teams to utilize reliability data for strategic product and release choices.
Promote widespread use of progressive delivery and automated safety gates, ensuring high velocity without compromising production stability.
Guide emerging product teams toward production readiness through expertise in observability, incident response, and scalable deployment health.
Collaborate with SRE, Platform, and Infrastructure teams to transform complex production requirements into intuitive, self-service platform features.
Direct the response to critical incidents and ensure the resulting post-mortem actions yield permanent improvements to the platform.
Prepare for an AI-driven development landscape by scaling our safety nets to handle an increased volume and frequency of code changes.
What excites us
Over 8 years of professional experience in backend systems, SRE, or platform engineering roles.
Proven track record of designing reliability programs—such as service maturity models or SLI frameworks—that achieved cross-team adoption.
Direct experience building or operating canary rollout systems, metric-gated analysis, or automated rollback infrastructure.
Technical proficiency in software development, with a preference for Go or similar systems languages.
Ability to drive organizational change and influence engineering culture without formal authority.
Sound technical judgment in high-stakes production scenarios, balancing user impact with developer velocity.
Prior exposure to Kubernetes, service mesh technologies, Prometheus, or ArgoCD is considered a strong asset.
Our culture is rooted in impact and collective growth. We seek technical leaders who resonate with our principles of inventing tomorrow and embracing openness.
Our mission at Plaid is to unlock financial freedom for everyone. To support that mission, we seek to build a diverse team of driven individuals who care deeply about making the financial ecosystem more equitable. We recognize that strong qualifications can come from both prior work experiences and lived experiences. We encourage you to apply to a role even if your experience doesn't fully match the job description. We are always looking for team members that will bring something unique to Plaid!
Plaid is proud to be an equal opportunity employer and values diversity at our company. We do not discriminate based on race, color, national origin, ethnicity, religion or religious belief, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, transgender status, sexual stereotypes, age, military or veteran status, disability, or other applicable legally protected characteristics. We also consider qualified applicants with criminal histories, consistent with applicable federal, state, and local laws. Plaid is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance with your application or interviews due to a disability, please let us know at [email protected].
Please review our Candidate Privacy Notice here.
Additional compensation in the form(s) of equity and/or commission are dependent on the position offered. Plaid provides a comprehensive benefit plan, including medical, dental, vision, and 401(k). Pay is based on factors such as (but not limited to) scope and responsibilities of the position, candidate's work experience and skillset, and location. Pay and benefits are subject to change at any time, consistent with the terms of any applicable compensation or benefit plans.
Redirects to Plaid's application page.
Other roles
More at Plaid.
Staff Software Engineer - Protect
Staff
Staff Software Engineer - Credit Insights
Staff
Staff Software Engineer - AI & Intelligent Tooling
Staff
TechOps Engineer
Mid
Senior Developer Relations Engineer
Senior