Staff Site Reliability Engineer

Site Reliability Engineer · Staff · Full Time

New York City, NYUSD 200k – 250k5mo ago
Apply for this role

Opens Tabs's application page

Role

What you'll do.

Tabs is seeking a Staff Site Reliability Engineer to lead infrastructure evolution and platform reliability for their AI-native revenue platform. The ideal candidate will be a senior technical leader who can design scalable systems, improve observability, and drive operational excellence across a high-growth technology startup.

Responsibilities

  • Infrastructure Management: Lead AWS infrastructure direction and platform evolution, including migration from ECS/Fargate to more scalable runtime environments
  • DevOps Excellence: Enhance CI/CD systems with focus on developer experience, safety, and advanced automation
  • Reliability Engineering: Define and evolve reliability standards, including SLIs, SLOs, and error budget management
  • Incident Response: Manage high-severity incidents, conduct thorough postmortems, and drive actionable improvements
  • System Design: Partner with engineering teams to design resilient, scalable, and observable distributed systems

Qualifications

What we look for.

Technical

  • Cloud Infrastructure

    Extensive experience managing production systems on AWS with platform-level change leadership

  • Programming Languages

    Strong software engineering skills in modern programming languages

  • Distributed Systems

    Proven expertise operating distributed systems at production scale

Education

  • Technical Degree

    Bachelor's or Master's degree in Computer Science, Software Engineering, or related technical field preferred

Experience

  • Professional Experience

    10+ years in SRE, infrastructure, or backend engineering roles

  • System Thinking

    Demonstrated ability to analyze systems holistically, considering risk, rollback strategies, and feedback loops

Skills

Required

  • AWS

    Advanced cloud infrastructure management and architecture skills

  • CI/CD

    Expertise in continuous integration and deployment systems

  • Observability

    Deep understanding of monitoring, logging, and tracing technologies

Preferred

  • Incident Management

    Nice to have

    Experience with structured incident response and postmortem methodologies

  • System Design

    Nice to have

    Strong skills in designing scalable, resilient distributed systems

Tech stack

Languages

PythonGo

Frameworks

GitHub Actions

Databases

AWS RDS

Tools

KubernetesPrometheusGrafana

Other

ECSFargate

Compensation

Pay and benefits.

Base·USD 200,000 – 250,000

Equity·Stock options

Benefits

  • Healthcare

    100% employer-covered monthly healthcare premium including medical, dental, and vision

  • Equity

    Competitive compensation package with stock options

  • Time Off

    Unlimited PTO with up to 12 weeks parental leave

  • Retirement

    401k retirement savings plan

  • Wellness

    Free One Medical Membership and Employee Assistance Program

Process

Interview steps.

  1. 01

    Initial Screening

    Phone or video call with recruiter to discuss background and role fit

  2. 02

    Technical Assessment

    Comprehensive technical evaluation of SRE and systems design skills

  3. 03

    Onsite/Virtual Interviews

    Multiple interview rounds covering technical expertise, system design, and cultural alignment

  4. 04

    Final Interview

    Meeting with engineering leadership to discuss role expectations and team integration

Full posting

Original listing.

Tabs is the leading AI-native revenue platform for modern finance and accounting teams. Tabs agents automates the entire contract-to-cash lifecycle, including billing, collections, revenue recognition, and reporting, to help teams eliminate manual work and accelerate cash flow.

High-growth companies like Cursor and Statsig rely on Tabs to generate invoices directly from contracts, reconcile payments in real time, and automate ASC 606 compliance.

Founded in 2023, Tabs has raised over $91 million from Lightspeed Venture Partners, General Catalyst, and Primary. The team is headquartered in New York and brings deep expertise in finance and AI.

About the Role

We’re looking for a Staff Site Reliability Engineer to lead the evolution of Tabs’ platform as we scale. In this role, you’ll operate as a senior individual contributor, partnering closely with engineering and product teams to design, build, and operate systems that are reliable, observable, and easy to develop on.

You’ll own our infrastructure direction, shape how we ship software, and set the standard for operational excellence across the company. This is a high-impact role for someone who enjoys solving complex systems problems, influencing architecture, and raising the reliability bar without becoming a gatekeeper.

What You’ll Own

  • AWS infrastructure direction and platform evolution, including the migration from ECS/Fargate toward a more modern, scalable runtime

  • CI/CD systems with a strong emphasis on developer experience, safety, and automation (GitHub Actions today; maturing CD tomorrow)

  • Ephemeral environments and preview deploys to speed iteration and increase confidence in changes

  • Observability standards across metrics, logs, and tracing, including alert hygiene, dashboards, and SLO development

  • Incident response, postmortems, and the reliability culture that surrounds them

What You’ll Do

  • Define and evolve reliability standards, SLIs, SLOs, and error budgets

  • Improve observability, alerting, and incident processes across services

  • Lead high-severity incidents and drive clear, actionable follow-ups

  • Partner with engineering teams to design resilient, scalable systems

  • Build automation to reduce toil and lower operational risk

  • Mentor engineers and influence best practices across teams

Who You Are

  • You’ve run production systems on AWS and can lead platform-level change

  • You think in systems: risk, rollback strategy, blast radius, and feedback loops

  • You treat CI/CD and environments as products that should be fast, reliable, and self-serve

  • You influence through trust and clarity rather than control

  • You balance pragmatism with long-term system health

  • You value learning from failure and improving processes over assigning blame

  • You communicate clearly and work well across teams

Experience

  • 10+ years in SRE, infrastructure, or backend engineering roles

  • Strong software engineering experience in one or more modern languages

  • Expertise operating distributed systems in production at scale

  • Deep experience with AWS, observability tooling, and CI/CD systems

  • Comfortable navigating ambiguity and setting direction in a fast-moving environment

Additional Information

This role is based onsite in our Soho office in New York City.

Perks and Benefits (Full-time Employees)

  • Competitive compensation and equity

  • Unlimited PTO

  • Up to 100% employer covered monthly healthcare premium (medical, dental, vision)

  • Lunch provided via Sharebite, plus dinner for any later office days.

  • Parental leave up to 12 weeks

  • Tax free commuter and parking benefits

  • Voluntary insurances (Life, Hospital, Critical Illness, Accident)

  • Employee Assistance Program (Rightway)

  • Free One Medical Membership

  • 401k

Tabs is an equal opportunity employer. We welcome teammates of all identities and do not discriminate on the basis of race, ethnicity, religion, gender identity, sexual orientation, age, disability, veteran status, or any other protected characteristic. We’re committed to creating an environment where everyone can grow, contribute, and feel comfortable being themselves.

Redirects to Tabs's application page.

Other roles

More at Tabs.