# ENGINEERING MANAGER - SRE
**Company:** [Xero](https://scaleengineer.com/companies/xero)
Engineering Manager for Xero's Site Reliability Engineering team, leading a growing APAC/EMEA SRE program focused on high-availability infrastructure, disaster recovery testing, and incident management at scale. You will manage a team of 2-4 engineers while driving a 70-80% reduction in customer-detected incidents through building mature HA/DR testing cadences, optimizing incident command protocols, and fostering a reliability-first engineering culture across geographically distributed regions.
**Role:** Engineering Manager
**Seniority:** Manager
**Locations:** AU: Melbourne: (260 Burwood Rd)
**Salary:** 150000–210000 NZD
[Apply](https://jobs.ashbyhq.com/xero/70d10228-148f-456e-b188-43d42c38e84b)
Canonical: https://scaleengineer.com/jobs/xero/engineering-manager-sre
---
## Responsibilities

- Lead SRE Program and Incident Response: Own and drive Xero's reliability programme across Asia-Pacific and EMEA regions, establishing incident command protocols with clear time-to-mitigation targets. Partner with observability specialists and platform teams to close systemic reliability gaps and implement a continuous high-availability and disaster recovery (HA/DR) testing cadence moving from ad hoc execution to reliable weekly operations.
- Build and Scale Engineering Team: Recruit, onboard, and mentor a growing SRE team from the current two engineers to at least four by year-end. Develop structured onboarding programs, conduct regular one-to-ones, and create clear career pathways while maintaining high incident response standards and fostering a reliability-first mindset across the engineering organization.
- Cross-Functional Stakeholder Management: Influence and coordinate with peer teams including Cloud Engineering, Database Reliability Engineering, and Shared Services to align on reliability objectives. Manage distributed teams across multiple time zones and regions, bringing calm and clarity to high-pressure incident response moments while forging durable relationships across functions with conviction and empathy.
- Drive Incident Detection and Prevention: Implement measurable improvements in incident detection, response, and prevention strategies targeting a 70-80% reduction in customer-detected incidents. Close cross-team reliability gaps identified between Observability, Tooling, and product engineering through structured collaboration and data-driven incident analysis.
- Shape Engineering Culture and Best Practices: Champion a culture where reliability is woven into the DNA of the business rather than treated as an afterthought. Mentor engineers on thoughtful problem-solving and ownership principles, establish incident management best practices, and report to the Head of Engineering for Reliability while coordinating globally with US-based Engineering Managers.
- Code-Level Problem Solving: Maintain hands-on engagement by diving into code to solve reliability problems, leveraging AI-assisted tooling and debugging techniques. Collaborate with engineering teams on technical solutions while demonstrating deep software engineering expertise to earn technical credibility and guide architectural reliability improvements.

## Requirements

### education

- {"name":"Bachelor's Degree in Computer Science or Related Field","description":"Formal education in Computer Science, Software Engineering, Systems Engineering, or equivalent demonstrable expertise through professional experience. Advanced degree or certifications in cloud architecture, infrastructure engineering, or SRE methodologies are advantageous."}

### technical

- {"name":"Site Reliability Engineering (SRE) Expertise","description":"Proven 5+ years of SRE or infrastructure leadership experience with demonstrated track record of building or maturing high-availability and disaster recovery (HA/DR) programmes. Must have hands-on experience managing incident command systems at scale and implementing reliability improvements across production environments."}
- {"name":"Software Engineering and Coding Proficiency","description":"Strong software engineering fundamentals with ability to read, write, and debug code in modern programming languages (Go, Python, or similar). Comfortable diving deep into codebases to solve reliability problems and implement infrastructure-as-code solutions for monitoring, alerting, and incident management."}
- {"name":"AI-Assisted Tooling and Modern DevOps","description":"Core competency with AI-assisted development tools, observability platforms, and cloud-native technologies. Experience with infrastructure monitoring, distributed systems debugging, and modern DevOps practices including containerization, Kubernetes, and cloud infrastructure platforms."}
- {"name":"Incident Management and Command Systems","description":"Deep expertise in incident command protocols, on-call rotations, and post-incident review (PIR) processes. Proven ability to establish and optimize incident management at scale across multiple time zones and geographic regions with measurable improvements in time-to-detection and time-to-mitigation metrics."}
- {"name":"Distributed Systems and Cloud Architecture","description":"Strong understanding of distributed systems principles, high-availability architecture patterns, failover mechanisms, and disaster recovery testing methodologies. Experience optimizing reliability across multi-region cloud deployments and managing complex infrastructure dependencies."}

### experience

- {"name":"Leadership of Engineering Teams","description":"Minimum 3-4 years managing engineering teams with responsibility for hiring, onboarding, performance management, and career development. Demonstrated success scaling teams during critical growth periods while maintaining quality standards and building strong team culture."}
- {"name":"Global Incident Response Leadership","description":"Experience leading incident response and management programs across geographically distributed teams spanning multiple time zones. Proven ability to establish incident protocols, lead blameless post-incident reviews, and drive measurable reductions in customer-impacting incidents (target: 70-80% reduction)."}
- {"name":"Cross-Functional Collaboration at Scale","description":"Track record of influencing peer teams without formal authority and building consensus across functions (Platform Engineering, Observability, Database teams, etc.). Experience closing reliability gaps through structured collaboration and working effectively with distributed stakeholders."}
- {"name":"Production Systems Ownership","description":"Hands-on ownership of production systems, on-call responsibilities, and incident investigation at scale. Experience with HA/DR testing programs, chaos engineering initiatives, or reliability improvements in high-traffic production environments serving thousands to millions of users."}

## Skills

### required

- {"name":"Leadership and Team Management","description":"Ability to recruit, mentor, and develop high-performing engineering teams. Strong communication skills with capability to inspire trust, provide constructive feedback, and build psychological safety within distributed teams operating across multiple regions."}
- {"name":"Incident Management and Crisis Communication","description":"Expertise in establishing incident command systems, leading incident response under pressure, and communicating clearly across stakeholders during high-severity outages. Ability to run effective post-incident reviews and drive organizational learning from reliability incidents."}
- {"name":"Systems Thinking and Problem Solving","description":"Deep analytical capability to identify root causes of systemic reliability issues, design holistic solutions, and implement improvements across complex distributed systems. Bias toward independent problem-solving and thinking beyond immediate scope."}
- {"name":"Stakeholder Engagement and Influence","description":"Exceptional ability to build relationships, communicate with empathy and conviction, and influence peer teams across Cloud Engineering, Observability, and Database teams without formal authority. Strong organizational awareness and political acumen in matrix environments."}
- {"name":"Software Development and Debugging","description":"Practical ability to write, read, and debug production code in Python, Go, or similar languages. Comfortable using AI-assisted coding tools and frameworks to solve infrastructure problems, implement monitoring solutions, and automate reliability improvements."}
- {"name":"Cloud Infrastructure and Observability","description":"Working knowledge of cloud platforms (AWS, GCP, or Azure), container orchestration (Kubernetes), infrastructure-as-code tools, and observability stacks (monitoring, logging, tracing). Ability to design and implement reliable, scalable infrastructure patterns."}

### preferred

- {"name":"Chaos Engineering and Testing","description":"Experience designing and executing chaos engineering experiments, game-day scenarios, or comprehensive HA/DR testing programs. Familiarity with tools like Gremlin, Chaos Toolkit, or custom chaos frameworks to validate system resilience."}
- {"name":"SRE Frameworks and Methodologies","description":"Formal knowledge of Google SRE principles, error budgets, SLO/SLI/SLA frameworks, and toil reduction initiatives. Familiarity with SRE-specific practices like runbook automation, monitoring as code, and observability-driven development."}
- {"name":"Multi-Region and Multi-Cloud Operations","description":"Experience managing production systems across multiple geographic regions and cloud providers. Understanding of compliance considerations, data residency requirements, and operational complexity of global infrastructure deployments."}
- {"name":"Financial Software or SaaS Platform Expertise","description":"Background working on financial technology, accounting platforms, SaaS infrastructure, or high-availability platforms serving small business users. Understanding of compliance, security, and availability requirements in regulated environments."}
- {"name":"Mentoring and Coaching","description":"Demonstrated commitment to engineering growth, experience mentoring junior engineers through career transitions, and track record of developing future leaders. Evidence of investing in team professional development and building inclusive engineering cultures."}

## Tech stack

### tools

### others

### databases

### languages

### frameworks

## Benefits

### benefits

- {"name":"Flexible Work Arrangements","description":"Work remotely with flexibility to join the team in person approximately once per quarter. Design your work schedule to suit your rhythm and your team's needs while maintaining strong collaboration and connection with distributed team members across APAC and EMEA regions."}
- {"name":"Professional Development and Career Growth","description":"Access to training programs, conference attendance budgets, and mentorship opportunities. Support for obtaining SRE certifications, cloud architecture credentials, or other professional development aligned with your career goals in reliability engineering leadership."}
- {"name":"Global Team and Impact","description":"Lead mission-critical reliability initiatives that directly impact customer experience for millions of small business users globally. Work with world-class engineering teams and have significant influence on Xero's engineering culture and operational resilience strategy."}
- {"name":"Health and Wellness Benefits","description":"Comprehensive health insurance coverage, mental health support, and wellness programs. Access to fitness resources and employee assistance programs supporting your physical and mental wellbeing."}
- {"name":"Inclusive and Diverse Culture","description":"Commitment to building inclusive teams with diverse perspectives. Emphasis on psychological safety, blameless post-incident reviews, and engineering culture that values thoughtful problem-solving and ownership regardless of background."}

## Compensation

- **max:** 210000
- **min:** 150000
- **currency:** NZD
- **stockOptions:** true

## Interview process

### steps

- {"name":"Recruiter Screening Call","description":"Initial 30-minute conversation with Xero's talent team to discuss your background in SRE leadership, team management experience, and alignment with the role's vision. Opportunity to ask questions about team structure, current initiatives, and working style expectations."}
- {"name":"Technical Leadership Assessment","description":"45-minute conversation with the Head of Engineering for Reliability and US-based Engineering Manager counterpart. Discuss your experience building HA/DR programs, managing incident response at scale, and specific examples of reliability improvements you've driven. Technical depth will be evaluated alongside leadership philosophy."}
- {"name":"Technical Deep Dive","description":"60-minute session with senior SRE engineers on the team. Walk through a complex infrastructure reliability challenge or incident scenario. Demonstrate your hands-on technical capability, problem-solving approach, and ability to explain complex concepts clearly to the team you'll be leading."}
- {"name":"Culture and Team Fit Discussion","description":"30-minute conversation with peer Engineering Managers and cross-functional leaders (Cloud Engineering, Observability, Database teams). Discuss stakeholder management philosophy, your approach to distributed team leadership, and how you build inclusive engineering cultures."}
- {"name":"Leadership Simulation and Take-Home Exercise","description":"Realistic scenario involving incident response decision-making and team prioritization. You may also receive a take-home technical or design exercise focused on reliability architecture, HA/DR testing, or incident command optimization."}
- {"name":"Final Conversation with Head of Reliability Engineering","description":"45-minute final discussion with your direct manager to align on role expectations, team vision, growth opportunities, and organizational context. Opportunity to discuss career trajectory and long-term impact you could have at Xero."}

## Full description
**The role and its impact**

We are seeking an Engineering Manager for our Site Reliability Engineering (SRE) team. In this leadership role, you will own and drive our reliability programme, focusing on high availability and disaster recovery (HA/DR) testing and incident management across our Asia-Pacific and EMEA regions. 

You will lead a growing team through a critical period of expansion, building a culture where reliability is woven into the DNA of the business, not treated as an afterthought.

Your work will have significant impact on Xero's customer experience and operational resilience. By partnering with platform teams, observability specialists, and cross-functional stakeholders, you will close systemic reliability gaps and help reduce customer-detected incidents by 70–80%. You will also shape how our engineering culture evolves, mentoring engineers and fostering an environment where thoughtful problem-solving and ownership are valued.

**The team and how they connect**

The Australia-based SRE team currently comprises two skilled engineers and is growing to at least four by year-end. The team operates collaboratively with Cloud Engineering, Observability, Database Reliability Engineering, and Shared Services to elevate Xero's reliability posture across our platform. You will report to the Head of Engineering for Reliability and partner closely with a US-based Engineering Manager who owns incident management for the EMEA/US time zone, ensuring seamless global coverage. Together, you will run Xero's reliability programme, working in lockstep to deliver measurable improvements in how the business detects, responds to, and prevents incidents.

**The team is currently working on**

• Building and maturing a continuous HA/DR testing cadence, moving from ad hoc to reliable weekly execution

• Establishing and optimising incident command protocols across APAC and EMEA regions, with a focus on time-to-mitigation against clear targets

• Closing cross-team reliability gaps identified between Observability, Tooling, and product engineering through structured collaboration

• Growing the team from two to four engineers whilst maintaining high incident response standards and driving a reliability-first mindset

**Where and how you can work**

We have no standing office attendance requirement; we expect you to join us in person approximately once per quarter. We offer the flexibility to work in a way that suits your rhythm and your team's needs, whilst maintaining the collaboration and connection that makes us stronger together.

**Here are some of the things we are looking for**

• You bring proven SRE or infrastructure leadership experience, ideally managing incident command at scale, with a demonstrable track record of building or maturing HA/DR programmes.

• You possess strong software engineering skills and are comfortable diving into code to solve reliability problems, not just process; fluency with AI-assisted tooling is a core capability, not a nice-to-have.

• Your agency and curiosity shine through—you think and act beyond the job description, bias toward independent problem-solving over waiting for direction, and bring a genuine enthusiasm for lifting reliability across the business.

• You excel at stakeholder management and can influence peer teams without formal authority, forging durable relationships across regions, platforms, and functions with both conviction and empathy.

• You are comfortable operating across time zones and managing distributed teams with ease, bringing calm and clarity to high-pressure incident response moments.

• You care about growing and developing engineers. You will shape the team's culture through structured onboarding, regular one-to-ones, and mentorship that helps people thrive and stay.

**Apply even if your experience isn't a perfect match! At Xero, we hire based on your skills, passion, and the unique perspective you can bring to enhance our culture and team.**
