ENGINEERING MANAGER - SRE
Engineering Manager · Manager · Full Time
Opens Xero's application page
Role
What you'll do.
Engineering Manager for Xero's Site Reliability Engineering team, leading a growing APAC/EMEA SRE program focused on high-availability infrastructure, disaster recovery testing, and incident management at scale. You will manage a team of 2-4 engineers while driving a 70-80% reduction in customer-detected incidents through building mature HA/DR testing cadences, optimizing incident command protocols, and fostering a reliability-first engineering culture across geographically distributed regions.
Responsibilities
- Lead SRE Program and Incident Response: Own and drive Xero's reliability programme across Asia-Pacific and EMEA regions, establishing incident command protocols with clear time-to-mitigation targets. Partner with observability specialists and platform teams to close systemic reliability gaps and implement a continuous high-availability and disaster recovery (HA/DR) testing cadence moving from ad hoc execution to reliable weekly operations.
- Build and Scale Engineering Team: Recruit, onboard, and mentor a growing SRE team from the current two engineers to at least four by year-end. Develop structured onboarding programs, conduct regular one-to-ones, and create clear career pathways while maintaining high incident response standards and fostering a reliability-first mindset across the engineering organization.
- Cross-Functional Stakeholder Management: Influence and coordinate with peer teams including Cloud Engineering, Database Reliability Engineering, and Shared Services to align on reliability objectives. Manage distributed teams across multiple time zones and regions, bringing calm and clarity to high-pressure incident response moments while forging durable relationships across functions with conviction and empathy.
- Drive Incident Detection and Prevention: Implement measurable improvements in incident detection, response, and prevention strategies targeting a 70-80% reduction in customer-detected incidents. Close cross-team reliability gaps identified between Observability, Tooling, and product engineering through structured collaboration and data-driven incident analysis.
- Shape Engineering Culture and Best Practices: Champion a culture where reliability is woven into the DNA of the business rather than treated as an afterthought. Mentor engineers on thoughtful problem-solving and ownership principles, establish incident management best practices, and report to the Head of Engineering for Reliability while coordinating globally with US-based Engineering Managers.
- Code-Level Problem Solving: Maintain hands-on engagement by diving into code to solve reliability problems, leveraging AI-assisted tooling and debugging techniques. Collaborate with engineering teams on technical solutions while demonstrating deep software engineering expertise to earn technical credibility and guide architectural reliability improvements.
Qualifications
What we look for.
Technical
Site Reliability Engineering (SRE) Expertise
Proven 5+ years of SRE or infrastructure leadership experience with demonstrated track record of building or maturing high-availability and disaster recovery (HA/DR) programmes. Must have hands-on experience managing incident command systems at scale and implementing reliability improvements across production environments.
Software Engineering and Coding Proficiency
Strong software engineering fundamentals with ability to read, write, and debug code in modern programming languages (Go, Python, or similar). Comfortable diving deep into codebases to solve reliability problems and implement infrastructure-as-code solutions for monitoring, alerting, and incident management.
AI-Assisted Tooling and Modern DevOps
Core competency with AI-assisted development tools, observability platforms, and cloud-native technologies. Experience with infrastructure monitoring, distributed systems debugging, and modern DevOps practices including containerization, Kubernetes, and cloud infrastructure platforms.
Incident Management and Command Systems
Deep expertise in incident command protocols, on-call rotations, and post-incident review (PIR) processes. Proven ability to establish and optimize incident management at scale across multiple time zones and geographic regions with measurable improvements in time-to-detection and time-to-mitigation metrics.
Distributed Systems and Cloud Architecture
Strong understanding of distributed systems principles, high-availability architecture patterns, failover mechanisms, and disaster recovery testing methodologies. Experience optimizing reliability across multi-region cloud deployments and managing complex infrastructure dependencies.
Education
Bachelor's Degree in Computer Science or Related Field
Formal education in Computer Science, Software Engineering, Systems Engineering, or equivalent demonstrable expertise through professional experience. Advanced degree or certifications in cloud architecture, infrastructure engineering, or SRE methodologies are advantageous.
Experience
Leadership of Engineering Teams
Minimum 3-4 years managing engineering teams with responsibility for hiring, onboarding, performance management, and career development. Demonstrated success scaling teams during critical growth periods while maintaining quality standards and building strong team culture.
Global Incident Response Leadership
Experience leading incident response and management programs across geographically distributed teams spanning multiple time zones. Proven ability to establish incident protocols, lead blameless post-incident reviews, and drive measurable reductions in customer-impacting incidents (target: 70-80% reduction).
Cross-Functional Collaboration at Scale
Track record of influencing peer teams without formal authority and building consensus across functions (Platform Engineering, Observability, Database teams, etc.). Experience closing reliability gaps through structured collaboration and working effectively with distributed stakeholders.
Production Systems Ownership
Hands-on ownership of production systems, on-call responsibilities, and incident investigation at scale. Experience with HA/DR testing programs, chaos engineering initiatives, or reliability improvements in high-traffic production environments serving thousands to millions of users.
Skills
Required
Leadership and Team Management
Ability to recruit, mentor, and develop high-performing engineering teams. Strong communication skills with capability to inspire trust, provide constructive feedback, and build psychological safety within distributed teams operating across multiple regions.
Incident Management and Crisis Communication
Expertise in establishing incident command systems, leading incident response under pressure, and communicating clearly across stakeholders during high-severity outages. Ability to run effective post-incident reviews and drive organizational learning from reliability incidents.
Systems Thinking and Problem Solving
Deep analytical capability to identify root causes of systemic reliability issues, design holistic solutions, and implement improvements across complex distributed systems. Bias toward independent problem-solving and thinking beyond immediate scope.
Stakeholder Engagement and Influence
Exceptional ability to build relationships, communicate with empathy and conviction, and influence peer teams across Cloud Engineering, Observability, and Database teams without formal authority. Strong organizational awareness and political acumen in matrix environments.
Software Development and Debugging
Practical ability to write, read, and debug production code in Python, Go, or similar languages. Comfortable using AI-assisted coding tools and frameworks to solve infrastructure problems, implement monitoring solutions, and automate reliability improvements.
Cloud Infrastructure and Observability
Working knowledge of cloud platforms (AWS, GCP, or Azure), container orchestration (Kubernetes), infrastructure-as-code tools, and observability stacks (monitoring, logging, tracing). Ability to design and implement reliable, scalable infrastructure patterns.
Preferred
Chaos Engineering and Testing
Nice to haveExperience designing and executing chaos engineering experiments, game-day scenarios, or comprehensive HA/DR testing programs. Familiarity with tools like Gremlin, Chaos Toolkit, or custom chaos frameworks to validate system resilience.
SRE Frameworks and Methodologies
Nice to haveFormal knowledge of Google SRE principles, error budgets, SLO/SLI/SLA frameworks, and toil reduction initiatives. Familiarity with SRE-specific practices like runbook automation, monitoring as code, and observability-driven development.
Multi-Region and Multi-Cloud Operations
Nice to haveExperience managing production systems across multiple geographic regions and cloud providers. Understanding of compliance considerations, data residency requirements, and operational complexity of global infrastructure deployments.
Financial Software or SaaS Platform Expertise
Nice to haveBackground working on financial technology, accounting platforms, SaaS infrastructure, or high-availability platforms serving small business users. Understanding of compliance, security, and availability requirements in regulated environments.
Mentoring and Coaching
Nice to haveDemonstrated commitment to engineering growth, experience mentoring junior engineers through career transitions, and track record of developing future leaders. Evidence of investing in team professional development and building inclusive engineering cultures.
Compensation
Pay and benefits.
Base·NZD 150,000 – 210,000
Equity·Stock options
Benefits
Flexible Work Arrangements
Work remotely with flexibility to join the team in person approximately once per quarter. Design your work schedule to suit your rhythm and your team's needs while maintaining strong collaboration and connection with distributed team members across APAC and EMEA regions.
Professional Development and Career Growth
Access to training programs, conference attendance budgets, and mentorship opportunities. Support for obtaining SRE certifications, cloud architecture credentials, or other professional development aligned with your career goals in reliability engineering leadership.
Global Team and Impact
Lead mission-critical reliability initiatives that directly impact customer experience for millions of small business users globally. Work with world-class engineering teams and have significant influence on Xero's engineering culture and operational resilience strategy.
Health and Wellness Benefits
Comprehensive health insurance coverage, mental health support, and wellness programs. Access to fitness resources and employee assistance programs supporting your physical and mental wellbeing.
Inclusive and Diverse Culture
Commitment to building inclusive teams with diverse perspectives. Emphasis on psychological safety, blameless post-incident reviews, and engineering culture that values thoughtful problem-solving and ownership regardless of background.
Process
Interview steps.
- 01
Recruiter Screening Call
Initial 30-minute conversation with Xero's talent team to discuss your background in SRE leadership, team management experience, and alignment with the role's vision. Opportunity to ask questions about team structure, current initiatives, and working style expectations.
- 02
Technical Leadership Assessment
45-minute conversation with the Head of Engineering for Reliability and US-based Engineering Manager counterpart. Discuss your experience building HA/DR programs, managing incident response at scale, and specific examples of reliability improvements you've driven. Technical depth will be evaluated alongside leadership philosophy.
- 03
Technical Deep Dive
60-minute session with senior SRE engineers on the team. Walk through a complex infrastructure reliability challenge or incident scenario. Demonstrate your hands-on technical capability, problem-solving approach, and ability to explain complex concepts clearly to the team you'll be leading.
- 04
Culture and Team Fit Discussion
30-minute conversation with peer Engineering Managers and cross-functional leaders (Cloud Engineering, Observability, Database teams). Discuss stakeholder management philosophy, your approach to distributed team leadership, and how you build inclusive engineering cultures.
- 05
Leadership Simulation and Take-Home Exercise
Realistic scenario involving incident response decision-making and team prioritization. You may also receive a take-home technical or design exercise focused on reliability architecture, HA/DR testing, or incident command optimization.
- 06
Final Conversation with Head of Reliability Engineering
45-minute final discussion with your direct manager to align on role expectations, team vision, growth opportunities, and organizational context. Opportunity to discuss career trajectory and long-term impact you could have at Xero.
Full posting
Original listing.
The role and its impact
We are seeking an Engineering Manager for our Site Reliability Engineering (SRE) team. In this leadership role, you will own and drive our reliability programme, focusing on high availability and disaster recovery (HA/DR) testing and incident management across our Asia-Pacific and EMEA regions.
You will lead a growing team through a critical period of expansion, building a culture where reliability is woven into the DNA of the business, not treated as an afterthought.
Your work will have significant impact on Xero's customer experience and operational resilience. By partnering with platform teams, observability specialists, and cross-functional stakeholders, you will close systemic reliability gaps and help reduce customer-detected incidents by 70–80%. You will also shape how our engineering culture evolves, mentoring engineers and fostering an environment where thoughtful problem-solving and ownership are valued.
The team and how they connect
The Australia-based SRE team currently comprises two skilled engineers and is growing to at least four by year-end. The team operates collaboratively with Cloud Engineering, Observability, Database Reliability Engineering, and Shared Services to elevate Xero's reliability posture across our platform. You will report to the Head of Engineering for Reliability and partner closely with a US-based Engineering Manager who owns incident management for the EMEA/US time zone, ensuring seamless global coverage. Together, you will run Xero's reliability programme, working in lockstep to deliver measurable improvements in how the business detects, responds to, and prevents incidents.
The team is currently working on
• Building and maturing a continuous HA/DR testing cadence, moving from ad hoc to reliable weekly execution
• Establishing and optimising incident command protocols across APAC and EMEA regions, with a focus on time-to-mitigation against clear targets
• Closing cross-team reliability gaps identified between Observability, Tooling, and product engineering through structured collaboration
• Growing the team from two to four engineers whilst maintaining high incident response standards and driving a reliability-first mindset
Where and how you can work
We have no standing office attendance requirement; we expect you to join us in person approximately once per quarter. We offer the flexibility to work in a way that suits your rhythm and your team's needs, whilst maintaining the collaboration and connection that makes us stronger together.
Here are some of the things we are looking for
• You bring proven SRE or infrastructure leadership experience, ideally managing incident command at scale, with a demonstrable track record of building or maturing HA/DR programmes.
• You possess strong software engineering skills and are comfortable diving into code to solve reliability problems, not just process; fluency with AI-assisted tooling is a core capability, not a nice-to-have.
• Your agency and curiosity shine through—you think and act beyond the job description, bias toward independent problem-solving over waiting for direction, and bring a genuine enthusiasm for lifting reliability across the business.
• You excel at stakeholder management and can influence peer teams without formal authority, forging durable relationships across regions, platforms, and functions with both conviction and empathy.
• You are comfortable operating across time zones and managing distributed teams with ease, bringing calm and clarity to high-pressure incident response moments.
• You care about growing and developing engineers. You will shape the team's culture through structured onboarding, regular one-to-ones, and mentorship that helps people thrive and stay.
Apply even if your experience isn't a perfect match! At Xero, we hire based on your skills, passion, and the unique perspective you can bring to enhance our culture and team.
Redirects to Xero's application page.
Other roles
More at Xero.
Lead Engineer - AI Workflows
Lead
Senior Security Engineer - Defence
Senior
Senior Security Engineer - Cloud Platform
Senior
Senior Search Engineer
Senior
Engineering Manager - Data
Manager