# Staff Site Reliability Engineer, Spend
**Company:** [Airwallex](https://scaleengineer.com/companies/airwallex)
As a Staff Site Reliability Engineer on Airwallex's Spend team, you'll architect and deliver scalable cloud infrastructure for critical financial services, leading infrastructure design for complex high-risk projects including new service launches and global data center migrations. This role requires 7+ years of SRE/DevOps experience with deep expertise in cloud platforms, Kubernetes, and production systems supporting high-availability financial compliance requirements, positioning you as a technical leader who bridges infrastructure and product delivery in a fast-paced fintech environment.
**Role:** Site Reliability Engineer
**Seniority:** Staff
**Locations:** SG - Singapore
**Salary:** 180000–280000 SGD
[Apply](https://jobs.ashbyhq.com/airwallex/54413125-c244-4f26-ae1b-85b86bded8a8)
Canonical: https://scaleengineer.com/jobs/airwallex/staff-site-reliability-engineer-spend
---
## Responsibilities

- Cloud Infrastructure Architecture and Implementation: Design and implement scalable cloud infrastructure solutions leveraging AWS or GCP technologies for new Spend platform services and product roadmap initiatives. Lead architectural decisions that balance performance, cost optimization, and operational excellence while supporting rapid product iteration and feature delivery.
- Product Team Collaboration and Embedded Reliability: Embed directly with Spend product development teams to drive reliability from design through deployment, providing technical guidance on architectural decisions, establishing performance baselines, and ensuring operational readiness before production launch. Act as a technical advisor on SRE best practices and help teams adopt reliability-first development methodologies.
- Incident Response Leadership and Observability: Lead incident response procedures for critical systems, establishing post-incident review processes and driving continuous improvement. Architect comprehensive observability solutions including distributed tracing, metrics collection, and alerting strategies. Build and maintain incident runbooks, on-call procedures, and escalation protocols across the Spend infrastructure portfolio.
- SLO Definition and DevOps Performance Management: Own Service Level Objectives (SLOs) and Key Performance Indicators (KPIs) at the team level, establishing reliability targets aligned with business requirements. Monitor DevOps performance metrics, capacity planning, and infrastructure utilization. Drive data-driven decisions regarding infrastructure investments and optimization priorities based on observed system behavior and forecast growth.
- Cross-Functional Compliance and Resilience: Collaborate with central DevOps and security teams to ensure infrastructure meets regulatory compliance requirements for fintech operations. Implement disaster recovery strategies, high-availability architectures, and security best practices. Maintain resilience patterns such as multi-region deployments, failover mechanisms, and backup systems for critical financial services.
- Infrastructure Modernization and Migration Leadership: Lead complex, high-risk infrastructure projects including global data center migrations, Kubernetes cluster upgrades, and modernization of data pipelines. Plan detailed migration strategies, manage dependencies across teams, and ensure zero-downtime transitions for production systems handling financial transactions.

## Requirements

### education

- {"name":"Bachelor's Degree in Computer Science, Software Engineering, or Related Field","description":"Formal educational foundation in computer science or software engineering providing core knowledge of systems design, computer architecture, algorithms, and data structures. Relevant disciplines include Computer Engineering, Information Systems, or Electrical Engineering."}

### technical

- {"name":"Cloud Platform Expertise","description":"Demonstrated proficiency with AWS, GCP, or equivalent cloud platforms. Deep understanding of compute services (EC2/GCE), networking, storage solutions, and cloud-native architecture patterns. Experience managing infrastructure-as-code tooling, API integrations, and multi-region deployments."}
- {"name":"Kubernetes and Container Orchestration","description":"Advanced Kubernetes operational expertise including cluster architecture, resource management, networking policies, and troubleshooting. Experience managing workloads at scale, implementing CI/CD pipelines with container registries, and performance optimization in containerized environments."}
- {"name":"Observability and Monitoring Systems","description":"Expertise implementing comprehensive observability stacks including metrics collection (Prometheus, Datadog, New Relic), distributed tracing (Jaeger, Zipkin), and centralized logging (ELK, Splunk). Ability to design alerting strategies, create meaningful dashboards, and translate metrics into actionable insights."}
- {"name":"Incident Response and Troubleshooting","description":"Proven ability to lead incident response for complex production issues, perform root cause analysis, and implement preventative measures. Experience establishing incident management processes, on-call rotations, and communication protocols. Strong analytical skills for diagnosing infrastructure, application, and network issues."}
- {"name":"Infrastructure-as-Code and Automation","description":"Advanced proficiency with IaC tools such as Terraform, CloudFormation, or equivalent. Experience automating infrastructure provisioning, configuration management, and deployment pipelines. Strong scripting capabilities in Bash, Python, or Go for operational automation and tooling development."}
- {"name":"Production Systems and High-Availability Architecture","description":"Deep experience supporting production systems with stringent uptime and reliability requirements. Understanding of high-availability patterns, load balancing, database replication, and graceful degradation. Experience managing systems subject to compliance audits and security requirements."}

### experience

- {"name":"7+ Years in SRE, DevOps, or Infrastructure Engineering","description":"Minimum seven years of professional experience in site reliability engineering, DevOps, platform engineering, or infrastructure-focused roles. Experience should demonstrate progression from operational support to strategic infrastructure design and cross-team leadership responsibilities."}
- {"name":"SRE Strategy and Leadership at Scale","description":"Proven ability to lead SRE strategy and initiatives for large-scale, cross-functional projects. Experience mentoring junior engineers, establishing standards, and driving organizational adoption of reliability best practices across multiple product teams."}
- {"name":"Fintech or Regulated Industry Experience","description":"Background in financial technology, banking, or other regulated industries such as healthcare or telecommunications. Familiarity with compliance requirements, audit processes, and the operational constraints of highly regulated sectors. Understanding of financial transaction systems and their reliability demands."}
- {"name":"Data Streaming and Analytics Pipeline Experience","description":"Familiarity with real-time data processing technologies such as Apache Kafka, AWS Kinesis, or Apache Flink. Experience managing analytics pipelines, data warehousing solutions, or financial data systems. Understanding of stream processing patterns and event-driven architectures."}

## Skills

### required

- {"name":"AWS or GCP Cloud Platform Administration","description":"Production-level expertise operating and optimizing cloud infrastructure on AWS or GCP. Advanced understanding of networking, IAM, cost optimization, and service configuration."}
- {"name":"Kubernetes Cluster Operations","description":"Production Kubernetes administration including deployment strategies, resource optimization, networking, security policies, and troubleshooting at scale."}
- {"name":"Distributed Systems Troubleshooting","description":"Ability to diagnose complex issues across distributed systems, analyze logs and metrics, and implement solutions that improve system reliability and performance."}
- {"name":"Terraform or CloudFormation","description":"Advanced proficiency writing and maintaining infrastructure-as-code for automated provisioning and management of cloud resources at scale."}
- {"name":"Python or Bash Scripting","description":"Strong scripting capabilities for automation, tooling development, and operational problem-solving in Unix-like environments."}
- {"name":"Incident Leadership and On-Call Management","description":"Experience leading incident response processes, conducting post-mortems, and establishing on-call protocols. Strong communication skills during high-pressure situations."}
- {"name":"Technical Leadership and Cross-Team Collaboration","description":"Demonstrated ability to influence technical decisions across teams, mentor engineers, and establish organizational standards for infrastructure and reliability practices."}

### preferred

- {"name":"Apache Kafka or Event Streaming Platforms","description":"Experience designing and operating Kafka clusters or equivalent event streaming systems. Understanding of message queue patterns, topic configuration, and consumer group management."}
- {"name":"Observability Stack Architecture","description":"Experience implementing end-to-end observability solutions with metrics, traces, and logs. Expertise with tools like Prometheus, Datadog, New Relic, or Splunk."}
- {"name":"Go Language Development","description":"Proficiency developing operational tooling and utilities in Go. Understanding of Go concurrency models and performance characteristics."}
- {"name":"Database Administration and Optimization","description":"Experience managing production databases including backups, replication, scaling strategies, and performance tuning. Familiarity with both relational and NoSQL databases."}
- {"name":"Multi-Region and Disaster Recovery Architecture","description":"Experience designing and implementing multi-region deployments, failover strategies, and disaster recovery procedures for critical systems."}
- {"name":"Fintech Compliance and Security Knowledge","description":"Understanding of financial services compliance frameworks such as PCI-DSS, SOC 2, or banking regulations. Knowledge of data protection and audit requirements in fintech."}
- {"name":"GitOps and Progressive Deployment Strategies","description":"Experience implementing GitOps workflows, blue-green deployments, canary releases, and feature flag management for safe infrastructure changes."}

## Tech stack

### tools

- {"name":"Prometheus","description":"Open-source time-series metrics database and monitoring system for collecting, storing, and querying infrastructure and application metrics."}
- {"name":"Grafana","description":"Visualization platform for creating dashboards from metrics data, enabling real-time monitoring and trend analysis of infrastructure health."}
- {"name":"Datadog or New Relic","description":"Comprehensive observability platforms providing metrics, logs, traces, and APM capabilities for end-to-end system monitoring."}
- {"name":"Jenkins or GitLab CI","description":"CI/CD orchestration platforms for automating build, test, and deployment pipelines. GitLab CI offers integrated container registry and GitOps capabilities."}
- {"name":"PagerDuty or Incident Management Systems","description":"On-call and incident management platforms for coordinating alerts, escalations, and incident response across distributed teams."}
- {"name":"Vault or Secrets Management","description":"Security-focused tooling for managing credentials, API keys, and sensitive configuration in infrastructure and applications."}
- {"name":"Slack or Communication Platforms","description":"Communication and notification infrastructure for alert routing, incident response coordination, and team collaboration."}

### others

- {"name":"Apache Kafka","description":"Distributed event streaming platform used for building real-time data pipelines and managing high-throughput message processing for financial data systems."}
- {"name":"AWS S3 and Cloud Storage","description":"Object storage and backup solutions for data persistence, disaster recovery, and analytics data warehousing."}
- {"name":"Service Mesh (Istio or Linkerd)","description":"Microservice communication layer providing observability, security policies, and traffic management for complex Kubernetes deployments."}
- {"name":"gRPC","description":"High-performance RPC framework used for inter-service communication in modern distributed fintech systems."}

### databases

- {"name":"PostgreSQL","description":"Relational database management system commonly used for transactional systems in fintech applications requiring ACID compliance."}
- {"name":"Redis","description":"In-memory data structure store used for caching, session management, and real-time data processing in high-performance infrastructure."}
- {"name":"Elasticsearch","description":"Distributed search and analytics engine used for centralized logging, log analysis, and operational observability in production systems."}
- {"name":"DynamoDB or NoSQL Databases","description":"Key-value and document stores used for scalable data storage, particularly for systems requiring high throughput and flexible schemas."}

### languages

- {"name":"Python","description":"Primary language for operational scripting, infrastructure automation, and tooling development. Used for cloud SDK interactions, log processing, and system utilities."}
- {"name":"Bash/Shell","description":"Essential for Unix/Linux system administration, CI/CD pipeline scripting, and operational command-line utilities in production environments."}
- {"name":"Go","description":"Increasingly used for building high-performance operational tools, Kubernetes controllers, and infrastructure utilities requiring concurrent execution."}
- {"name":"YAML","description":"Core configuration language for Kubernetes manifests, infrastructure-as-code templates, and CI/CD pipeline definitions."}

### frameworks

- {"name":"Kubernetes","description":"Primary container orchestration platform for managing microservices at scale. Central to infrastructure architecture and application deployment strategy."}
- {"name":"Terraform","description":"Infrastructure-as-code framework for declarative provisioning and management of cloud resources across AWS and GCP environments."}
- {"name":"CloudFormation","description":"AWS-native infrastructure-as-code service for AWS resource orchestration and stack management as an alternative or complementary to Terraform."}
- {"name":"Docker","description":"Container runtime and image standard used across the infrastructure for application packaging and Kubernetes deployments."}

## Benefits

### benefits

- {"name":"Competitive Equity Package","description":"Participate in Airwallex's equity compensation program as part of the overall remuneration, aligning personal financial success with company growth and long-term value creation."}
- {"name":"Professional Development and Learning","description":"Access to continuous learning opportunities, conference attendance, training budgets, and opportunities to work on cutting-edge fintech infrastructure technologies alongside world-class engineers."}
- {"name":"Collaborative Engineering Culture","description":"Join a diverse engineering team of 2,300+ professionals across 27 global offices, with emphasis on technical craftsmanship, ownership, and continuous innovation in financial technology."}
- {"name":"Impact-Driven Work","description":"Solve complex, high-visibility infrastructure challenges that directly impact 250,000+ businesses worldwide, including enterprises like Brex, Navan, Qantas, and SHEIN."}
- {"name":"Global Scale and Infrastructure","description":"Architect and operate infrastructure serving global financial services at scale, with exposure to multi-region deployments, high-availability systems, and complex compliance requirements across jurisdictions."}
- {"name":"Singapore Location Benefits","description":"Based in Singapore with access to quality of life, competitive healthcare, public transportation, and positioning in Asia-Pacific's fintech innovation hub."}
- {"name":"Career Advancement Opportunities","description":"Staff-level position with clear pathways to Principal Engineer and leadership roles within Airwallex's rapidly growing infrastructure and platform engineering organization."}

## Compensation

- **max:** 280000
- **min:** 180000
- **currency:** SGD
- **stockOptions:** true

## Interview process

### steps

## Full description
## **About Airwallex**

Airwallex is the only unified payments and financial platform for global businesses. Powered by our unique combination of proprietary infrastructure and software, we empower over 250,000 businesses worldwide – including Brex, Navan, Qantas, SHEIN and many more – with fully integrated solutions to manage everything from business accounts, payments, spend management and treasury, to embedded finance at a global scale.

Proudly founded in Melbourne, we have a team of over 2,300 of the brightest and most innovative people in tech across 27 offices around the globe. Valued at US$11 billion and backed by world-leading investors including T. Rowe Price, Visa, Mastercard, Robinhood Ventures, Sequoia, Salesforce Ventures, DST Global, and Lone Pine Capital, Airwallex is leading the charge in building the global payments and financial platform of the future. If you’re ready to do the most ambitious work of your career, join us.

## **Attributes We Value**

We hire successful builders with founder-like energy who want real impact, accelerated learning, and true ownership. You bring strong role-related expertise and sharp thinking, and you’re motivated by our mission and [operating principles](https://www.airwallex.com/us/operating-principles). You move fast with good judgment, dig deep with curiosity, and make decisions from first principles, balancing speed and rigor.

You're humble and collaborative; turn zero‑to‑one ideas into real products, and you “get stuff done” end-to-end. You use AI to work smarter and solve problems faster. Here, you’ll tackle complex, high‑visibility problems with exceptional teammates and grow your career as we build the future of global banking. If that sounds like you, let’s build what’s next.

## **About the team**

The Engineering team at Airwallex is a diverse group of innovators, builders, and problem solvers, driven by a mission to empower businesses to operate anywhere, anytime. We thrive in a collaborative and fast-paced environment, where we're constantly pushing the boundaries of what's possible in the financial technology space. As a team, we value technical craftsmanship, continuous learning, and a strong sense of ownership, working together to build scalable, reliable, and secure products that empower businesses of all sizes to grow without borders.

  
Our SRE team is breaking new engineering ground and we have the opportunity to define innovative solutions for a number of challenges, paving the way for other teams to follow in our footsteps. This team is responsible for the availability, performance, monitoring and capacity planning of our Global services.

## **What you’ll do**

As a Staff Site Reliability Engineer, you’ll work closely with product teams in Spend to deliver and maintain scalable, reliable cloud infrastructure in support of key product initiatives. Aligned to the roadmap, you’ll lead on infrastructure design and delivery for complex, high-risk projects such as launching new services, executing global data centre migrations, and modernising data pipelines.

**This role is based in Singapore.**

### **Responsibilities:**

* Architect and implement cloud infrastructure for new services and roadmap initiatives.
* Embed with development teams to drive reliability, performance, and operational readiness.
* Lead incident response, observability, and automation across critical systems.
* Own team-level SLOs, runbooks, and DevOps performance metrics.
* Collaborate with central DevOps and security teams to ensure compliance and resilience.

## **Who you are**

We're looking for people who meet the minimum qualifications for this role. The preferred qualifications are great to have, but are not mandatory.

### **Minimum qualifications:**

* Minimum 7 years in an SRE, DevOps, or infrastructure-focused engineering role.
* Bachelor degree in Computer Science, Software Engineering, or a related field.
* Expertise in cloud platforms (AWS/GCP), Kubernetes, observability, and incident response.
* Able to lead SRE strategy for large-scale, cross-functional projects.
* Strong experience supporting production systems with high availability and compliance requirements.
* Proven ability to work closely with developers and guide reliability best practices.

### **Preferred qualifications:**

* Experience in a fintech or similarly regulated industry.
* Familiarity with data streaming, analytics pipelines, or financial data systems.

#Singapore 

## **Applicant Safety Policy: Fraud and Third-Party Recruiters**

_To protect you from recruitment scams, please be aware that Airwallex will not ask for bank details, sensitive ID numbers (i.e. passport), or any form of payment during the application or interview process. All official communication will come from an @_[_airwallex.com_](http://airwallex.com) _email address. Please apply only through_ [_careers.airwallex.com_](http://careers.airwallex.com) _or our official LinkedIn page._

_Airwallex does not accept unsolicited resumes from search firms/recruiters. Airwallex will not pay any fees to search firms/recruiters if a candidate is submitted by a search firm/recruiter unless an agreement has been entered into with respect to specific open position(s). Search firms/recruiters submitting resumes to Airwallex on an unsolicited basis shall be deemed to accept this condition, regardless of any other provision to the contrary._

## **Equal opportunity**

Airwallex is proud to be an equal opportunity employer. We value diversity and anyone seeking employment at Airwallex is considered based on merit, qualifications, competence and talent. We don’t regard color, religion, race, national origin, sexual orientation, ancestry, citizenship, sex, marital or family status, disability, gender, or any other legally protected status when making our hiring decisions. If you have a disability or special need that requires accommodation, please let us know.
