Staff Site Reliability Engineer, Spend
Site Reliability Engineer · Staff · Full Time
Opens Airwallex's application page
Role
What you'll do.
As a Staff Site Reliability Engineer on Airwallex's Spend team, you'll architect and deliver scalable cloud infrastructure for critical financial services, leading infrastructure design for complex high-risk projects including new service launches and global data center migrations. This role requires 7+ years of SRE/DevOps experience with deep expertise in cloud platforms, Kubernetes, and production systems supporting high-availability financial compliance requirements, positioning you as a technical leader who bridges infrastructure and product delivery in a fast-paced fintech environment.
Responsibilities
- Cloud Infrastructure Architecture and Implementation: Design and implement scalable cloud infrastructure solutions leveraging AWS or GCP technologies for new Spend platform services and product roadmap initiatives. Lead architectural decisions that balance performance, cost optimization, and operational excellence while supporting rapid product iteration and feature delivery.
- Product Team Collaboration and Embedded Reliability: Embed directly with Spend product development teams to drive reliability from design through deployment, providing technical guidance on architectural decisions, establishing performance baselines, and ensuring operational readiness before production launch. Act as a technical advisor on SRE best practices and help teams adopt reliability-first development methodologies.
- Incident Response Leadership and Observability: Lead incident response procedures for critical systems, establishing post-incident review processes and driving continuous improvement. Architect comprehensive observability solutions including distributed tracing, metrics collection, and alerting strategies. Build and maintain incident runbooks, on-call procedures, and escalation protocols across the Spend infrastructure portfolio.
- SLO Definition and DevOps Performance Management: Own Service Level Objectives (SLOs) and Key Performance Indicators (KPIs) at the team level, establishing reliability targets aligned with business requirements. Monitor DevOps performance metrics, capacity planning, and infrastructure utilization. Drive data-driven decisions regarding infrastructure investments and optimization priorities based on observed system behavior and forecast growth.
- Cross-Functional Compliance and Resilience: Collaborate with central DevOps and security teams to ensure infrastructure meets regulatory compliance requirements for fintech operations. Implement disaster recovery strategies, high-availability architectures, and security best practices. Maintain resilience patterns such as multi-region deployments, failover mechanisms, and backup systems for critical financial services.
- Infrastructure Modernization and Migration Leadership: Lead complex, high-risk infrastructure projects including global data center migrations, Kubernetes cluster upgrades, and modernization of data pipelines. Plan detailed migration strategies, manage dependencies across teams, and ensure zero-downtime transitions for production systems handling financial transactions.
Qualifications
What we look for.
Technical
Cloud Platform Expertise
Demonstrated proficiency with AWS, GCP, or equivalent cloud platforms. Deep understanding of compute services (EC2/GCE), networking, storage solutions, and cloud-native architecture patterns. Experience managing infrastructure-as-code tooling, API integrations, and multi-region deployments.
Kubernetes and Container Orchestration
Advanced Kubernetes operational expertise including cluster architecture, resource management, networking policies, and troubleshooting. Experience managing workloads at scale, implementing CI/CD pipelines with container registries, and performance optimization in containerized environments.
Observability and Monitoring Systems
Expertise implementing comprehensive observability stacks including metrics collection (Prometheus, Datadog, New Relic), distributed tracing (Jaeger, Zipkin), and centralized logging (ELK, Splunk). Ability to design alerting strategies, create meaningful dashboards, and translate metrics into actionable insights.
Incident Response and Troubleshooting
Proven ability to lead incident response for complex production issues, perform root cause analysis, and implement preventative measures. Experience establishing incident management processes, on-call rotations, and communication protocols. Strong analytical skills for diagnosing infrastructure, application, and network issues.
Infrastructure-as-Code and Automation
Advanced proficiency with IaC tools such as Terraform, CloudFormation, or equivalent. Experience automating infrastructure provisioning, configuration management, and deployment pipelines. Strong scripting capabilities in Bash, Python, or Go for operational automation and tooling development.
Production Systems and High-Availability Architecture
Deep experience supporting production systems with stringent uptime and reliability requirements. Understanding of high-availability patterns, load balancing, database replication, and graceful degradation. Experience managing systems subject to compliance audits and security requirements.
Education
Bachelor's Degree in Computer Science, Software Engineering, or Related Field
Formal educational foundation in computer science or software engineering providing core knowledge of systems design, computer architecture, algorithms, and data structures. Relevant disciplines include Computer Engineering, Information Systems, or Electrical Engineering.
Experience
7+ Years in SRE, DevOps, or Infrastructure Engineering
Minimum seven years of professional experience in site reliability engineering, DevOps, platform engineering, or infrastructure-focused roles. Experience should demonstrate progression from operational support to strategic infrastructure design and cross-team leadership responsibilities.
SRE Strategy and Leadership at Scale
Proven ability to lead SRE strategy and initiatives for large-scale, cross-functional projects. Experience mentoring junior engineers, establishing standards, and driving organizational adoption of reliability best practices across multiple product teams.
Fintech or Regulated Industry Experience
Background in financial technology, banking, or other regulated industries such as healthcare or telecommunications. Familiarity with compliance requirements, audit processes, and the operational constraints of highly regulated sectors. Understanding of financial transaction systems and their reliability demands.
Data Streaming and Analytics Pipeline Experience
Familiarity with real-time data processing technologies such as Apache Kafka, AWS Kinesis, or Apache Flink. Experience managing analytics pipelines, data warehousing solutions, or financial data systems. Understanding of stream processing patterns and event-driven architectures.
Skills
Required
AWS or GCP Cloud Platform Administration
Production-level expertise operating and optimizing cloud infrastructure on AWS or GCP. Advanced understanding of networking, IAM, cost optimization, and service configuration.
Kubernetes Cluster Operations
Production Kubernetes administration including deployment strategies, resource optimization, networking, security policies, and troubleshooting at scale.
Distributed Systems Troubleshooting
Ability to diagnose complex issues across distributed systems, analyze logs and metrics, and implement solutions that improve system reliability and performance.
Terraform or CloudFormation
Advanced proficiency writing and maintaining infrastructure-as-code for automated provisioning and management of cloud resources at scale.
Python or Bash Scripting
Strong scripting capabilities for automation, tooling development, and operational problem-solving in Unix-like environments.
Incident Leadership and On-Call Management
Experience leading incident response processes, conducting post-mortems, and establishing on-call protocols. Strong communication skills during high-pressure situations.
Technical Leadership and Cross-Team Collaboration
Demonstrated ability to influence technical decisions across teams, mentor engineers, and establish organizational standards for infrastructure and reliability practices.
Preferred
Apache Kafka or Event Streaming Platforms
Nice to haveExperience designing and operating Kafka clusters or equivalent event streaming systems. Understanding of message queue patterns, topic configuration, and consumer group management.
Observability Stack Architecture
Nice to haveExperience implementing end-to-end observability solutions with metrics, traces, and logs. Expertise with tools like Prometheus, Datadog, New Relic, or Splunk.
Go Language Development
Nice to haveProficiency developing operational tooling and utilities in Go. Understanding of Go concurrency models and performance characteristics.
Database Administration and Optimization
Nice to haveExperience managing production databases including backups, replication, scaling strategies, and performance tuning. Familiarity with both relational and NoSQL databases.
Multi-Region and Disaster Recovery Architecture
Nice to haveExperience designing and implementing multi-region deployments, failover strategies, and disaster recovery procedures for critical systems.
Fintech Compliance and Security Knowledge
Nice to haveUnderstanding of financial services compliance frameworks such as PCI-DSS, SOC 2, or banking regulations. Knowledge of data protection and audit requirements in fintech.
GitOps and Progressive Deployment Strategies
Nice to haveExperience implementing GitOps workflows, blue-green deployments, canary releases, and feature flag management for safe infrastructure changes.
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·SGD 180,000 – 280,000
Equity·Stock options
Benefits
Competitive Equity Package
Participate in Airwallex's equity compensation program as part of the overall remuneration, aligning personal financial success with company growth and long-term value creation.
Professional Development and Learning
Access to continuous learning opportunities, conference attendance, training budgets, and opportunities to work on cutting-edge fintech infrastructure technologies alongside world-class engineers.
Collaborative Engineering Culture
Join a diverse engineering team of 2,300+ professionals across 27 global offices, with emphasis on technical craftsmanship, ownership, and continuous innovation in financial technology.
Impact-Driven Work
Solve complex, high-visibility infrastructure challenges that directly impact 250,000+ businesses worldwide, including enterprises like Brex, Navan, Qantas, and SHEIN.
Global Scale and Infrastructure
Architect and operate infrastructure serving global financial services at scale, with exposure to multi-region deployments, high-availability systems, and complex compliance requirements across jurisdictions.
Singapore Location Benefits
Based in Singapore with access to quality of life, competitive healthcare, public transportation, and positioning in Asia-Pacific's fintech innovation hub.
Career Advancement Opportunities
Staff-level position with clear pathways to Principal Engineer and leadership roles within Airwallex's rapidly growing infrastructure and platform engineering organization.
Full posting
Original listing.
About Airwallex
Airwallex is the only unified payments and financial platform for global businesses. Powered by our unique combination of proprietary infrastructure and software, we empower over 250,000 businesses worldwide – including Brex, Navan, Qantas, SHEIN and many more – with fully integrated solutions to manage everything from business accounts, payments, spend management and treasury, to embedded finance at a global scale.
Proudly founded in Melbourne, we have a team of over 2,300 of the brightest and most innovative people in tech across 27 offices around the globe. Valued at US$11 billion and backed by world-leading investors including T. Rowe Price, Visa, Mastercard, Robinhood Ventures, Sequoia, Salesforce Ventures, DST Global, and Lone Pine Capital, Airwallex is leading the charge in building the global payments and financial platform of the future. If you’re ready to do the most ambitious work of your career, join us.
Attributes We Value
We hire successful builders with founder-like energy who want real impact, accelerated learning, and true ownership. You bring strong role-related expertise and sharp thinking, and you’re motivated by our mission and operating principles. You move fast with good judgment, dig deep with curiosity, and make decisions from first principles, balancing speed and rigor.
You're humble and collaborative; turn zero‑to‑one ideas into real products, and you “get stuff done” end-to-end. You use AI to work smarter and solve problems faster. Here, you’ll tackle complex, high‑visibility problems with exceptional teammates and grow your career as we build the future of global banking. If that sounds like you, let’s build what’s next.
About the team
The Engineering team at Airwallex is a diverse group of innovators, builders, and problem solvers, driven by a mission to empower businesses to operate anywhere, anytime. We thrive in a collaborative and fast-paced environment, where we're constantly pushing the boundaries of what's possible in the financial technology space. As a team, we value technical craftsmanship, continuous learning, and a strong sense of ownership, working together to build scalable, reliable, and secure products that empower businesses of all sizes to grow without borders.
Our SRE team is breaking new engineering ground and we have the opportunity to define innovative solutions for a number of challenges, paving the way for other teams to follow in our footsteps. This team is responsible for the availability, performance, monitoring and capacity planning of our Global services.
What you’ll do
As a Staff Site Reliability Engineer, you’ll work closely with product teams in Spend to deliver and maintain scalable, reliable cloud infrastructure in support of key product initiatives. Aligned to the roadmap, you’ll lead on infrastructure design and delivery for complex, high-risk projects such as launching new services, executing global data centre migrations, and modernising data pipelines.
This role is based in Singapore.
Responsibilities:
Architect and implement cloud infrastructure for new services and roadmap initiatives.
Embed with development teams to drive reliability, performance, and operational readiness.
Lead incident response, observability, and automation across critical systems.
Own team-level SLOs, runbooks, and DevOps performance metrics.
Collaborate with central DevOps and security teams to ensure compliance and resilience.
Who you are
We're looking for people who meet the minimum qualifications for this role. The preferred qualifications are great to have, but are not mandatory.
Minimum qualifications:
Minimum 7 years in an SRE, DevOps, or infrastructure-focused engineering role.
Bachelor degree in Computer Science, Software Engineering, or a related field.
Expertise in cloud platforms (AWS/GCP), Kubernetes, observability, and incident response.
Able to lead SRE strategy for large-scale, cross-functional projects.
Strong experience supporting production systems with high availability and compliance requirements.
Proven ability to work closely with developers and guide reliability best practices.
Preferred qualifications:
Experience in a fintech or similarly regulated industry.
Familiarity with data streaming, analytics pipelines, or financial data systems.
#Singapore
Applicant Safety Policy: Fraud and Third-Party Recruiters
To protect you from recruitment scams, please be aware that Airwallex will not ask for bank details, sensitive ID numbers (i.e. passport), or any form of payment during the application or interview process. All official communication will come from an @airwallex.com email address. Please apply only through careers.airwallex.com or our official LinkedIn page.
Airwallex does not accept unsolicited resumes from search firms/recruiters. Airwallex will not pay any fees to search firms/recruiters if a candidate is submitted by a search firm/recruiter unless an agreement has been entered into with respect to specific open position(s). Search firms/recruiters submitting resumes to Airwallex on an unsolicited basis shall be deemed to accept this condition, regardless of any other provision to the contrary.
Equal opportunity
Airwallex is proud to be an equal opportunity employer. We value diversity and anyone seeking employment at Airwallex is considered based on merit, qualifications, competence and talent. We don’t regard color, religion, race, national origin, sexual orientation, ancestry, citizenship, sex, marital or family status, disability, gender, or any other legally protected status when making our hiring decisions. If you have a disability or special need that requires accommodation, please let us know.
Redirects to Airwallex's application page.
Other roles
More at Airwallex.
Senior Software Engineer, Data Platform & AI Enablement
Senior
Staff Software Engineer, Payments
Staff
Staff Data Scientist, Algorithm (Risk Product – AML & Financial Crime)
Staff
Senior Network Engineer, Infrastructure
Senior
Manager, Data Engineering
Manager