(Senior) Software Engineer, AI Infrastructure

Backend Engineer · Senior · Full Time

SG - SingaporeSGD 180k – 280k2d ago
Apply for this role

Opens Airwallex's application page

Role

What you'll do.

Join Airwallex's Infrastructure & Productivity team to pioneer an AI agent ecosystem that automates infrastructure operations across global engineering environments. As a (Senior) Software Engineer on the Infrastructure AI team, you'll design goal-oriented AI agents using the Quartermaster platform to autonomously handle SRE, DevOps, and DBA workflows, incident response, database management, and CI/CD operations for hundreds of engineers across multiple time zones. This role requires expertise in agentic AI systems, cloud infrastructure (AWS, GCP, Aliyun), container orchestration (Kubernetes), and production infrastructure management.

Responsibilities

  • Design and Build Goal-Oriented Infrastructure AI Agents: Architect and implement autonomous AI agents on the Quartermaster platform that intelligently handle complex SRE, DevOps, and DBA workflows. These agents will autonomously investigate production incidents and execute remediation strategies, provision infrastructure resources at scale, orchestrate sophisticated deployment processes, manage database operations including schema migrations and optimization, and handle capacity planning and resource management across global infrastructure environments.
  • Engineer Agent Architectures for Production Safety and Reliability: Design robust agent loops incorporating planning algorithms, multi-step reasoning capabilities, intelligent tool-use orchestration, comprehensive error recovery mechanisms, and human-in-the-loop escalation systems. Since infrastructure agents directly interact with production systems, you will implement safety guardrails, audit trails, state management, and predictability measures to ensure agents operate reliably and transparently in mission-critical environments.
  • Integrate AI Agents with Enterprise Infrastructure Tooling: Connect agents to real-world infrastructure systems including Kubernetes clusters for container orchestration, Terraform for infrastructure-as-code provisioning, cloud platform APIs (AWS, GCP, Aliyun) for resource management, observability and monitoring platforms for system insights, database management systems, and CI/CD pipelines. Design clean, composable tool interfaces that enable agents to safely interact with these complex systems.
  • Establish Agent Performance Evaluation Frameworks: Define and measure comprehensive metrics for agent effectiveness including task completion rates, accuracy benchmarks, time-to-resolution improvements, escalation rates, and cost efficiency gains. Build sophisticated evaluation frameworks and automated testing systems to rigorously validate whether agents reliably achieve their operational goals and identify improvement opportunities.
  • Collaborate Cross-Functionally Across Infrastructure Teams: Partner closely with Site Reliability Engineers, DevOps teams, and platform engineering groups to identify the highest-value workflows for intelligent automation. Translate real operational pain points and domain expertise into practical agent capabilities, ensuring the AI systems you build address genuine infrastructure challenges faced by distributed global teams.
  • Shape the Technical Vision for Infrastructure AI: Contribute strategically to Airwallex's technical direction for AI-driven infrastructure automation. Stay at the forefront of agentic AI developments, LLM capabilities, and autonomous systems research. Bring innovative ideas to the team, influence architectural decisions, and help define the path toward increasingly autonomous, intelligent infrastructure operations that transform how the company manages global scale.

Qualifications

What we look for.

Technical

  • Backend Software Engineering Proficiency

    Expert-level experience with backend programming languages including Python, Go, Java, or Kotlin. Demonstrated ability to write clean, maintainable, production-grade code with strong understanding of software architecture patterns, design principles, and best practices in systems programming.

  • Cloud Infrastructure and Containerization

    Hands-on practical experience with major cloud platforms (AWS, GCP, or Aliyun), Kubernetes container orchestration, and infrastructure-as-code tools like Terraform. Deep understanding of cloud-native architectures, networking, service deployments, and operational patterns required to work effectively with infrastructure systems.

  • AI Agent Systems Development

    Core requirement: Demonstrated hands-on experience designing and building agentic AI systems that autonomously plan, reason, and execute multi-step tasks using Large Language Models. Proficiency in agent loop design, tool-use orchestration, agent state and memory management, and evaluation methodologies for autonomous systems.

  • Systems Architecture and Design

    Strong system design capabilities with proven ability to architect reliable, safe, and scalable agent systems that interact with production infrastructure. Experience reasoning about trade-offs between safety, performance, autonomy, and human oversight in complex distributed systems.

  • DevOps and CI/CD Infrastructure

    Solid understanding of CI/CD pipeline architecture, deployment orchestration, version control systems, and infrastructure operations. Familiarity with how modern engineering organizations structure infrastructure automation and the operational challenges these agents will help solve.

Education

  • Computer Science or Related Field

    Bachelor's degree in Computer Science, Software Engineering, or equivalent technical discipline. Equivalent practical experience in software engineering roles demonstrating mastery of core computer science concepts is also valued.

Experience

  • 5+ Years Backend Software Engineering

    Minimum five years of professional software engineering experience with demonstrated expertise in backend systems development, infrastructure engineering, or platform engineering roles. Track record of shipping production systems at scale and solving complex architectural challenges.

  • AI/ML Systems Experience

    Proven experience building machine learning or AI-driven systems, preferably with hands-on work on agentic AI architectures. Understanding of LLM capabilities and limitations, prompt engineering, and how to build reliable systems on top of language models.

  • Production Infrastructure Operations

    Experience working with production infrastructure at meaningful scale — either as an infrastructure engineer, SRE, DevOps engineer, or platform engineer. Familiarity with the operational challenges, incident patterns, and automation opportunities that this role addresses.

  • Cross-Functional Collaboration in Global Teams

    Track record of working effectively across engineering disciplines and organizational boundaries in distributed, multi-timezone environments. Demonstrated ability to communicate technical concepts clearly and build consensus across diverse stakeholder groups.

Skills

Required

  • Python

    Expert proficiency in Python for building production backend systems, with strong understanding of async programming patterns, package management, and performance optimization for large-scale applications.

  • Go

    Strong experience with Go for systems programming, microservices development, and performance-critical infrastructure tooling. Understanding of Go's concurrency model and ecosystem for DevOps tools.

  • Kubernetes

    Deep hands-on experience with Kubernetes cluster management, deployment strategies, resource management, networking, and troubleshooting. Understanding of Kubernetes APIs and custom resource management.

  • Terraform

    Practical expertise with Terraform for infrastructure-as-code provisioning, state management, module design, and managing complex cloud infrastructure through declarative configuration.

  • AWS or GCP

    Comprehensive experience with major cloud platforms (AWS or Google Cloud Platform), including compute services, networking, storage, managed databases, and cloud APIs relevant to infrastructure automation.

  • AI Agent Architecture

    Core competency in designing agentic AI systems including planning algorithms, reasoning loops, tool orchestration, prompt engineering for agent behavior, and state management for autonomous systems.

  • LLM APIs and Frameworks

    Hands-on experience integrating with Large Language Model APIs (OpenAI GPT, Anthropic Claude, etc.) and working with agent orchestration frameworks like LangGraph, CrewAI, or AutoGen for building goal-oriented autonomous systems.

  • System Design and Architecture

    Advanced capability in designing complex distributed systems with emphasis on reliability, safety, scalability, and observability. Experience building systems that interact with production infrastructure.

Preferred

  • LangGraph

    Nice to have

    Experience building stateful agent applications with LangGraph, including managing complex agentic workflows, multi-step reasoning, and tool interactions in production environments.

  • CrewAI or AutoGen

    Nice to have

    Familiarity with agent orchestration frameworks like CrewAI for multi-agent coordination or AutoGen for collaborative autonomous agents. Understanding of framework capabilities and limitations.

  • Observability and Monitoring Systems

    Nice to have

    Experience with observability platforms (Grafana, Prometheus, Datadog, ELK stack) either building monitoring infrastructure or architecting systems that agents can query for operational insights and decision-making.

  • Database Administration and Automation

    Nice to have

    Knowledge of database operations including schema migrations, query optimization, backup and recovery procedures, replication management. Experience automating these workflows or understanding DBA operational patterns.

  • Site Reliability Engineering (SRE) Practices

    Nice to have

    Familiarity with SRE disciplines including incident response procedures, runbook design, postmortem analysis, SLO/SLI definition, and alerting strategies. Understanding the operational domain these agents will transform.

  • Aliyun Cloud Platform

    Nice to have

    Experience with Alibaba Cloud (Aliyun) infrastructure, APIs, and services. Valuable for working with Airwallex's multi-cloud infrastructure spanning AWS, GCP, and Aliyun.

  • Agent Reliability and Safety Engineering

    Nice to have

    Forward-thinking perspective on building safe, reliable autonomous systems. Opinions on agent evaluation methodologies, safety guardrails, human oversight mechanisms, and the trajectory of autonomous infrastructure systems.

  • Incident Response and Troubleshooting

    Nice to have

    Strong troubleshooting skills and experience investigating complex production incidents. Ability to think through incident investigation workflows that can be automated by intelligent agents.

Tech stack

Languages

PythonGoJava or Kotlin

Frameworks

LangGraphCrewAIAutoGenQuartermaster

Databases

PostgreSQLRedisMongoDB

Tools

KubernetesTerraformAWS APIsGoogle Cloud APIsAliyun APIsGrafanaPrometheusDatadogELK Stack (Elasticsearch, Logstash, Kibana)OpenAI GPT APIAnthropic Claude APICI/CD Systems

Other

Agentic AI PrinciplesProduction Infrastructure OperationsLLM Integration PatternsDistributed Systems ArchitectureCloud-Native Architecture

Compensation

Pay and benefits.

Base·SGD 180,000 – 280,000

Equity·Stock options

Full posting

Original listing.

About Airwallex

Airwallex is the only unified payments and financial platform for global businesses. Powered by our unique combination of proprietary infrastructure and software, we empower over 250,000 businesses worldwide – including Brex, Navan, Qantas, SHEIN and many more – with fully integrated solutions to manage everything from business accounts, payments, spend management and treasury, to embedded finance at a global scale.

Proudly founded in Melbourne, we have a team of over 2,300 of the brightest and most innovative people in tech across 27 offices around the globe. Valued at US$11 billion and backed by world-leading investors including T. Rowe Price, Visa, Mastercard, Robinhood Ventures, Sequoia, Salesforce Ventures, DST Global, and Lone Pine Capital, Airwallex is leading the charge in building the global payments and financial platform of the future. If you’re ready to do the most ambitious work of your career, join us.

 

Attributes We Value

We hire successful builders with founder-like energy who want real impact, accelerated learning, and true ownership. You bring strong role-related expertise and sharp thinking, and you’re motivated by our mission and operating principles. You move fast with good judgment, dig deep with curiosity, and make decisions from first principles, balancing speed and rigor.

You're humble and collaborative; turn zero‑to‑one ideas into real products, and you “get stuff done” end-to-end. You use AI to work smarter and solve problems faster. Here, you’ll tackle complex, high‑visibility problems with exceptional teammates and grow your career as we build the future of global banking. If that sounds like you, let’s build what’s next.

 

About the Team

Airwallex’s Infrastructure & Productivity team is on a mission to turn everyday software engineers into superheroes. The team builds and maintains the internal platforms, tooling, and CI/CD infrastructure that power Airwallex’s global engineering organization, serving hundreds of engineers across multiple offices and time zones.

The team is now pioneering a bold new direction: an AI agent ecosystem where goal-oriented agents continuously operate across every repository — automating infrastructure operations, incident response, database management, and DevOps workflows. Built on Quartermaster, our internal agent platform, these agents autonomously plan, reason, and execute multi-step tasks against real infrastructure. You will be joining at the ground floor of this initiative, building the intelligent systems that fundamentally change how infrastructure work gets done at Airwallex.

Location: Singapore

 

What You’ll Do

As a (Senior) Software Engineer on the Infrastructure AI team, you will design, build, and operate goal-oriented AI agents that automate SRE, DevOps, and DBA workflows across Airwallex’s global infrastructure. You won’t be doing operations work yourself — you’ll be building the AI systems that do it. Your agents will autonomously investigate incidents, provision infrastructure, manage databases, execute deployments, and handle the operational tasks that traditionally require human intervention.

You will work on top of Quartermaster, Airwallex’s internal agent platform, building domain-specific agents that interact with real infrastructure tooling (Kubernetes, Terraform, cloud APIs, databases, monitoring systems). This role requires you to think deeply about agent architecture — planning and reasoning loops, tool selection and execution, safety guardrails, human-in-the-loop escalation, and how to make autonomous systems reliable in production infrastructure environments.

 

Responsibilities

  • Build goal-oriented infrastructure AI agents: Design and implement autonomous agents on the Quartermaster platform that handle SRE, DevOps, and DBA workflows — including incident investigation and remediation, infrastructure provisioning, deployment orchestration, database operations, and capacity management.

  • Design agent architectures for safety and reliability: Build robust agent loops with planning, reasoning, tool-use, error recovery, and human-in-the-loop escalation. Infrastructure agents touch production systems — they must be safe, auditable, and predictable.

  • Integrate agents with infrastructure tooling: Give agents the ability to interact with real systems — Kubernetes, Terraform, cloud APIs (AWS, GCP, Aliyun), monitoring/observability platforms, databases, and CI/CD pipelines — through well-designed tool interfaces.

  • Evaluate and improve agent performance: Define metrics for agent effectiveness (task completion rate, accuracy, time-to-resolution, escalation rate). Build evaluation frameworks to measure whether agents reliably achieve their goals.

  • Collaborate cross-functionally: Partner with SRE, DevOps, and platform engineering teams to identify the highest-value workflows for agent automation. Understand real operational pain points and translate them into agent capabilities.

  • Shape the future of infrastructure AI at Airwallex: Contribute to the technical vision for how AI agents will transform infrastructure operations. Stay at the forefront of agentic AI developments and bring new ideas into the team.

 

Who You Are:

We’re looking for people who meet the minimum requirements for this role. The preferred qualifications are great to have, but are not mandatory:

  • 5+ years of professional software engineering experience, with strong proficiency in backend languages such as Python, Go, Java, or Kotlin.

  • Infrastructure and cloud knowledge: Practical experience with cloud platforms (AWS, GCP, or Aliyun), container orchestration (Kubernetes), infrastructure-as-code (Terraform), and CI/CD systems. You understand the domain the agents will operate in.

  • Strong system design skills with the ability to architect reliable, safe, and scalable agent systems that interact with production infrastructure.

  • Excellent communication and collaboration skills, with a track record of working effectively across cross-functional teams in a global environment.

 

Preferred Qualifications:

  • Hands-on experience building agentic AI systems: You have designed and built AI agents that autonomously plan, reason, and execute multi-step tasks using LLMs. This includes agent loop design, tool-use orchestration, and managing agent state and memory. This is a core requirement, not a nice-to-have.

  • Experience with agent orchestration frameworks (e.g., LangGraph, CrewAI, AutoGen, or custom agent frameworks) and LLM APIs (OpenAI, Anthropic, etc.).

  • Experience with observability and monitoring systems (Grafana, Prometheus, Datadog, ELK stack) — either building them or building agents that interact with them.

  • Experience with database administration or automation (schema migrations, query optimization, backup/recovery) — knowledge that helps build effective DBA agents.

  • Understanding of SRE practices: incident response, runbooks, postmortems, SLOs/SLIs — the operational domain these agents will automate.

  • A forward-thinking mindset about where AI agents are headed. You have opinions on agent reliability, safety, evaluation, and the path toward increasingly autonomous infrastructure systems.

Applicant Safety Policy: Fraud and Third-Party Recruiters

To protect you from recruitment scams, please be aware that Airwallex will not ask for bank details, sensitive ID numbers (i.e. passport), or any form of payment during the application or interview process. All official communication will come from an @airwallex.com email address. Please apply only through careers.airwallex.com or our official LinkedIn page.

Airwallex does not accept unsolicited resumes from search firms/recruiters. Airwallex will not pay any fees to search firms/recruiters if a candidate is submitted by a search firm/recruiter unless an agreement has been entered into with respect to specific open position(s). Search firms/recruiters submitting resumes to Airwallex on an unsolicited basis shall be deemed to accept this condition, regardless of any other provision to the contrary.

Equal opportunity

Airwallex is proud to be an equal opportunity employer. We value diversity and anyone seeking employment at Airwallex is considered based on merit, qualifications, competence and talent. We don’t regard color, religion, race, national origin, sexual orientation, ancestry, citizenship, sex, marital or family status, disability, gender, or any other legally protected status when making our hiring decisions. If you have a disability or special need that requires accommodation, please let us know.

Redirects to Airwallex's application page.

Other roles

More at Airwallex.

View all 145 roles