# Software Engineer, Machine Learning Infrastructure
**Company:** [Deliveroo](https://scaleengineer.com/companies/deliveroo)
Join Deliveroo's GenAI Platform team to build production-grade infrastructure for generative AI across DoorDash, Wolt, and Deliveroo. This role focuses on scaling open-weight LLM and VLM infrastructure, including real-time GPU serving, high-throughput batch inference, and distributed fine-tuning pipelines while optimizing cost and latency. You'll design systems that power AI agents, automation, and personalization while working across model serving frameworks, GPU autoscaling, backend services, and observability in a fast-moving technical environment.
**Role:** Machine Learning Infrastructure Engineer
**Seniority:** Mid
**Locations:** London - The River Building HQ
**Salary:** 140000–200000 GBP
[Apply](https://jobs.ashbyhq.com/deliveroo/58485b61-4c8e-4996-89bc-685d5e154b94)
Canonical: https://scaleengineer.com/jobs/deliveroo/software-engineer-machine-learning-infrastructure
---
## Responsibilities

- Design and build production GPU serving infrastructure: Architect and implement real-time GPU serving endpoints, high-throughput batch inference pipelines, and autoscaling systems for open-weight LLMs and VLMs. Optimize for cost, latency, and throughput while handling multi-model deployments across inference engines like vLLM and SGLang.
- Develop and optimize model fine-tuning infrastructure: Build distributed fine-tuning and training pipelines supporting SFT, DPO, RLHF, and LoRA techniques on autoscaling GPUs. Handle data preparation, evaluation frameworks, and production-ready training orchestration that supports rapid experimentation.
- Implement GPU autoscaling and resource optimization: Engineer intelligent GPU autoscaling systems, optimize GPU utilization rates, and implement cost attribution mechanisms. Focus on KV-cache optimization, quantization strategies (FP8/INT8/AWQ/GPTQ), and multi-node distributed inference patterns.
- Own platform gateway and observability systems: Develop and maintain LLM Gateway and Agent Gateway components, implement comprehensive observability including tracing and monitoring, build evals infrastructure and guardrails, and establish production reliability standards including SLOs and incident playbooks.
- Partner across engineering organizations: Collaborate closely with ML engineers, product engineers, data scientists, and platform teams across DoorDash, Wolt, and Deliveroo to translate business requirements into scalable platform primitives and infrastructure abstractions.
- Push cost and performance frontiers: Continuously optimize inference and fine-tuning systems to deliver cost and latency wins, turning days-long batch jobs into hours while reducing inference costs by multiples. Evaluate emerging model architectures, vendor offerings, and optimization techniques.
- Establish production excellence standards: Build platforms that support rapid experimentation while maintaining production standards for reliability, performance, monitoring, and operational excellence. Develop playbooks, runbooks, and automated remediation capabilities.

## Requirements

### education

- {"name":"Bachelor's degree in Computer Science","description":"BSc in Computer Science, Computer Engineering, or equivalent discipline providing foundational knowledge in algorithms, data structures, and systems design."}
- {"name":"Master's degree or higher (preferred)","description":"MSc or PhD in Computer Science, Machine Learning, or related field demonstrating advanced research capabilities and deep theoretical understanding of distributed systems or machine learning infrastructure."}

### technical

- {"name":"Backend engineering in Python","description":"Strong proficiency in Python for building scalable backend services, distributed systems, and data pipelines. Experience designing APIs, handling concurrency, and optimizing performance for production workloads."}
- {"name":"Distributed systems architecture","description":"Deep understanding of distributed systems concepts including consensus, partitioning, replication, and orchestration. Experience designing systems that scale horizontally across multiple nodes and handle network failures gracefully."}
- {"name":"Production LLM serving and inference","description":"Hands-on production experience with LLM inference, model serving frameworks (vLLM, SGLang, TensorRT-LLM), and optimization techniques including batching, autoscaling, KV-cache management, and quantization methods."}
- {"name":"GPU computing and CUDA fundamentals","description":"Understanding of GPU architecture, memory hierarchies, CUDA programming basics, and GPU-specific optimization strategies for machine learning workloads including distributed multi-GPU training and inference."}
- {"name":"Production observability and debugging","description":"Proficiency in production monitoring, logging, tracing, and debugging complex distributed systems. Experience with metrics collection, anomaly detection, and performance profiling at scale."}
- {"name":"Open-weight model fine-tuning","description":"Production experience with fine-tuning open-weight models including supervised fine-tuning (SFT), DPO, RLHF, and LoRA techniques. Understanding of training dynamics, convergence, and evaluation methodologies."}
- {"name":"Cloud infrastructure and containerization","description":"Experience with Kubernetes orchestration, cloud platforms (AWS/GCP), containerization strategies, and serverless/elastic GPU platforms. Knowledge of infrastructure-as-code and deployment automation."}

### experience

- {"name":"3+ years software engineering industry experience","description":"Minimum three years of professional software engineering experience building production systems, preferably in backend infrastructure, platform engineering, or machine learning systems."}
- {"name":"Production-scale data infrastructure","description":"Track record of building and operating production services, APIs, data pipelines, or ML infrastructure serving significant scale. Demonstrated ability to handle reliability, performance optimization, and operational excellence."}
- {"name":"Production system operations experience","description":"Real-world experience operating systems in production including incident response, performance optimization, cost optimization, debugging complex failures, and establishing reliability standards."}
- {"name":"AI coding tools proficiency","description":"Proficiency using modern AI coding assistants (Claude, Codex, Cursor) throughout the full software development lifecycle including design, code generation, testing, monitoring, and deployment."}

## Skills

### required

- {"name":"Python backend development","description":"Expert-level Python programming for building scalable backend services, production APIs, and data processing pipelines with focus on performance and reliability."}
- {"name":"Distributed systems design","description":"Ability to architect and implement distributed systems with consideration for scalability, fault tolerance, consistency models, and operational complexity."}
- {"name":"Production LLM infrastructure","description":"Practical experience building or operating LLM serving infrastructure in production environments, including inference optimization, model routing, and resource management."}
- {"name":"GPU and CUDA fundamentals","description":"Working knowledge of GPU architecture, CUDA programming basics, and how to optimize machine learning computations for GPU execution and memory efficiency."}
- {"name":"Cloud infrastructure (AWS/GCP)","description":"Hands-on experience deploying and managing infrastructure on AWS or GCP including compute services, networking, storage, and cost optimization."}
- {"name":"Kubernetes and container orchestration","description":"Proficiency with Kubernetes for container orchestration, service deployment, resource management, and production operations."}
- {"name":"Systems observability","description":"Experience implementing monitoring, logging, tracing, and alerting for production systems. Ability to debug performance issues and establish SLOs."}
- {"name":"Fine-tuning techniques for LLMs","description":"Practical knowledge of supervised fine-tuning (SFT), direct preference optimization (DPO), reinforcement learning from human feedback (RLHF), and LoRA-based adaptation methods."}

### preferred

- {"name":"vLLM and SGLang framework experience","description":"Production experience with vLLM, SGLang, or TensorRT-LLM inference engines for optimizing LLM serving performance, batching, and multi-GPU inference."}
- {"name":"LLM quantization techniques","description":"Experience with quantization methods including FP8, INT8, AWQ, GPTQ, and other techniques for reducing model size and inference latency while maintaining accuracy."}
- {"name":"Multi-node distributed training","description":"Experience implementing or operating distributed training systems using frameworks like PyTorch Distributed, DeepSpeed, or Ray for large-scale fine-tuning and training."}
- {"name":"Vector databases and semantic search","description":"Familiarity with vector databases (Pinecone, Weaviate, Milvus), embedding systems, and RAG architectures for retrieval-augmented generation applications."}
- {"name":"LLM evaluation and evals infrastructure","description":"Experience building evaluation frameworks, running LLM benchmarks, and establishing metrics for model quality assessment."}
- {"name":"AI agents and MCP servers","description":"Production experience building AI agents, orchestration frameworks, or Model Context Protocol (MCP) servers that coordinate multiple tools and models."}
- {"name":"API gateway and routing systems","description":"Experience building or operating API gateways, load balancers, or intelligent routing systems that handle vendor abstraction and dynamic routing decisions."}
- {"name":"Cost attribution and metering","description":"Experience implementing cost tracking, usage metering, and attribution systems for multi-tenant infrastructure serving business units."}
- {"name":"Serverless and elastic GPU platforms","description":"Experience with serverless computing platforms or elastic GPU services (Modal, Anyscale, Lambda Labs) for dynamic workload scaling."}
- {"name":"Performance profiling and optimization","description":"Advanced skills in profiling tools, identifying bottlenecks in GPU code, and implementing targeted optimizations for latency and throughput."}

## Tech stack

### tools

- {"name":"Kubernetes","description":"Container orchestration platform for deploying, scaling, and managing containerized ML infrastructure and backend services across GPU clusters."}
- {"name":"AWS (EC2, S3, SageMaker)","description":"Cloud platform services including compute instances (particularly GPU instances), object storage for models, and managed ML services."}
- {"name":"GCP (Compute Engine, Cloud Storage, TPUs)","description":"Alternative cloud provider with Compute Engine instances, managed storage, and tensor processing unit support for large-scale training."}
- {"name":"Prometheus & Grafana","description":"Monitoring and observability stack for collecting metrics, setting up alerts, and visualizing system performance and health."}
- {"name":"Jaeger/Datadog","description":"Distributed tracing platforms for debugging complex request flows through the LLM serving infrastructure and identifying performance bottlenecks."}
- {"name":"Docker","description":"Container technology for packaging services and models consistently across development and production environments."}
- {"name":"Modal","description":"Serverless GPU computing platform enabling elastic scaling of inference and fine-tuning workloads without infrastructure management."}
- {"name":"Git/GitHub","description":"Version control system for infrastructure-as-code, service code, and collaborative development workflows."}

### others

- {"name":"GPU profiling and optimization tools","description":"Expertise with NVIDIA Nsight, cuDNN profiling, and other GPU-specific tools for identifying optimization opportunities in inference and training workloads."}
- {"name":"AI coding assistants","description":"Proficiency with Claude, Codex, Cursor, and similar AI-powered code generation and completion tools throughout the development lifecycle."}
- {"name":"Bash and shell scripting","description":"Scripting capabilities for infrastructure automation, deployment scripts, and operational tooling supporting platform reliability."}
- {"name":"Terraform/Infrastructure-as-Code","description":"Infrastructure provisioning and management using declarative infrastructure-as-code approaches for reproducible infrastructure deployment."}
- {"name":"Model evaluation frameworks","description":"Tools and libraries for benchmarking model quality including MTEB, LMEval, and custom evaluation harnesses specific to application use cases."}

### databases

- {"name":"PostgreSQL","description":"Primary relational database for storing infrastructure state, metrics, and operational data with support for complex queries and strong consistency."}
- {"name":"Redis","description":"In-memory cache and store for high-performance session management, rate limiting, and real-time metrics aggregation."}
- {"name":"Vector databases (Pinecone/Weaviate/Milvus)","description":"Specialized databases for storing and querying embeddings, supporting semantic search and RAG pipeline infrastructure."}
- {"name":"TimescaleDB","description":"Time-series database extension for PostgreSQL enabling efficient metrics storage and time-series analysis for observability systems."}

### languages

- {"name":"Python","description":"Primary programming language for backend services, data pipelines, model serving, and infrastructure tooling. Used extensively for LLM serving implementations and GPU computing orchestration."}
- {"name":"SQL","description":"For data querying, metrics aggregation, and backend service state management in relational databases supporting infrastructure components."}
- {"name":"CUDA/C++","description":"For performance-critical GPU kernels, custom inference optimizations, and interfacing with GPU computing primitives at low levels."}

### frameworks

- {"name":"vLLM","description":"Open-source LLM serving engine providing high-throughput inference with continuous batching, LoRA support, and multi-GPU inference optimization."}
- {"name":"SGLang","description":"Structured generation language and inference engine for efficient LLM serving with complex control flow and batching optimization."}
- {"name":"TensorRT-LLM","description":"NVIDIA's inference optimization framework for low-latency, high-throughput LLM serving with quantization and multi-GPU support."}
- {"name":"PyTorch Distributed","description":"Distributed training framework for multi-GPU and multi-node fine-tuning operations including SFT, DPO, RLHF, and LoRA implementations."}
- {"name":"DeepSpeed","description":"Training optimization library providing distributed training, quantization, and inference optimization capabilities for large-scale model training and serving."}
- {"name":"FastAPI","description":"Modern Python web framework for building high-performance APIs and backend services for the platform gateway and related components."}
- {"name":"Ray","description":"Distributed computing framework for orchestrating distributed training, inference, and batch processing workloads across GPU clusters."}

## Benefits

### benefits

- {"name":"Equity and stock options","description":"Participate in Deliveroo's growth through competitive equity grants allowing you to build long-term wealth alongside the company."}
- {"name":"Comprehensive health and wellness benefits","description":"Health insurance coverage including medical, dental, and vision plans with options for you and your family members."}
- {"name":"Flexible working arrangements","description":"Remote work flexibility and flexible scheduling to support work-life balance while collaborating across distributed teams."}
- {"name":"Professional development budget","description":"Annual learning and development budget for conferences, courses, certifications, and training programs to grow your technical expertise."}
- {"name":"Generous time off policy","description":"Competitive paid time off including vacation days, sick leave, and parental leave to ensure adequate rest and personal time."}
- {"name":"Pension and retirement planning","description":"Employer-contributed pension scheme supporting long-term financial security and retirement planning."}
- {"name":"Mental health and wellbeing support","description":"Access to mental health resources, counseling services, and wellness programs supporting holistic employee wellbeing."}
- {"name":"Parental leave and family support","description":"Comprehensive parental leave policies for various family situations and childcare support options."}
- {"name":"Commuter benefits and transportation","description":"Support for transportation costs including commuter passes and transit benefits for employees commuting to offices."}
- {"name":"Diversity and inclusion initiatives","description":"Active commitment to fostering an inclusive workplace with employee resource groups, diversity programs, and inclusive hiring practices."}

## Compensation

- **max:** 200000
- **min:** 140000
- **currency:** GBP
- **stockOptions:** true

## Interview process

### steps

- {"name":"Initial screening call","description":"30-minute phone screening with a recruiter covering your background, experience with ML infrastructure, and initial technical alignment with the role requirements."}
- {"name":"Technical screening interview","description":"60-90 minute technical conversation with an ML Infrastructure engineer focusing on your production experience with distributed systems, LLM serving, and GPU infrastructure. Expect discussion of architectural decisions and optimization techniques."}
- {"name":"System design deep dive","description":"90-minute interview centered on designing large-scale ML infrastructure systems. You'll discuss tradeoffs in model serving architectures, GPU autoscaling strategies, and how you'd approach cost/performance optimization challenges."}
- {"name":"Production experience panel","description":"Interview with multiple engineers on the GenAI Platform team exploring your hands-on experience operating production systems, incident response capabilities, and how you approach reliability and observability."}
- {"name":"Cross-functional collaboration discussion","description":"Conversation with stakeholders from product, data science, or adjacent platform teams to assess your ability to translate business requirements into platform abstractions and work effectively across organizations."}
- {"name":"Final conversation with hiring manager","description":"Discussion with the engineering manager covering team dynamics, long-term career goals, and fit for Deliveroo's culture of technical excellence and operational ownership."}

## Full description
# **Software Engineer, Machine Learning Infrastructure - Generative AI**

## **About the Team**

Deliveroo's GenAI Platform team sits within Machine Learning Platform and builds the shared infrastructure that helps DoorDash, Wolt, and Deliveroo teams safely bring GenAI-powered products, agents, automation, and personalization to production. Our mission is to increase the velocity of business impact from GenAI. A central pillar of that work is running frontier open-weight LLMs and VLMs (such as GLM, Qwen, Kimi, and DeepSeek) ourselves — real-time GPU serving, high-throughput batch inference, and fine-tuning on autoscaling GPUs — delivering large cost and latency wins (for example, a billion embeddings produced roughly 20× cheaper and visual models served roughly 72% cheaper). We also own core platform surfaces including the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution.

## **About the Role**

You will join a small, high-leverage team building production infrastructure for Generative AI at Deliveroo and DoorDash, with a primary focus on our open-weights model platform spanning inference and fine-tuning: real-time GPU serving, high-throughput batch inference, and model fine-tuning. You’ll work across model serving and inference engines, fine-tuning and training pipelines, GPU autoscaling and utilization, batch pipelines, backend services, and observability. This role is ideal for an engineer who enjoys pushing the cost/performance frontier of GPU inference and fine-tuning in a fast-moving technical area where product needs, model capabilities, vendor ecosystems, and cost/performance tradeoffs are evolving quickly.

## **You’re excited about this opportunity because you will…**

* Build the infrastructure that helps Deliveroo teams move GenAI ideas from prototype to production, increasing the velocity of business impact from AI across the company.
* Work on our open-weights serving stack — real-time GPU endpoints, high-throughput batch inference, and fine-tuning (SFT/DPO/LoRA) — alongside the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution.
* Design scalable, high-performance systems for model serving, batch inference, GPU autoscaling, and fine-tuning that power real customer and internal automation use cases
* Push the cost and latency frontier of GPU inference — turning batch jobs that took days into hours and cutting inference cost by multiples — while giving product teams a clean choice across open-weight and closed-source models with reliability, fallback, observability, and cost controls built in.
* Build platforms that support rapid experimentation while meeting production standards for latency, scale, monitoring, SLOs, playbooks, and operational excellence.
* Partner closely with ML engineers, product engineers, data scientists, and platform teams across DoorDash, Wolt, and Deliveroo to turn emerging GenAI capabilities into durable platform primitives.
* Shape the future of the centralized GenAI platform — including emerging directions such as reinforcement learning (RLHF/RLVR), agent optimization, and other post-training and agentic techniques — enabling the next generation of AI-powered products, agents, automation, and personalization.

## **We’re excited about you because you have…**

* BSc, MSc, or PhD in Computer Science or equivalent
* 3+ years of industry experience in software engineering
* Strong backend engineering fundamentals, especially in Python and distributed systems.
* Experience building production services, APIs, data pipelines, or ML infrastructure at scale.
* Experience operating systems in production, including observability, debugging, reliability, incident response, and performance/cost optimization.
* Hands-on experience with LLM inference and/or fine-tuning of open-weight models in production — serving (latency, throughput, batching, autoscaling, GPU utilization) and/or fine-tuning (SFT/DPO/LoRA).
* Ability to work across ambiguous, fast-moving technical areas and turn customer use cases into reusable platform capabilities
* Proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) in the full software development lifecycle, including designing, generating code, testing, monitoring and releasing software

## **Nice To Haves**

* Experience with LLM inference engines and serving frameworks (e.g., vLLM, SGLang, TensorRT-LLM) in production
* Experience with distributed/multi-node fine-tuning and training pipelines (SFT, DPO/RLHF, LoRA), including data preparation and evaluation
* GPU performance work — multi-node/distributed inference, KV-cache/memory optimization, quantization (FP8/INT8/AWQ/GPTQ), or cold-start/throughput tuning
* Experience with Kubernetes, cloud infrastructure (AWS/GCP), GPUs, serverless/elastic GPU platforms (e.g., Modal), or high-throughput batch systems
* Experience with LLM gateways, model routing, vendor abstraction, or cost attribution
* Experience building developer platforms, internal platforms, or self-serve infrastructure
* Experience building and deploying AI agents or MCP servers in production
* Experience with eval systems, LLM observability, tracing, RAG, search, or vector databases

## **Diversity, Equity and Inclusion**

At Deliveroo, we know that a great workplace reflects the world around us and that true diversity and inclusion make us stronger, more creative, and better at what we do. We’re committed to fostering an environment where everyone can do their best work and feel they belong.

We believe in equality of opportunity and welcome candidates from all backgrounds regardless of age, gender, ethnicity, disability, sexual orientation, gender identity, socio-economic background, religion, or belief.

If you have a disability or long-term health condition and need support to apply for one of our roles, or require any reasonable adjustments during the recruitment process, you’ll have the opportunity to let us know once you’ve submitted your application. We’ll share details on how to request support so we can ensure you have a fair and equitable experience.

If you’re excited about making a real impact in a fast-moving marketplace and growing your career alongside ambitious, supportive teams, we’d love to hear from you!
