Deliveroo

Software Engineer, Machine Learning Infrastructure

Deliveroo5 days ago
Location

London - The River Building HQ

Type

Full Time

Salary

GBP 140,000 – 200,000

Level

Mid

Role

Machine Learning Infrastructure Engineer

Posted

Jul 20, 2026

Full TimeMid

The role

Summary

Join Deliveroo's GenAI Platform team to build production-grade infrastructure for generative AI across DoorDash, Wolt, and Deliveroo. This role focuses on scaling open-weight LLM and VLM infrastructure, including real-time GPU serving, high-throughput batch inference, and distributed fine-tuning pipelines while optimizing cost and latency. You'll design systems that power AI agents, automation, and personalization while working across model serving frameworks, GPU autoscaling, backend services, and observability in a fast-moving technical environment.

What you'll do

Design and build production GPU serving infrastructure: Architect and implement real-time GPU serving endpoints, high-throughput batch inference pipelines, and autoscaling systems for open-weight LLMs and VLMs. Optimize for cost, latency, and throughput while handling multi-model deployments across inference engines like vLLM and SGLang.
Develop and optimize model fine-tuning infrastructure: Build distributed fine-tuning and training pipelines supporting SFT, DPO, RLHF, and LoRA techniques on autoscaling GPUs. Handle data preparation, evaluation frameworks, and production-ready training orchestration that supports rapid experimentation.
Implement GPU autoscaling and resource optimization: Engineer intelligent GPU autoscaling systems, optimize GPU utilization rates, and implement cost attribution mechanisms. Focus on KV-cache optimization, quantization strategies (FP8/INT8/AWQ/GPTQ), and multi-node distributed inference patterns.
Own platform gateway and observability systems: Develop and maintain LLM Gateway and Agent Gateway components, implement comprehensive observability including tracing and monitoring, build evals infrastructure and guardrails, and establish production reliability standards including SLOs and incident playbooks.
Partner across engineering organizations: Collaborate closely with ML engineers, product engineers, data scientists, and platform teams across DoorDash, Wolt, and Deliveroo to translate business requirements into scalable platform primitives and infrastructure abstractions.
Push cost and performance frontiers: Continuously optimize inference and fine-tuning systems to deliver cost and latency wins, turning days-long batch jobs into hours while reducing inference costs by multiples. Evaluate emerging model architectures, vendor offerings, and optimization techniques.
Establish production excellence standards: Build platforms that support rapid experimentation while maintaining production standards for reliability, performance, monitoring, and operational excellence. Develop playbooks, runbooks, and automated remediation capabilities.

What we look for

Technical

Backend engineering in PythonStrong proficiency in Python for building scalable backend services, distributed systems, and data pipelines. Experience designing APIs, handling concurrency, and optimizing performance for production workloads.
Distributed systems architectureDeep understanding of distributed systems concepts including consensus, partitioning, replication, and orchestration. Experience designing systems that scale horizontally across multiple nodes and handle network failures gracefully.
Production LLM serving and inferenceHands-on production experience with LLM inference, model serving frameworks (vLLM, SGLang, TensorRT-LLM), and optimization techniques including batching, autoscaling, KV-cache management, and quantization methods.
GPU computing and CUDA fundamentalsUnderstanding of GPU architecture, memory hierarchies, CUDA programming basics, and GPU-specific optimization strategies for machine learning workloads including distributed multi-GPU training and inference.
Production observability and debuggingProficiency in production monitoring, logging, tracing, and debugging complex distributed systems. Experience with metrics collection, anomaly detection, and performance profiling at scale.
Open-weight model fine-tuningProduction experience with fine-tuning open-weight models including supervised fine-tuning (SFT), DPO, RLHF, and LoRA techniques. Understanding of training dynamics, convergence, and evaluation methodologies.
Cloud infrastructure and containerizationExperience with Kubernetes orchestration, cloud platforms (AWS/GCP), containerization strategies, and serverless/elastic GPU platforms. Knowledge of infrastructure-as-code and deployment automation.

Education

Bachelor's degree in Computer ScienceBSc in Computer Science, Computer Engineering, or equivalent discipline providing foundational knowledge in algorithms, data structures, and systems design.
Master's degree or higher (preferred)MSc or PhD in Computer Science, Machine Learning, or related field demonstrating advanced research capabilities and deep theoretical understanding of distributed systems or machine learning infrastructure.

Experience

3+ years software engineering industry experienceMinimum three years of professional software engineering experience building production systems, preferably in backend infrastructure, platform engineering, or machine learning systems.
Production-scale data infrastructureTrack record of building and operating production services, APIs, data pipelines, or ML infrastructure serving significant scale. Demonstrated ability to handle reliability, performance optimization, and operational excellence.
Production system operations experienceReal-world experience operating systems in production including incident response, performance optimization, cost optimization, debugging complex failures, and establishing reliability standards.
AI coding tools proficiencyProficiency using modern AI coding assistants (Claude, Codex, Cursor) throughout the full software development lifecycle including design, code generation, testing, monitoring, and deployment.

Skills

Required skills

Python backend developmentExpert-level Python programming for building scalable backend services, production APIs, and data processing pipelines with focus on performance and reliability.
Distributed systems designAbility to architect and implement distributed systems with consideration for scalability, fault tolerance, consistency models, and operational complexity.
Production LLM infrastructurePractical experience building or operating LLM serving infrastructure in production environments, including inference optimization, model routing, and resource management.
GPU and CUDA fundamentalsWorking knowledge of GPU architecture, CUDA programming basics, and how to optimize machine learning computations for GPU execution and memory efficiency.
Cloud infrastructure (AWS/GCP)Hands-on experience deploying and managing infrastructure on AWS or GCP including compute services, networking, storage, and cost optimization.
Kubernetes and container orchestrationProficiency with Kubernetes for container orchestration, service deployment, resource management, and production operations.
Systems observabilityExperience implementing monitoring, logging, tracing, and alerting for production systems. Ability to debug performance issues and establish SLOs.
Fine-tuning techniques for LLMsPractical knowledge of supervised fine-tuning (SFT), direct preference optimization (DPO), reinforcement learning from human feedback (RLHF), and LoRA-based adaptation methods.

Nice to have

vLLM and SGLang framework experienceProduction experience with vLLM, SGLang, or TensorRT-LLM inference engines for optimizing LLM serving performance, batching, and multi-GPU inference.
LLM quantization techniquesExperience with quantization methods including FP8, INT8, AWQ, GPTQ, and other techniques for reducing model size and inference latency while maintaining accuracy.
Multi-node distributed trainingExperience implementing or operating distributed training systems using frameworks like PyTorch Distributed, DeepSpeed, or Ray for large-scale fine-tuning and training.
Vector databases and semantic searchFamiliarity with vector databases (Pinecone, Weaviate, Milvus), embedding systems, and RAG architectures for retrieval-augmented generation applications.
LLM evaluation and evals infrastructureExperience building evaluation frameworks, running LLM benchmarks, and establishing metrics for model quality assessment.
AI agents and MCP serversProduction experience building AI agents, orchestration frameworks, or Model Context Protocol (MCP) servers that coordinate multiple tools and models.
API gateway and routing systemsExperience building or operating API gateways, load balancers, or intelligent routing systems that handle vendor abstraction and dynamic routing decisions.
Cost attribution and meteringExperience implementing cost tracking, usage metering, and attribution systems for multi-tenant infrastructure serving business units.
Serverless and elastic GPU platformsExperience with serverless computing platforms or elastic GPU services (Modal, Anyscale, Lambda Labs) for dynamic workload scaling.
Performance profiling and optimizationAdvanced skills in profiling tools, identifying bottlenecks in GPU code, and implementing targeted optimizations for latency and throughput.

Compensation & benefits

Salary

GBP 140,000 – 200,000 (annual)

Stock options

Available

Benefits

Equity and stock options

Participate in Deliveroo's growth through competitive equity grants allowing you to build long-term wealth alongside the company.

Comprehensive health and wellness benefits

Health insurance coverage including medical, dental, and vision plans with options for you and your family members.

Flexible working arrangements

Remote work flexibility and flexible scheduling to support work-life balance while collaborating across distributed teams.

Professional development budget

Annual learning and development budget for conferences, courses, certifications, and training programs to grow your technical expertise.

Generous time off policy

Competitive paid time off including vacation days, sick leave, and parental leave to ensure adequate rest and personal time.

Pension and retirement planning

Employer-contributed pension scheme supporting long-term financial security and retirement planning.

Mental health and wellbeing support

Access to mental health resources, counseling services, and wellness programs supporting holistic employee wellbeing.

Parental leave and family support

Comprehensive parental leave policies for various family situations and childcare support options.

Commuter benefits and transportation

Support for transportation costs including commuter passes and transit benefits for employees commuting to offices.

Diversity and inclusion initiatives

Active commitment to fostering an inclusive workplace with employee resource groups, diversity programs, and inclusive hiring practices.


Interview process

  1. 1
    Initial screening call 30-minute phone screening with a recruiter covering your background, experience with ML infrastructure, and initial technical alignment with the role requirements.
  2. 2
    Technical screening interview 60-90 minute technical conversation with an ML Infrastructure engineer focusing on your production experience with distributed systems, LLM serving, and GPU infrastructure. Expect discussion of architectural decisions and optimization techniques.
  3. 3
    System design deep dive 90-minute interview centered on designing large-scale ML infrastructure systems. You'll discuss tradeoffs in model serving architectures, GPU autoscaling strategies, and how you'd approach cost/performance optimization challenges.
  4. 4
    Production experience panel Interview with multiple engineers on the GenAI Platform team exploring your hands-on experience operating production systems, incident response capabilities, and how you approach reliability and observability.
  5. 5
    Cross-functional collaboration discussion Conversation with stakeholders from product, data science, or adjacent platform teams to assess your ability to translate business requirements into platform abstractions and work effectively across organizations.
  6. 6
    Final conversation with hiring manager Discussion with the engineering manager covering team dynamics, long-term career goals, and fit for Deliveroo's culture of technical excellence and operational ownership.

Apply for this position

You'll be redirected to the company's application page


Deliveroo

Deliveroo

View all jobs

Deliveroo is a British multinational online food delivery company operating a platform for ordering from restaurants and grocers.

London, England, United KingdomFounded 2013deliveroo.co.uk

Tech Stack

Languages
PythonSQLCUDA/C++
Frameworks
vLLMSGLangTensorRT-LLMPyTorch DistributedDeepSpeedFastAPIRay
Databases
PostgreSQLRedisVector databases (Pinecone/Weaviate/Milvus)TimescaleDB
Tools
KubernetesAWS (EC2, S3, SageMaker)GCP (Compute Engine, Cloud Storage, TPUs)Prometheus & GrafanaJaeger/DatadogDockerModalGit/GitHub
Other
GPU profiling and optimization toolsAI coding assistantsBash and shell scriptingTerraform/Infrastructure-as-CodeModel evaluation frameworks

Interview Guides

12 guides available for Deliveroo

Apply Now