Deliveroo

Software Engineer, GenAI Platform

Deliveroo5 days ago
Location

London - The River Building HQ

Type

Full Time

Salary

GBP 140,000 – 190,000

Level

Mid

Role

Backend Engineer

Posted

Jul 20, 2026

Full TimeMid

The role

Summary

Join Deliveroo's GenAI Platform team to build production-grade infrastructure for generative AI products and agents serving DoorDash, Wolt, and Deliveroo. This role focuses on designing and operating scalable systems for LLM inference, fine-tuning, and GPU optimization, working across real-time serving, batch processing, and model training pipelines. You'll need 3+ years of backend engineering experience with demonstrated expertise in distributed systems, production ML infrastructure, and hands-on LLM operations to push the cost and latency frontier of GPU-accelerated workloads at scale.

What you'll do

Design & Build LLM Inference Infrastructure: Architect and implement production-grade inference serving systems for open-weight LLMs and VLMs including real-time GPU endpoints, request batching optimization, latency reduction strategies, and multi-model routing. Optimize systems for concurrent throughput while maintaining strict SLOs and implementing intelligent fallback mechanisms across model providers.
Develop Fine-tuning & Training Pipelines: Build scalable distributed fine-tuning infrastructure supporting SFT, DPO, RLHF, and LoRA techniques on multi-node GPU clusters. Implement data preparation workflows, evaluation frameworks, and training orchestration systems enabling rapid model customization and optimization for Deliveroo, DoorDash, and Wolt use cases.
Optimize GPU Utilization & Cost: Profile, benchmark, and optimize GPU workload efficiency across inference and training pipelines. Implement advanced optimization techniques including quantization (FP8/INT8/AWQ/GPTQ), KV-cache management, and autoscaling policies. Drive measurable cost reductions and latency improvements, targeting multiple order-of-magnitude improvements in cost-per-inference or training throughput.
Build & Operate LLM Gateway: Develop centralized gateway infrastructure providing unified API access to diverse model providers and open-weight deployments. Implement model routing logic, vendor abstraction, cost attribution, SLA enforcement, and request-level observability enabling product teams to seamlessly switch between models and optimize for cost/performance tradeoffs.
Implement Production Observability: Design comprehensive observability systems for monitoring inference performance, GPU health, pipeline status, and operational metrics. Build dashboards, alerting rules, and incident playbooks enabling rapid diagnosis and resolution of production issues. Implement cost attribution tracking providing granular visibility into consumption across teams and models.
Develop Platform Primitives & APIs: Create clean, well-documented platform surfaces including the Agent Gateway, evals infrastructure, guardrails systems, and cost attribution tools. Design APIs balancing power and usability enabling ML engineers and product teams to build on top of the platform with safety, cost controls, and production reliability built-in.
Collaborate on Agentic & Post-training Capabilities: Partner with ML and product teams to understand emerging GenAI use cases and translate requirements into platform capabilities. Support implementation of advanced techniques including reinforcement learning optimization (RLHF/RLVR), AI agent deployment and monitoring, and next-generation personalization features.
Maintain System Reliability & Performance: Own operational excellence for GenAI platform infrastructure ensuring high availability, quick incident response, and continuous performance optimization. Implement SLO targets, maintain runbooks, and drive post-incident learning. Balance velocity and stability while supporting rapid experimentation and deployment cycles.

What we look for

Technical

Backend Service ArchitectureProven experience designing and building production-grade backend services including API design, error handling, and integration patterns for ML infrastructure.
Distributed Systems FundamentalsDeep understanding of distributed computing concepts including consensus, fault tolerance, consistency models, and scalability patterns applied to high-throughput systems.
LLM Inference Stack KnowledgeHands-on production experience with the complete LLM inference stack covering request batching, KV-cache management, attention optimization, and multi-GPU coordination.
Model Fine-tuning ImplementationDirect experience implementing supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), direct preference optimization (DPO), and LoRA techniques at production scale.
GPU Optimization ExpertiseDemonstrated ability to optimize GPU workloads for latency, throughput, and utilization including memory profiling, kernel optimization, and multi-GPU scaling strategies.
Production ObservabilityExpertise instrumenting systems with comprehensive logging, metrics, tracing, and alerting to maintain visibility into model serving performance and infrastructure health.

Education

Bachelor's DegreeBSc in Computer Science, Mathematics, Engineering, or equivalent field providing foundational knowledge in algorithms, systems design, and computational theory.
Advanced Degree (Optional)MSc or PhD in Computer Science, Machine Learning, or related field demonstrating advanced expertise in specialized areas such as distributed systems or deep learning optimization.

Experience

3+ Years Backend EngineeringMinimum three years of professional software engineering experience with focus on backend systems, distributed architectures, and production service development.
Production ML InfrastructureDirect experience building or operating ML infrastructure systems at scale including data pipelines, training systems, or inference platforms serving production workloads.
LLM Production OperationsHands-on production experience with large language models spanning inference serving (latency, throughput, autoscaling, GPU utilization) and fine-tuning operations (SFT, DPO, LoRA).
System Reliability & OperationsProven track record operating complex systems in production including incident response, performance debugging, cost optimization, and maintaining service reliability standards.
Cross-functional CollaborationExperience working across ambiguous, rapidly evolving technical areas while translating customer use cases and requirements into scalable, reusable platform primitives.

Skills

Required skills

Python ProgrammingExpert-level proficiency in Python for building production backend services, data pipelines, and ML infrastructure. Experience with async programming, testing frameworks, and performance optimization essential for GPU inference and batch processing systems.
Distributed Systems DesignStrong understanding of distributed system architecture including load balancing, fault tolerance, consensus protocols, and scalability patterns. Must be able to design systems handling high-throughput inference and multi-node GPU coordination.
LLM Inference OperationsHands-on production experience with large language model inference including real-time serving, latency optimization, throughput batching, GPU utilization management, and autoscaling. Proficiency with serving different model formats and sizes.
GPU Systems & CUDAPractical experience optimizing GPU workloads including memory management, kernel utilization, multi-GPU coordination, and performance profiling. Understanding of CUDA fundamentals and GPU constraints essential for inference and fine-tuning operations.
Production Systems ReliabilityProven experience building and operating production services with emphasis on observability, monitoring, incident response, debugging, and performance optimization. Strong grasp of SLOs, alerting, and operational excellence practices.
Fine-tuning & Training PipelinesHands-on experience implementing and scaling fine-tuning pipelines using techniques like SFT, DPO, LoRA, and RLHF. Must understand data preparation, evaluation frameworks, distributed training coordination, and training/inference tradeoffs.

Nice to have

LLM Serving FrameworksProduction experience with specialized LLM serving engines like vLLM, SGLang, or TensorRT-LLM. Familiarity with request batching, KV-cache optimization, token streaming, and dynamic batching essential for high-throughput systems.
Model Quantization & OptimizationExperience with quantization techniques (FP8, INT8, AWQ, GPTQ) and model optimization strategies. Understanding of accuracy-latency-throughput tradeoffs in production quantized models.
Kubernetes & Container OrchestrationProduction Kubernetes experience including resource management, autoscaling policies, GPU resource allocation, workload scheduling, and troubleshooting cluster operations at scale.
Cloud Infrastructure ManagementProduction-level experience with AWS or GCP including GPU instance management, cost optimization, network configuration, storage solutions, and infrastructure-as-code practices.
LLM Gateway & RoutingExperience building or operating LLM gateways with model routing logic, vendor abstraction layers, fallback handling, cost attribution, and request routing optimization.
AI Agent DeploymentExperience building, testing, and deploying AI agents or model context protocol (MCP) servers in production environments with focus on reliability and observability.
AI-Assisted Development ToolsProficiency with modern AI coding assistants like Claude Code, Cursor, or GitHub Copilot throughout the full development lifecycle including code generation, testing, debugging, and deployment automation.
Vector Databases & RAG SystemsProduction experience with vector databases, retrieval-augmented generation (RAG) systems, semantic search, and LLM observability tools for monitoring model behavior and performance.

Compensation & benefits

Salary

GBP 140,000 – 190,000 (annual)


Apply for this position

You'll be redirected to the company's application page


Deliveroo

Deliveroo

View all jobs

Deliveroo is a British multinational online food delivery company operating a platform for ordering from restaurants and grocers.

London, England, United KingdomFounded 2013deliveroo.co.uk

Tech Stack

Languages
Python
Frameworks
vLLMSGLangTensorRT-LLMFastAPIRay
Databases
Vector Databases (Pinecone/Weaviate/Milvus)PostgreSQL
Tools
KubernetesAWS/GCPPrometheus & GrafanaTerraform/CloudFormationDockerPyTorch/Hugging Face Transformers
Other
CUDA/GPU ProgrammingDistributed Training & InferenceModel QuantizationCost & Performance Optimization

Interview Guides

12 guides available for Deliveroo

Apply Now