Software Engineer, GenAI Platform
Backend Engineer · Mid · Full Time
Opens Deliveroo's application page
Role
What you'll do.
Join Deliveroo's GenAI Platform team to build production-grade infrastructure for generative AI products and agents serving DoorDash, Wolt, and Deliveroo. This role focuses on designing and operating scalable systems for LLM inference, fine-tuning, and GPU optimization, working across real-time serving, batch processing, and model training pipelines. You'll need 3+ years of backend engineering experience with demonstrated expertise in distributed systems, production ML infrastructure, and hands-on LLM operations to push the cost and latency frontier of GPU-accelerated workloads at scale.
Responsibilities
- Design & Build LLM Inference Infrastructure: Architect and implement production-grade inference serving systems for open-weight LLMs and VLMs including real-time GPU endpoints, request batching optimization, latency reduction strategies, and multi-model routing. Optimize systems for concurrent throughput while maintaining strict SLOs and implementing intelligent fallback mechanisms across model providers.
- Develop Fine-tuning & Training Pipelines: Build scalable distributed fine-tuning infrastructure supporting SFT, DPO, RLHF, and LoRA techniques on multi-node GPU clusters. Implement data preparation workflows, evaluation frameworks, and training orchestration systems enabling rapid model customization and optimization for Deliveroo, DoorDash, and Wolt use cases.
- Optimize GPU Utilization & Cost: Profile, benchmark, and optimize GPU workload efficiency across inference and training pipelines. Implement advanced optimization techniques including quantization (FP8/INT8/AWQ/GPTQ), KV-cache management, and autoscaling policies. Drive measurable cost reductions and latency improvements, targeting multiple order-of-magnitude improvements in cost-per-inference or training throughput.
- Build & Operate LLM Gateway: Develop centralized gateway infrastructure providing unified API access to diverse model providers and open-weight deployments. Implement model routing logic, vendor abstraction, cost attribution, SLA enforcement, and request-level observability enabling product teams to seamlessly switch between models and optimize for cost/performance tradeoffs.
- Implement Production Observability: Design comprehensive observability systems for monitoring inference performance, GPU health, pipeline status, and operational metrics. Build dashboards, alerting rules, and incident playbooks enabling rapid diagnosis and resolution of production issues. Implement cost attribution tracking providing granular visibility into consumption across teams and models.
- Develop Platform Primitives & APIs: Create clean, well-documented platform surfaces including the Agent Gateway, evals infrastructure, guardrails systems, and cost attribution tools. Design APIs balancing power and usability enabling ML engineers and product teams to build on top of the platform with safety, cost controls, and production reliability built-in.
- Collaborate on Agentic & Post-training Capabilities: Partner with ML and product teams to understand emerging GenAI use cases and translate requirements into platform capabilities. Support implementation of advanced techniques including reinforcement learning optimization (RLHF/RLVR), AI agent deployment and monitoring, and next-generation personalization features.
- Maintain System Reliability & Performance: Own operational excellence for GenAI platform infrastructure ensuring high availability, quick incident response, and continuous performance optimization. Implement SLO targets, maintain runbooks, and drive post-incident learning. Balance velocity and stability while supporting rapid experimentation and deployment cycles.
Qualifications
What we look for.
Technical
Backend Service Architecture
Proven experience designing and building production-grade backend services including API design, error handling, and integration patterns for ML infrastructure.
Distributed Systems Fundamentals
Deep understanding of distributed computing concepts including consensus, fault tolerance, consistency models, and scalability patterns applied to high-throughput systems.
LLM Inference Stack Knowledge
Hands-on production experience with the complete LLM inference stack covering request batching, KV-cache management, attention optimization, and multi-GPU coordination.
Model Fine-tuning Implementation
Direct experience implementing supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), direct preference optimization (DPO), and LoRA techniques at production scale.
GPU Optimization Expertise
Demonstrated ability to optimize GPU workloads for latency, throughput, and utilization including memory profiling, kernel optimization, and multi-GPU scaling strategies.
Production Observability
Expertise instrumenting systems with comprehensive logging, metrics, tracing, and alerting to maintain visibility into model serving performance and infrastructure health.
Education
Bachelor's Degree
BSc in Computer Science, Mathematics, Engineering, or equivalent field providing foundational knowledge in algorithms, systems design, and computational theory.
Advanced Degree (Optional)
MSc or PhD in Computer Science, Machine Learning, or related field demonstrating advanced expertise in specialized areas such as distributed systems or deep learning optimization.
Experience
3+ Years Backend Engineering
Minimum three years of professional software engineering experience with focus on backend systems, distributed architectures, and production service development.
Production ML Infrastructure
Direct experience building or operating ML infrastructure systems at scale including data pipelines, training systems, or inference platforms serving production workloads.
LLM Production Operations
Hands-on production experience with large language models spanning inference serving (latency, throughput, autoscaling, GPU utilization) and fine-tuning operations (SFT, DPO, LoRA).
System Reliability & Operations
Proven track record operating complex systems in production including incident response, performance debugging, cost optimization, and maintaining service reliability standards.
Cross-functional Collaboration
Experience working across ambiguous, rapidly evolving technical areas while translating customer use cases and requirements into scalable, reusable platform primitives.
Skills
Required
Python Programming
Expert-level proficiency in Python for building production backend services, data pipelines, and ML infrastructure. Experience with async programming, testing frameworks, and performance optimization essential for GPU inference and batch processing systems.
Distributed Systems Design
Strong understanding of distributed system architecture including load balancing, fault tolerance, consensus protocols, and scalability patterns. Must be able to design systems handling high-throughput inference and multi-node GPU coordination.
LLM Inference Operations
Hands-on production experience with large language model inference including real-time serving, latency optimization, throughput batching, GPU utilization management, and autoscaling. Proficiency with serving different model formats and sizes.
GPU Systems & CUDA
Practical experience optimizing GPU workloads including memory management, kernel utilization, multi-GPU coordination, and performance profiling. Understanding of CUDA fundamentals and GPU constraints essential for inference and fine-tuning operations.
Production Systems Reliability
Proven experience building and operating production services with emphasis on observability, monitoring, incident response, debugging, and performance optimization. Strong grasp of SLOs, alerting, and operational excellence practices.
Fine-tuning & Training Pipelines
Hands-on experience implementing and scaling fine-tuning pipelines using techniques like SFT, DPO, LoRA, and RLHF. Must understand data preparation, evaluation frameworks, distributed training coordination, and training/inference tradeoffs.
Preferred
LLM Serving Frameworks
Nice to haveProduction experience with specialized LLM serving engines like vLLM, SGLang, or TensorRT-LLM. Familiarity with request batching, KV-cache optimization, token streaming, and dynamic batching essential for high-throughput systems.
Model Quantization & Optimization
Nice to haveExperience with quantization techniques (FP8, INT8, AWQ, GPTQ) and model optimization strategies. Understanding of accuracy-latency-throughput tradeoffs in production quantized models.
Kubernetes & Container Orchestration
Nice to haveProduction Kubernetes experience including resource management, autoscaling policies, GPU resource allocation, workload scheduling, and troubleshooting cluster operations at scale.
Cloud Infrastructure Management
Nice to haveProduction-level experience with AWS or GCP including GPU instance management, cost optimization, network configuration, storage solutions, and infrastructure-as-code practices.
LLM Gateway & Routing
Nice to haveExperience building or operating LLM gateways with model routing logic, vendor abstraction layers, fallback handling, cost attribution, and request routing optimization.
AI Agent Deployment
Nice to haveExperience building, testing, and deploying AI agents or model context protocol (MCP) servers in production environments with focus on reliability and observability.
AI-Assisted Development Tools
Nice to haveProficiency with modern AI coding assistants like Claude Code, Cursor, or GitHub Copilot throughout the full development lifecycle including code generation, testing, debugging, and deployment automation.
Vector Databases & RAG Systems
Nice to haveProduction experience with vector databases, retrieval-augmented generation (RAG) systems, semantic search, and LLM observability tools for monitoring model behavior and performance.
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·GBP 140,000 – 190,000
Full posting
Original listing.
Software Engineer, GenAI Platform
About the Team
Deliveroo's GenAI Platform team sits within Machine Learning Platform and builds the shared infrastructure that helps DoorDash, Wolt, and Deliveroo teams safely bring GenAI-powered products, agents, automation, and personalization to production. Our mission is to increase the velocity of business impact from GenAI. A central pillar of that work is running frontier open-weight LLMs and VLMs (such as GLM, Qwen, Kimi, and DeepSeek) ourselves — real-time GPU serving, high-throughput batch inference, and fine-tuning on autoscaling GPUs — delivering large cost and latency wins (for example, a billion embeddings produced roughly 20× cheaper and visual models served roughly 72% cheaper). We also own core platform surfaces including the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution.
About the Role
You will join a small, high-leverage team building production infrastructure for Generative AI at Deliveroo and DoorDash, with a primary focus on our open-weights model platform spanning inference and fine-tuning: real-time GPU serving, high-throughput batch inference, and model fine-tuning. You’ll work across model serving and inference engines, fine-tuning and training pipelines, GPU autoscaling and utilization, batch pipelines, backend services, and observability. This role is ideal for an engineer who enjoys pushing the cost/performance frontier of GPU inference and fine-tuning in a fast-moving technical area where product needs, model capabilities, vendor ecosystems, and cost/performance tradeoffs are evolving quickly.
You’re excited about this opportunity because you will…
Build the infrastructure that helps Deliveroo teams move GenAI ideas from prototype to production, increasing the velocity of business impact from AI across the company.
Work on our open-weights serving stack — real-time GPU endpoints, high-throughput batch inference, and fine-tuning (SFT/DPO/LoRA) — alongside the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution.
Design scalable, high-performance systems for model serving, batch inference, GPU autoscaling, and fine-tuning that power real customer and internal automation use cases
Push the cost and latency frontier of GPU inference — turning batch jobs that took days into hours and cutting inference cost by multiples — while giving product teams a clean choice across open-weight and closed-source models with reliability, fallback, observability, and cost controls built in.
Build platforms that support rapid experimentation while meeting production standards for latency, scale, monitoring, SLOs, playbooks, and operational excellence.
Partner closely with ML engineers, product engineers, data scientists, and platform teams across DoorDash, Wolt, and Deliveroo to turn emerging GenAI capabilities into durable platform primitives.
Shape the future of the centralized GenAI platform — including emerging directions such as reinforcement learning (RLHF/RLVR), agent optimization, and other post-training and agentic techniques — enabling the next generation of AI-powered products, agents, automation, and personalization.
We’re excited about you because you have…
BSc, MSc, or PhD in Computer Science or equivalent
3+ years of industry experience in software engineering
Strong backend engineering fundamentals, especially in Python and distributed systems.
Experience building production services, APIs, data pipelines, or ML infrastructure at scale.
Experience operating systems in production, including observability, debugging, reliability, incident response, and performance/cost optimization.
Hands-on experience with LLM inference and/or fine-tuning of open-weight models in production — serving (latency, throughput, batching, autoscaling, GPU utilization) and/or fine-tuning (SFT/DPO/LoRA).
Ability to work across ambiguous, fast-moving technical areas and turn customer use cases into reusable platform capabilities
Proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) in the full software development lifecycle, including designing, generating code, testing, monitoring and releasing software
Nice To Haves
Experience with LLM inference engines and serving frameworks (e.g., vLLM, SGLang, TensorRT-LLM) in production
Experience with distributed/multi-node fine-tuning and training pipelines (SFT, DPO/RLHF, LoRA), including data preparation and evaluation
GPU performance work — multi-node/distributed inference, KV-cache/memory optimization, quantization (FP8/INT8/AWQ/GPTQ), or cold-start/throughput tuning
Experience with Kubernetes, cloud infrastructure (AWS/GCP), GPUs, serverless/elastic GPU platforms (e.g., Modal), or high-throughput batch systems
Experience with LLM gateways, model routing, vendor abstraction, or cost attribution
Experience building developer platforms, internal platforms, or self-serve infrastructure
Experience building and deploying AI agents or MCP servers in production
Experience with eval systems, LLM observability, tracing, RAG, search, or vector databases
Diversity, Equity and Inclusion
At Deliveroo, we know that a great workplace reflects the world around us and that true diversity and inclusion make us stronger, more creative, and better at what we do. We’re committed to fostering an environment where everyone can do their best work and feel they belong.
We believe in equality of opportunity and welcome candidates from all backgrounds regardless of age, gender, ethnicity, disability, sexual orientation, gender identity, socio-economic background, religion, or belief.
If you have a disability or long-term health condition and need support to apply for one of our roles, or require any reasonable adjustments during the recruitment process, you’ll have the opportunity to let us know once you’ve submitted your application. We’ll share details on how to request support so we can ensure you have a fair and equitable experience.
If you’re excited about making a real impact in a fast-moving marketplace and growing your career alongside ambitious, supportive teams, we’d love to hear from you!
Redirects to Deliveroo's application page.
Other roles
More at Deliveroo.
Software Engineer, Backend (BI Platform)
Mid
BI Engineer
Mid
Senior Software Engineer
Senior
Senior Software Engineer - Full Stack (Merchants)
Senior
Staff Software Engineer
Staff