# Software Engineer, GenAI Platform
**Company:** [Deliveroo](https://scaleengineer.com/companies/deliveroo)
Join Deliveroo's GenAI Platform team to build production-grade infrastructure for generative AI products and agents serving DoorDash, Wolt, and Deliveroo. This role focuses on designing and operating scalable systems for LLM inference, fine-tuning, and GPU optimization, working across real-time serving, batch processing, and model training pipelines. You'll need 3+ years of backend engineering experience with demonstrated expertise in distributed systems, production ML infrastructure, and hands-on LLM operations to push the cost and latency frontier of GPU-accelerated workloads at scale.
**Role:** Backend Engineer
**Seniority:** Mid
**Locations:** London - The River Building HQ
**Salary:** 140000–190000 GBP
[Apply](https://jobs.ashbyhq.com/deliveroo/78c5338a-6a37-4f5c-91c3-02707cab0726)
Canonical: https://scaleengineer.com/jobs/deliveroo/software-engineer-genai-platform
---
## Responsibilities

- Design & Build LLM Inference Infrastructure: Architect and implement production-grade inference serving systems for open-weight LLMs and VLMs including real-time GPU endpoints, request batching optimization, latency reduction strategies, and multi-model routing. Optimize systems for concurrent throughput while maintaining strict SLOs and implementing intelligent fallback mechanisms across model providers.
- Develop Fine-tuning & Training Pipelines: Build scalable distributed fine-tuning infrastructure supporting SFT, DPO, RLHF, and LoRA techniques on multi-node GPU clusters. Implement data preparation workflows, evaluation frameworks, and training orchestration systems enabling rapid model customization and optimization for Deliveroo, DoorDash, and Wolt use cases.
- Optimize GPU Utilization & Cost: Profile, benchmark, and optimize GPU workload efficiency across inference and training pipelines. Implement advanced optimization techniques including quantization (FP8/INT8/AWQ/GPTQ), KV-cache management, and autoscaling policies. Drive measurable cost reductions and latency improvements, targeting multiple order-of-magnitude improvements in cost-per-inference or training throughput.
- Build & Operate LLM Gateway: Develop centralized gateway infrastructure providing unified API access to diverse model providers and open-weight deployments. Implement model routing logic, vendor abstraction, cost attribution, SLA enforcement, and request-level observability enabling product teams to seamlessly switch between models and optimize for cost/performance tradeoffs.
- Implement Production Observability: Design comprehensive observability systems for monitoring inference performance, GPU health, pipeline status, and operational metrics. Build dashboards, alerting rules, and incident playbooks enabling rapid diagnosis and resolution of production issues. Implement cost attribution tracking providing granular visibility into consumption across teams and models.
- Develop Platform Primitives & APIs: Create clean, well-documented platform surfaces including the Agent Gateway, evals infrastructure, guardrails systems, and cost attribution tools. Design APIs balancing power and usability enabling ML engineers and product teams to build on top of the platform with safety, cost controls, and production reliability built-in.
- Collaborate on Agentic & Post-training Capabilities: Partner with ML and product teams to understand emerging GenAI use cases and translate requirements into platform capabilities. Support implementation of advanced techniques including reinforcement learning optimization (RLHF/RLVR), AI agent deployment and monitoring, and next-generation personalization features.
- Maintain System Reliability & Performance: Own operational excellence for GenAI platform infrastructure ensuring high availability, quick incident response, and continuous performance optimization. Implement SLO targets, maintain runbooks, and drive post-incident learning. Balance velocity and stability while supporting rapid experimentation and deployment cycles.

## Requirements

### education

- {"name":"Bachelor's Degree","description":"BSc in Computer Science, Mathematics, Engineering, or equivalent field providing foundational knowledge in algorithms, systems design, and computational theory."}
- {"name":"Advanced Degree (Optional)","description":"MSc or PhD in Computer Science, Machine Learning, or related field demonstrating advanced expertise in specialized areas such as distributed systems or deep learning optimization."}

### technical

- {"name":"Backend Service Architecture","description":"Proven experience designing and building production-grade backend services including API design, error handling, and integration patterns for ML infrastructure."}
- {"name":"Distributed Systems Fundamentals","description":"Deep understanding of distributed computing concepts including consensus, fault tolerance, consistency models, and scalability patterns applied to high-throughput systems."}
- {"name":"LLM Inference Stack Knowledge","description":"Hands-on production experience with the complete LLM inference stack covering request batching, KV-cache management, attention optimization, and multi-GPU coordination."}
- {"name":"Model Fine-tuning Implementation","description":"Direct experience implementing supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), direct preference optimization (DPO), and LoRA techniques at production scale."}
- {"name":"GPU Optimization Expertise","description":"Demonstrated ability to optimize GPU workloads for latency, throughput, and utilization including memory profiling, kernel optimization, and multi-GPU scaling strategies."}
- {"name":"Production Observability","description":"Expertise instrumenting systems with comprehensive logging, metrics, tracing, and alerting to maintain visibility into model serving performance and infrastructure health."}

### experience

- {"name":"3+ Years Backend Engineering","description":"Minimum three years of professional software engineering experience with focus on backend systems, distributed architectures, and production service development."}
- {"name":"Production ML Infrastructure","description":"Direct experience building or operating ML infrastructure systems at scale including data pipelines, training systems, or inference platforms serving production workloads."}
- {"name":"LLM Production Operations","description":"Hands-on production experience with large language models spanning inference serving (latency, throughput, autoscaling, GPU utilization) and fine-tuning operations (SFT, DPO, LoRA)."}
- {"name":"System Reliability & Operations","description":"Proven track record operating complex systems in production including incident response, performance debugging, cost optimization, and maintaining service reliability standards."}
- {"name":"Cross-functional Collaboration","description":"Experience working across ambiguous, rapidly evolving technical areas while translating customer use cases and requirements into scalable, reusable platform primitives."}

## Skills

### required

- {"name":"Python Programming","description":"Expert-level proficiency in Python for building production backend services, data pipelines, and ML infrastructure. Experience with async programming, testing frameworks, and performance optimization essential for GPU inference and batch processing systems."}
- {"name":"Distributed Systems Design","description":"Strong understanding of distributed system architecture including load balancing, fault tolerance, consensus protocols, and scalability patterns. Must be able to design systems handling high-throughput inference and multi-node GPU coordination."}
- {"name":"LLM Inference Operations","description":"Hands-on production experience with large language model inference including real-time serving, latency optimization, throughput batching, GPU utilization management, and autoscaling. Proficiency with serving different model formats and sizes."}
- {"name":"GPU Systems & CUDA","description":"Practical experience optimizing GPU workloads including memory management, kernel utilization, multi-GPU coordination, and performance profiling. Understanding of CUDA fundamentals and GPU constraints essential for inference and fine-tuning operations."}
- {"name":"Production Systems Reliability","description":"Proven experience building and operating production services with emphasis on observability, monitoring, incident response, debugging, and performance optimization. Strong grasp of SLOs, alerting, and operational excellence practices."}
- {"name":"Fine-tuning & Training Pipelines","description":"Hands-on experience implementing and scaling fine-tuning pipelines using techniques like SFT, DPO, LoRA, and RLHF. Must understand data preparation, evaluation frameworks, distributed training coordination, and training/inference tradeoffs."}

### preferred

- {"name":"LLM Serving Frameworks","description":"Production experience with specialized LLM serving engines like vLLM, SGLang, or TensorRT-LLM. Familiarity with request batching, KV-cache optimization, token streaming, and dynamic batching essential for high-throughput systems."}
- {"name":"Model Quantization & Optimization","description":"Experience with quantization techniques (FP8, INT8, AWQ, GPTQ) and model optimization strategies. Understanding of accuracy-latency-throughput tradeoffs in production quantized models."}
- {"name":"Kubernetes & Container Orchestration","description":"Production Kubernetes experience including resource management, autoscaling policies, GPU resource allocation, workload scheduling, and troubleshooting cluster operations at scale."}
- {"name":"Cloud Infrastructure Management","description":"Production-level experience with AWS or GCP including GPU instance management, cost optimization, network configuration, storage solutions, and infrastructure-as-code practices."}
- {"name":"LLM Gateway & Routing","description":"Experience building or operating LLM gateways with model routing logic, vendor abstraction layers, fallback handling, cost attribution, and request routing optimization."}
- {"name":"AI Agent Deployment","description":"Experience building, testing, and deploying AI agents or model context protocol (MCP) servers in production environments with focus on reliability and observability."}
- {"name":"AI-Assisted Development Tools","description":"Proficiency with modern AI coding assistants like Claude Code, Cursor, or GitHub Copilot throughout the full development lifecycle including code generation, testing, debugging, and deployment automation."}
- {"name":"Vector Databases & RAG Systems","description":"Production experience with vector databases, retrieval-augmented generation (RAG) systems, semantic search, and LLM observability tools for monitoring model behavior and performance."}

## Tech stack

### tools

- {"name":"Kubernetes","description":"Container orchestration platform for deploying, scaling, and managing GPU workloads across multi-node clusters with automated resource allocation and autoscaling."}
- {"name":"AWS/GCP","description":"Cloud infrastructure platforms for GPU instance management, compute resource provisioning, cost optimization, and elastic scaling of inference and training workloads."}
- {"name":"Prometheus & Grafana","description":"Monitoring and observability stack for tracking system metrics, GPU utilization, inference latency, throughput, and cost attribution across the platform."}
- {"name":"Terraform/CloudFormation","description":"Infrastructure-as-code tools for managing cloud resources, GPU cluster configuration, autoscaling policies, and reproducible infrastructure deployment."}
- {"name":"Docker","description":"Containerization platform for packaging inference servers, fine-tuning pipelines, and platform services with consistent environments across deployments."}
- {"name":"PyTorch/Hugging Face Transformers","description":"Deep learning frameworks and model libraries for implementing fine-tuning pipelines, model loading, and training operations."}

### others

- {"name":"CUDA/GPU Programming","description":"Low-level GPU programming for optimizing inference kernels, memory management, and performance tuning of deep learning workloads."}
- {"name":"Distributed Training & Inference","description":"Multi-node GPU coordination including data parallelism, tensor parallelism, pipeline parallelism for scaling both training and inference across clusters."}
- {"name":"Model Quantization","description":"Techniques for reducing model size and latency through FP8, INT8, AWQ, and GPTQ quantization while maintaining accuracy in production deployments."}
- {"name":"Cost & Performance Optimization","description":"Expertise in profiling, benchmarking, and optimizing GPU utilization, inference latency, batch throughput, and total cost of ownership for large-scale model serving."}

### databases

- {"name":"Vector Databases (Pinecone/Weaviate/Milvus)","description":"Storage and retrieval systems for embeddings supporting RAG pipelines and semantic search functionality within GenAI applications."}
- {"name":"PostgreSQL","description":"Relational database for metadata storage, cost attribution tracking, job scheduling, and operational data in the GenAI platform."}

### languages

- {"name":"Python","description":"Primary language for backend services, ML infrastructure, batch processing, and GPU workload management. Essential for writing inference servers, fine-tuning pipelines, and automation scripts."}

### frameworks

- {"name":"vLLM","description":"High-performance LLM serving framework for real-time inference with dynamic batching, optimized KV-cache management, and multi-GPU/multi-node support."}
- {"name":"SGLang","description":"Structured generation framework for efficient LLM serving with complex output constraints and efficient token scheduling."}
- {"name":"TensorRT-LLM","description":"NVIDIA's optimized inference library for quantized and optimized LLM serving on NVIDIA hardware with advanced kernel optimization."}
- {"name":"FastAPI","description":"Modern Python web framework for building high-performance APIs powering inference endpoints, gateways, and platform services."}
- {"name":"Ray","description":"Distributed computing framework for scaling LLM inference, fine-tuning jobs, and batch processing across multi-node GPU clusters."}

## Benefits

### benefits

## Compensation

- **max:** 0
- **min:** 0
- **currency:** 
- **stockOptions:** false

## Interview process

### steps

## Full description
# **Software Engineer, GenAI Platform**

## **About the Team**

Deliveroo's GenAI Platform team sits within Machine Learning Platform and builds the shared infrastructure that helps DoorDash, Wolt, and Deliveroo teams safely bring GenAI-powered products, agents, automation, and personalization to production. Our mission is to increase the velocity of business impact from GenAI. A central pillar of that work is running frontier open-weight LLMs and VLMs (such as GLM, Qwen, Kimi, and DeepSeek) ourselves — real-time GPU serving, high-throughput batch inference, and fine-tuning on autoscaling GPUs — delivering large cost and latency wins (for example, a billion embeddings produced roughly 20× cheaper and visual models served roughly 72% cheaper). We also own core platform surfaces including the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution.

## **About the Role**

You will join a small, high-leverage team building production infrastructure for Generative AI at Deliveroo and DoorDash, with a primary focus on our open-weights model platform spanning inference and fine-tuning: real-time GPU serving, high-throughput batch inference, and model fine-tuning. You’ll work across model serving and inference engines, fine-tuning and training pipelines, GPU autoscaling and utilization, batch pipelines, backend services, and observability. This role is ideal for an engineer who enjoys pushing the cost/performance frontier of GPU inference and fine-tuning in a fast-moving technical area where product needs, model capabilities, vendor ecosystems, and cost/performance tradeoffs are evolving quickly.

## **You’re excited about this opportunity because you will…**

* Build the infrastructure that helps Deliveroo teams move GenAI ideas from prototype to production, increasing the velocity of business impact from AI across the company.
* Work on our open-weights serving stack — real-time GPU endpoints, high-throughput batch inference, and fine-tuning (SFT/DPO/LoRA) — alongside the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution.
* Design scalable, high-performance systems for model serving, batch inference, GPU autoscaling, and fine-tuning that power real customer and internal automation use cases
* Push the cost and latency frontier of GPU inference — turning batch jobs that took days into hours and cutting inference cost by multiples — while giving product teams a clean choice across open-weight and closed-source models with reliability, fallback, observability, and cost controls built in.
* Build platforms that support rapid experimentation while meeting production standards for latency, scale, monitoring, SLOs, playbooks, and operational excellence.
* Partner closely with ML engineers, product engineers, data scientists, and platform teams across DoorDash, Wolt, and Deliveroo to turn emerging GenAI capabilities into durable platform primitives.
* Shape the future of the centralized GenAI platform — including emerging directions such as reinforcement learning (RLHF/RLVR), agent optimization, and other post-training and agentic techniques — enabling the next generation of AI-powered products, agents, automation, and personalization.

## **We’re excited about you because you have…**

* BSc, MSc, or PhD in Computer Science or equivalent
* 3+ years of industry experience in software engineering
* Strong backend engineering fundamentals, especially in Python and distributed systems.
* Experience building production services, APIs, data pipelines, or ML infrastructure at scale.
* Experience operating systems in production, including observability, debugging, reliability, incident response, and performance/cost optimization.
* Hands-on experience with LLM inference and/or fine-tuning of open-weight models in production — serving (latency, throughput, batching, autoscaling, GPU utilization) and/or fine-tuning (SFT/DPO/LoRA).
* Ability to work across ambiguous, fast-moving technical areas and turn customer use cases into reusable platform capabilities
* Proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) in the full software development lifecycle, including designing, generating code, testing, monitoring and releasing software

## **Nice To Haves**

* Experience with LLM inference engines and serving frameworks (e.g., vLLM, SGLang, TensorRT-LLM) in production
* Experience with distributed/multi-node fine-tuning and training pipelines (SFT, DPO/RLHF, LoRA), including data preparation and evaluation
* GPU performance work — multi-node/distributed inference, KV-cache/memory optimization, quantization (FP8/INT8/AWQ/GPTQ), or cold-start/throughput tuning
* Experience with Kubernetes, cloud infrastructure (AWS/GCP), GPUs, serverless/elastic GPU platforms (e.g., Modal), or high-throughput batch systems
* Experience with LLM gateways, model routing, vendor abstraction, or cost attribution
* Experience building developer platforms, internal platforms, or self-serve infrastructure
* Experience building and deploying AI agents or MCP servers in production
* Experience with eval systems, LLM observability, tracing, RAG, search, or vector databases

## **Diversity, Equity and Inclusion**

At Deliveroo, we know that a great workplace reflects the world around us and that true diversity and inclusion make us stronger, more creative, and better at what we do. We’re committed to fostering an environment where everyone can do their best work and feel they belong.

We believe in equality of opportunity and welcome candidates from all backgrounds regardless of age, gender, ethnicity, disability, sexual orientation, gender identity, socio-economic background, religion, or belief.

If you have a disability or long-term health condition and need support to apply for one of our roles, or require any reasonable adjustments during the recruitment process, you’ll have the opportunity to let us know once you’ve submitted your application. We’ll share details on how to request support so we can ensure you have a fair and equitable experience.

If you’re excited about making a real impact in a fast-moving marketplace and growing your career alongside ambitious, supportive teams, we’d love to hear from you!
