Senior Software Engineer, GenAI Platform

Senior Software Engineer · Senior · Full Time

London - The River Building HQGBP 180k – 280k2w ago
Apply for this role

Opens Deliveroo's application page

Role

What you'll do.

Join Deliveroo's GenAI Platform team as a Senior Software Engineer to architect and lead production infrastructure for generative AI serving over 500 million users across DoorDash, Wolt, and Deliveroo. You'll design and own large-scale GPU inference and fine-tuning systems, optimize cost and latency of open-weight LLM serving, and mentor engineers while setting technical direction for enterprise-grade machine learning infrastructure. This role demands deep expertise in distributed systems, LLM inference optimization, and production ML infrastructure at scale.

Responsibilities

  • Lead Infrastructure Architecture for GenAI Platform: Design and architect scalable, production-grade infrastructure for real-time GPU serving, high-throughput batch inference, and model fine-tuning (SFT/DPO/LoRA). Establish technical standards and best practices for serving open-weight LLMs and VLMs (GLM, Qwen, Kimi, DeepSeek) while maintaining reliability, observability, and cost efficiency at scale across multiple billion-request production systems.
  • Own and Evolve Model Serving Stack: Lead the evolution of the open-weights serving platform including real-time GPU endpoints, batch inference pipelines, and fine-tuning infrastructure. Manage the LLM Gateway, Agent Gateway, evaluation infrastructure, guardrails, and cost attribution systems. Drive technical decisions on model serving frameworks, inference engines, and vendor selection to deliver 20-72% cost reductions and sub-100ms latency improvements.
  • Optimize Cost and Latency at the GPU Frontier: Push production GPU inference and fine-tuning cost-performance boundaries through KV-cache optimization, quantization strategies (FP8/INT8/AWQ/GPTQ), multi-node distributed inference, autoscaling tuning, and batch scheduling. Transform multi-day batch jobs into hours while maintaining SLOs and providing product teams with reliable fallback strategies and real-time cost controls.
  • Design Systems for Rapid Experimentation and Production Excellence: Build platforms enabling fast ML experimentation cycles while meeting strict production standards for latency, throughput, monitoring, SLOs, incident playbooks, and operational excellence. Implement comprehensive observability, tracing, and debugging tooling to support end-to-end visibility across model serving, inference pipelines, and GPU resource utilization.
  • Cross-Functional Technical Leadership and Mentoring: Partner closely with ML engineers, product engineers, data scientists, and platform teams across DoorDash, Wolt, and Deliveroo to translate emerging GenAI capabilities into durable, reusable platform primitives. Mentor senior engineers on infrastructure design patterns, distributed systems best practices, and production reliability. Raise technical bar for infrastructure quality and operational standards across the organization.
  • Define Future Technical Direction for AI Infrastructure: Set strategic technical direction for next-generation GenAI platform capabilities including reinforcement learning infrastructure (RLHF/RLVR), agentic systems optimization, post-training techniques, and emerging model architectures. Evaluate and integrate new inference engines, model serving frameworks, and GPU optimization techniques as the ecosystem evolves rapidly.
  • Production Operations and Incident Response: Lead incident response and post-incident review processes for critical infrastructure. Establish and maintain operational runbooks, monitoring dashboards, alerting strategies, and performance debugging practices. Drive culture of reliability engineering, cost optimization, and continuous improvement in production systems serving billions of inference requests.
  • Develop GPU Infrastructure and Resource Management: Design and implement GPU autoscaling strategies, resource allocation algorithms, and cluster management systems that maximize utilization while meeting latency SLOs. Own capacity planning, cost tracking, and vendor negotiation for GPU resources. Implement advanced scheduling and batching strategies for both real-time and batch workloads.
  • Champion Platform Adoption and Developer Experience: Build self-serve infrastructure tooling and developer-friendly abstractions that enable 200+ internal teams to safely deploy GenAI products into production. Create clear abstractions across open-weight and closed-source models, implement vendor abstraction layers, and provide comprehensive documentation, examples, and support to accelerate platform adoption company-wide.
  • Evaluate and Integrate Emerging Technologies: Research and evaluate emerging LLM inference engines (vLLM, SGLang, TensorRT-LLM), distributed training frameworks, GPU optimization techniques, and fine-tuning methodologies. Conduct technical POCs, benchmark studies, and cost-benefit analyses. Guide strategic decisions on infrastructure investments and technology partnerships to maintain competitive cost and latency advantages.
  • Implement Observability and Cost Attribution: Design comprehensive monitoring, tracing, and observability systems for end-to-end visibility into model inference, resource utilization, and cost drivers. Build accurate cost attribution models enabling internal teams to understand and optimize their AI infrastructure spending. Implement SLO-based monitoring, alert fatigue reduction, and incident correlation systems.
  • Architect AI Agent and MCP Server Infrastructure: Build infrastructure supporting production AI agents, multi-turn conversations, tool integrations, and Model Context Protocol (MCP) server deployments. Design agent orchestration, state management, and streaming response systems. Implement reliability patterns for long-running agent interactions and complex multi-step workflows at scale.
  • Lead Vector Database and RAG Infrastructure: Own evaluation systems, RAG infrastructure, vector database integration, and semantic search capabilities supporting GenAI products. Design efficient retrieval pipelines, embedding generation systems, and knowledge base management. Implement evaluation frameworks assessing model quality, retrieval relevance, and end-to-end system performance.
  • Mentor Senior Engineers on AI Infrastructure: Conduct technical interviews assessing infrastructure expertise, system design capability, and operational maturity. Mentor mid-career engineers advancing toward staff-level infrastructure roles. Provide technical guidance on architectural decisions, performance optimization techniques, and production reliability patterns in distributed systems and ML infrastructure contexts.
  • Contribute to Platform Engineering Standards: Establish standards for infrastructure documentation, architectural decision records (ADRs), design reviews, and technical RFCs. Drive adoption of infrastructure-as-code practices, configuration management, and GitOps workflows. Build reusable infrastructure libraries, SDKs, and internal platforms accelerating velocity across the GenAI ecosystem.
  • Optimize for Business Impact and Cost Efficiency: Balance infrastructure investments across cost, performance, latency, and developer experience. Work with finance and product teams to understand unit economics of AI infrastructure. Drive initiatives that deliver 10-100x efficiency improvements through architectural innovation, vendor selection, and resource optimization.
  • Build Guardrails and Safety Infrastructure: Design and implement safety guardrails, content moderation pipelines, rate limiting, and anomaly detection systems protecting against abuse and ensuring responsible AI deployment. Build infrastructure supporting compliance requirements, audit logging, and governance controls for production AI systems. Implement cost controls preventing runaway spending on model inference and fine-tuning.
  • Establish Kubernetes and Cloud Operations Excellence: Lead Kubernetes cluster management, AWS/GCP cloud infrastructure optimization, serverless GPU platforms integration (Modal, Lambda Labs), and cost optimization strategies. Implement advanced networking, storage, and compute configurations supporting high-performance GPU workloads. Drive cloud cost allocation, FinOps practices, and vendor negotiation for optimal pricing.
  • Drive Continuous Performance Improvement: Establish performance testing frameworks, benchmarking pipelines, and profiling practices for GPU inference and fine-tuning workloads. Drive quarterly performance optimization initiatives targeting 10-20% latency reductions and cost improvements. Implement A/B testing infrastructure for evaluating serving stack improvements, model routing strategies, and inference optimization techniques.
  • Support Developer AI Tools Integration: Integrate AI coding assistants (Claude Code, GitHub Copilot, Cursor) into development workflows and infrastructure tooling. Support team adoption of AI-powered code generation, testing, and debugging tools. Balance productivity gains with security, reliability, and quality requirements for mission-critical infrastructure.
  • Build Data Pipelines for Model Fine-Tuning: Design robust data pipelines supporting model fine-tuning and training infrastructure. Implement data preparation, validation, augmentation, and quality assurance systems. Build end-to-end pipelines supporting SFT, DPO, LoRA, and emerging fine-tuning methodologies. Implement reproducibility, versioning, and audit trail systems for model training.
  • Collaborate on Advanced Fine-Tuning Techniques: Lead architecture and implementation of advanced model customization including distributed multi-node fine-tuning, preference learning (DPO/IPO), reinforcement learning optimization (RLHF/RLVR), and LoRA-based efficient fine-tuning. Optimize for throughput, cost, and model quality. Implement experimentation frameworks enabling rapid iteration on fine-tuning strategies.
  • Drive Technical Excellence and Code Quality: Establish code review standards, testing practices, and deployment safety mechanisms for infrastructure code. Promote infrastructure-as-code practices, automated testing, and continuous deployment pipelines. Drive refactoring initiatives improving codebase maintainability, reducing technical debt, and enhancing developer velocity across the platform team.
  • Plan Capacity and Resource Allocation: Conduct capacity planning for GPU resources, compute clusters, and network infrastructure. Project demand across model types, inference patterns, and seasonal peaks. Optimize resource allocation balancing cost, latency, and reliability. Lead negotiations with GPU vendors and cloud providers for optimal pricing and availability guarantees.
  • Implement Vendor Abstraction and Model Routing: Design LLM gateway architecture abstracting across open-weight and closed-source model providers. Implement intelligent routing strategies optimizing for latency, cost, and availability. Build fallback mechanisms ensuring reliability when primary models unavailable. Implement cost-based routing optimizing for unit economics while maintaining performance SLOs.

Compensation

Pay and benefits.

Base·GBP 180,000 – 280,000

Full posting

Original listing.

Senior Software Engineer, Machine Learning Infrastructure - Generative AI

About the Team

Deliveroo's GenAI Platform team sits within Machine Learning Platform and builds the shared infrastructure that helps DoorDash, Wolt, and Deliveroo teams safely bring GenAI-powered products, agents, automation, and personalisation to production. Our mission is to increase the velocity of business impact from GenAI. A central pillar of that work is running frontier open-weight LLMs and VLMs (such as GLM, Qwen, Kimi, and DeepSeek) ourselves — real-time GPU serving, high-throughput batch inference, and fine-tuning on autoscaling GPUs — delivering large cost and latency wins (for example, a billion embeddings produced roughly 20× cheaper and visual models served roughly 72% cheaper). We also own core platform surfaces including the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution.

About the Role

You will join a small, high-leverage team building production infrastructure for Generative AI at Deliveroo and DoorDash, leading the design and architecture of our open-weights model platform spanning inference and fine-tuning: real-time GPU serving, high-throughput batch inference, and model fine-tuning. You’ll set technical direction across model serving and inference engines, fine-tuning and training pipelines, GPU autoscaling and utilization, batch pipelines, backend services, and observability, and mentor engineers as you go. This role is ideal for a senior engineer who enjoys owning ambiguous, high-impact systems and pushing the cost/performance frontier of GPU inference and fine-tuning in a fast-moving technical area where product needs, model capabilities, vendor ecosystems, and cost/performance tradeoffs are evolving quickly.

You’re excited about this opportunity because you will…

  • Lead the design of infrastructure that helps teams move GenAI ideas from prototype to production, increasing the velocity of business impact from AI across the company.

  • Own and evolve our open-weights serving stack — real-time GPU endpoints, high-throughput batch inference, and fine-tuning (SFT/DPO/LoRA) — alongside the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution.

  • Architect scalable, high-performance systems for model serving, batch inference, GPU autoscaling, and fine-tuning that power real customer and internal automation use cases

  • Push the cost and latency frontier of GPU inference — turning batch jobs that took days into hours and cutting inference cost by multiples — while giving product teams a clean choice across open-weight and closed-source models with reliability, fallback, observability, and cost controls built in.

  • Build platforms that support rapid experimentation while meeting production standards for latency, scale, monitoring, SLOs, playbooks, and operational excellence.

  • Partner closely with — and raise the technical bar for — ML engineers, product engineers, data scientists, and platform teams across DoorDash, Wolt, and Deliveroo to turn emerging GenAI capabilities into durable platform primitives.

  • Set technical direction for the future of the company's centralised GenAI platform — including emerging directions such as reinforcement learning (RLHF/RLVR), agent optimisation, and other post-training and agentic techniques — enabling the next generation of AI-powered products, agents, automation, and personalisation.

We’re excited about you because…

  • BSc, MSc or PhD in Computer Science or equivalent

  • 5+ years of industry experience in software engineering

  • Deep backend engineering fundamentals, especially in Python and distributed systems.

  • Track record of designing and owning production services, APIs, data pipelines, or ML infrastructure at scale.

  • Experience operating systems in production, including observability, debugging, reliability, incident response, and performance/cost optimization.

  • Deep hands-on experience with LLM inference and/or fine-tuning of open-weight models in production — serving (latency, throughput, batching, autoscaling, GPU utilization) and/or fine-tuning (SFT/DPO/LoRA).

  • Demonstrated technical leadership: leading design across ambiguous, fast-moving technical areas, mentoring engineers, and turning customer use cases into reusable platform capabilities

  • Proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) in the full software development lifecycle, including designing, generating code, testing, monitoring and releasing software

Nice To Haves

  • Experience with LLM inference engines and serving frameworks (e.g., vLLM, SGLang, TensorRT-LLM) in production

  • Experience with distributed/multi-node fine-tuning and training pipelines (SFT, DPO/RLHF, LoRA), including data preparation and evaluation

  • GPU performance work — multi-node/distributed inference, KV-cache/memory optimisation, quantisation (FP8/INT8/AWQ/GPTQ), or cold-start/throughput tuning

  • Experience with Kubernetes, cloud infrastructure (AWS/GCP), GPUs, serverless/elastic GPU platforms (e.g., Modal), or high-throughput batch systems

  • Experience with LLM gateways, model routing, vendor abstraction, or cost attribution

  • Experience building developer platforms, internal platforms, or self-serve infrastructure

  • Experience building and deploying AI agents or MCP servers in production

  • Experience with eval systems, LLM observability, tracing, RAG, search, or vector databases

Diversity, Equity and Inclusion

At Deliveroo, we know that a great workplace reflects the world around us and that true diversity and inclusion make us stronger, more creative, and better at what we do. We’re committed to fostering an environment where everyone can do their best work and feel they belong.

We believe in equality of opportunity and welcome candidates from all backgrounds regardless of age, gender, ethnicity, disability, sexual orientation, gender identity, socio-economic background, religion, or belief.

If you have a disability or long-term health condition and need support to apply for one of our roles, or require any reasonable adjustments during the recruitment process, you’ll have the opportunity to let us know once you’ve submitted your application. We’ll share details on how to request support so we can ensure you have a fair and equitable experience.

If you’re excited about making a real impact in a fast-moving marketplace and growing your career alongside ambitious, supportive teams, we’d love to hear from you!

Redirects to Deliveroo's application page.

Other roles

More at Deliveroo.

View all 20 roles