Software Engineer, Machine Learning Infrastructure

Machine Learning Infrastructure Engineer · Mid · Full Time

London - The River Building HQGBP 140k – 200k1mo ago
Apply for this role

Opens Deliveroo's application page

Role

What you'll do.

Join Deliveroo's GenAI Platform team to build production-grade infrastructure for generative AI across DoorDash, Wolt, and Deliveroo. This role focuses on scaling open-weight LLM and VLM infrastructure, including real-time GPU serving, high-throughput batch inference, and distributed fine-tuning pipelines while optimizing cost and latency. You'll design systems that power AI agents, automation, and personalization while working across model serving frameworks, GPU autoscaling, backend services, and observability in a fast-moving technical environment.

Responsibilities

  • Design and build production GPU serving infrastructure: Architect and implement real-time GPU serving endpoints, high-throughput batch inference pipelines, and autoscaling systems for open-weight LLMs and VLMs. Optimize for cost, latency, and throughput while handling multi-model deployments across inference engines like vLLM and SGLang.
  • Develop and optimize model fine-tuning infrastructure: Build distributed fine-tuning and training pipelines supporting SFT, DPO, RLHF, and LoRA techniques on autoscaling GPUs. Handle data preparation, evaluation frameworks, and production-ready training orchestration that supports rapid experimentation.
  • Implement GPU autoscaling and resource optimization: Engineer intelligent GPU autoscaling systems, optimize GPU utilization rates, and implement cost attribution mechanisms. Focus on KV-cache optimization, quantization strategies (FP8/INT8/AWQ/GPTQ), and multi-node distributed inference patterns.
  • Own platform gateway and observability systems: Develop and maintain LLM Gateway and Agent Gateway components, implement comprehensive observability including tracing and monitoring, build evals infrastructure and guardrails, and establish production reliability standards including SLOs and incident playbooks.
  • Partner across engineering organizations: Collaborate closely with ML engineers, product engineers, data scientists, and platform teams across DoorDash, Wolt, and Deliveroo to translate business requirements into scalable platform primitives and infrastructure abstractions.
  • Push cost and performance frontiers: Continuously optimize inference and fine-tuning systems to deliver cost and latency wins, turning days-long batch jobs into hours while reducing inference costs by multiples. Evaluate emerging model architectures, vendor offerings, and optimization techniques.
  • Establish production excellence standards: Build platforms that support rapid experimentation while maintaining production standards for reliability, performance, monitoring, and operational excellence. Develop playbooks, runbooks, and automated remediation capabilities.

Qualifications

What we look for.

Technical

  • Backend engineering in Python

    Strong proficiency in Python for building scalable backend services, distributed systems, and data pipelines. Experience designing APIs, handling concurrency, and optimizing performance for production workloads.

  • Distributed systems architecture

    Deep understanding of distributed systems concepts including consensus, partitioning, replication, and orchestration. Experience designing systems that scale horizontally across multiple nodes and handle network failures gracefully.

  • Production LLM serving and inference

    Hands-on production experience with LLM inference, model serving frameworks (vLLM, SGLang, TensorRT-LLM), and optimization techniques including batching, autoscaling, KV-cache management, and quantization methods.

  • GPU computing and CUDA fundamentals

    Understanding of GPU architecture, memory hierarchies, CUDA programming basics, and GPU-specific optimization strategies for machine learning workloads including distributed multi-GPU training and inference.

  • Production observability and debugging

    Proficiency in production monitoring, logging, tracing, and debugging complex distributed systems. Experience with metrics collection, anomaly detection, and performance profiling at scale.

  • Open-weight model fine-tuning

    Production experience with fine-tuning open-weight models including supervised fine-tuning (SFT), DPO, RLHF, and LoRA techniques. Understanding of training dynamics, convergence, and evaluation methodologies.

  • Cloud infrastructure and containerization

    Experience with Kubernetes orchestration, cloud platforms (AWS/GCP), containerization strategies, and serverless/elastic GPU platforms. Knowledge of infrastructure-as-code and deployment automation.

Education

  • Bachelor's degree in Computer Science

    BSc in Computer Science, Computer Engineering, or equivalent discipline providing foundational knowledge in algorithms, data structures, and systems design.

  • Master's degree or higher (preferred)

    MSc or PhD in Computer Science, Machine Learning, or related field demonstrating advanced research capabilities and deep theoretical understanding of distributed systems or machine learning infrastructure.

Experience

  • 3+ years software engineering industry experience

    Minimum three years of professional software engineering experience building production systems, preferably in backend infrastructure, platform engineering, or machine learning systems.

  • Production-scale data infrastructure

    Track record of building and operating production services, APIs, data pipelines, or ML infrastructure serving significant scale. Demonstrated ability to handle reliability, performance optimization, and operational excellence.

  • Production system operations experience

    Real-world experience operating systems in production including incident response, performance optimization, cost optimization, debugging complex failures, and establishing reliability standards.

  • AI coding tools proficiency

    Proficiency using modern AI coding assistants (Claude, Codex, Cursor) throughout the full software development lifecycle including design, code generation, testing, monitoring, and deployment.

Skills

Required

  • Python backend development

    Expert-level Python programming for building scalable backend services, production APIs, and data processing pipelines with focus on performance and reliability.

  • Distributed systems design

    Ability to architect and implement distributed systems with consideration for scalability, fault tolerance, consistency models, and operational complexity.

  • Production LLM infrastructure

    Practical experience building or operating LLM serving infrastructure in production environments, including inference optimization, model routing, and resource management.

  • GPU and CUDA fundamentals

    Working knowledge of GPU architecture, CUDA programming basics, and how to optimize machine learning computations for GPU execution and memory efficiency.

  • Cloud infrastructure (AWS/GCP)

    Hands-on experience deploying and managing infrastructure on AWS or GCP including compute services, networking, storage, and cost optimization.

  • Kubernetes and container orchestration

    Proficiency with Kubernetes for container orchestration, service deployment, resource management, and production operations.

  • Systems observability

    Experience implementing monitoring, logging, tracing, and alerting for production systems. Ability to debug performance issues and establish SLOs.

  • Fine-tuning techniques for LLMs

    Practical knowledge of supervised fine-tuning (SFT), direct preference optimization (DPO), reinforcement learning from human feedback (RLHF), and LoRA-based adaptation methods.

Preferred

  • vLLM and SGLang framework experience

    Nice to have

    Production experience with vLLM, SGLang, or TensorRT-LLM inference engines for optimizing LLM serving performance, batching, and multi-GPU inference.

  • LLM quantization techniques

    Nice to have

    Experience with quantization methods including FP8, INT8, AWQ, GPTQ, and other techniques for reducing model size and inference latency while maintaining accuracy.

  • Multi-node distributed training

    Nice to have

    Experience implementing or operating distributed training systems using frameworks like PyTorch Distributed, DeepSpeed, or Ray for large-scale fine-tuning and training.

  • Vector databases and semantic search

    Nice to have

    Familiarity with vector databases (Pinecone, Weaviate, Milvus), embedding systems, and RAG architectures for retrieval-augmented generation applications.

  • LLM evaluation and evals infrastructure

    Nice to have

    Experience building evaluation frameworks, running LLM benchmarks, and establishing metrics for model quality assessment.

  • AI agents and MCP servers

    Nice to have

    Production experience building AI agents, orchestration frameworks, or Model Context Protocol (MCP) servers that coordinate multiple tools and models.

  • API gateway and routing systems

    Nice to have

    Experience building or operating API gateways, load balancers, or intelligent routing systems that handle vendor abstraction and dynamic routing decisions.

  • Cost attribution and metering

    Nice to have

    Experience implementing cost tracking, usage metering, and attribution systems for multi-tenant infrastructure serving business units.

  • Serverless and elastic GPU platforms

    Nice to have

    Experience with serverless computing platforms or elastic GPU services (Modal, Anyscale, Lambda Labs) for dynamic workload scaling.

  • Performance profiling and optimization

    Nice to have

    Advanced skills in profiling tools, identifying bottlenecks in GPU code, and implementing targeted optimizations for latency and throughput.

Tech stack

Languages

PythonSQLCUDA/C++

Frameworks

vLLMSGLangTensorRT-LLMPyTorch DistributedDeepSpeedFastAPIRay

Databases

PostgreSQLRedisVector databases (Pinecone/Weaviate/Milvus)TimescaleDB

Tools

KubernetesAWS (EC2, S3, SageMaker)GCP (Compute Engine, Cloud Storage, TPUs)Prometheus & GrafanaJaeger/DatadogDockerModalGit/GitHub

Other

GPU profiling and optimization toolsAI coding assistantsBash and shell scriptingTerraform/Infrastructure-as-CodeModel evaluation frameworks

Compensation

Pay and benefits.

Base·GBP 140,000 – 200,000

Equity·Stock options

Benefits

  • Equity and stock options

    Participate in Deliveroo's growth through competitive equity grants allowing you to build long-term wealth alongside the company.

  • Comprehensive health and wellness benefits

    Health insurance coverage including medical, dental, and vision plans with options for you and your family members.

  • Flexible working arrangements

    Remote work flexibility and flexible scheduling to support work-life balance while collaborating across distributed teams.

  • Professional development budget

    Annual learning and development budget for conferences, courses, certifications, and training programs to grow your technical expertise.

  • Generous time off policy

    Competitive paid time off including vacation days, sick leave, and parental leave to ensure adequate rest and personal time.

  • Pension and retirement planning

    Employer-contributed pension scheme supporting long-term financial security and retirement planning.

  • Mental health and wellbeing support

    Access to mental health resources, counseling services, and wellness programs supporting holistic employee wellbeing.

  • Parental leave and family support

    Comprehensive parental leave policies for various family situations and childcare support options.

  • Commuter benefits and transportation

    Support for transportation costs including commuter passes and transit benefits for employees commuting to offices.

  • Diversity and inclusion initiatives

    Active commitment to fostering an inclusive workplace with employee resource groups, diversity programs, and inclusive hiring practices.

Process

Interview steps.

  1. 01

    Initial screening call

    30-minute phone screening with a recruiter covering your background, experience with ML infrastructure, and initial technical alignment with the role requirements.

  2. 02

    Technical screening interview

    60-90 minute technical conversation with an ML Infrastructure engineer focusing on your production experience with distributed systems, LLM serving, and GPU infrastructure. Expect discussion of architectural decisions and optimization techniques.

  3. 03

    System design deep dive

    90-minute interview centered on designing large-scale ML infrastructure systems. You'll discuss tradeoffs in model serving architectures, GPU autoscaling strategies, and how you'd approach cost/performance optimization challenges.

  4. 04

    Production experience panel

    Interview with multiple engineers on the GenAI Platform team exploring your hands-on experience operating production systems, incident response capabilities, and how you approach reliability and observability.

  5. 05

    Cross-functional collaboration discussion

    Conversation with stakeholders from product, data science, or adjacent platform teams to assess your ability to translate business requirements into platform abstractions and work effectively across organizations.

  6. 06

    Final conversation with hiring manager

    Discussion with the engineering manager covering team dynamics, long-term career goals, and fit for Deliveroo's culture of technical excellence and operational ownership.

Full posting

Original listing.

Software Engineer, Machine Learning Infrastructure - Generative AI

About the Team

Deliveroo's GenAI Platform team sits within Machine Learning Platform and builds the shared infrastructure that helps DoorDash, Wolt, and Deliveroo teams safely bring GenAI-powered products, agents, automation, and personalization to production. Our mission is to increase the velocity of business impact from GenAI. A central pillar of that work is running frontier open-weight LLMs and VLMs (such as GLM, Qwen, Kimi, and DeepSeek) ourselves — real-time GPU serving, high-throughput batch inference, and fine-tuning on autoscaling GPUs — delivering large cost and latency wins (for example, a billion embeddings produced roughly 20× cheaper and visual models served roughly 72% cheaper). We also own core platform surfaces including the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution.

About the Role

You will join a small, high-leverage team building production infrastructure for Generative AI at Deliveroo and DoorDash, with a primary focus on our open-weights model platform spanning inference and fine-tuning: real-time GPU serving, high-throughput batch inference, and model fine-tuning. You’ll work across model serving and inference engines, fine-tuning and training pipelines, GPU autoscaling and utilization, batch pipelines, backend services, and observability. This role is ideal for an engineer who enjoys pushing the cost/performance frontier of GPU inference and fine-tuning in a fast-moving technical area where product needs, model capabilities, vendor ecosystems, and cost/performance tradeoffs are evolving quickly.

You’re excited about this opportunity because you will…

  • Build the infrastructure that helps Deliveroo teams move GenAI ideas from prototype to production, increasing the velocity of business impact from AI across the company.

  • Work on our open-weights serving stack — real-time GPU endpoints, high-throughput batch inference, and fine-tuning (SFT/DPO/LoRA) — alongside the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution.

  • Design scalable, high-performance systems for model serving, batch inference, GPU autoscaling, and fine-tuning that power real customer and internal automation use cases

  • Push the cost and latency frontier of GPU inference — turning batch jobs that took days into hours and cutting inference cost by multiples — while giving product teams a clean choice across open-weight and closed-source models with reliability, fallback, observability, and cost controls built in.

  • Build platforms that support rapid experimentation while meeting production standards for latency, scale, monitoring, SLOs, playbooks, and operational excellence.

  • Partner closely with ML engineers, product engineers, data scientists, and platform teams across DoorDash, Wolt, and Deliveroo to turn emerging GenAI capabilities into durable platform primitives.

  • Shape the future of the centralized GenAI platform — including emerging directions such as reinforcement learning (RLHF/RLVR), agent optimization, and other post-training and agentic techniques — enabling the next generation of AI-powered products, agents, automation, and personalization.

We’re excited about you because you have…

  • BSc, MSc, or PhD in Computer Science or equivalent

  • 3+ years of industry experience in software engineering

  • Strong backend engineering fundamentals, especially in Python and distributed systems.

  • Experience building production services, APIs, data pipelines, or ML infrastructure at scale.

  • Experience operating systems in production, including observability, debugging, reliability, incident response, and performance/cost optimization.

  • Hands-on experience with LLM inference and/or fine-tuning of open-weight models in production — serving (latency, throughput, batching, autoscaling, GPU utilization) and/or fine-tuning (SFT/DPO/LoRA).

  • Ability to work across ambiguous, fast-moving technical areas and turn customer use cases into reusable platform capabilities

  • Proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) in the full software development lifecycle, including designing, generating code, testing, monitoring and releasing software

Nice To Haves

  • Experience with LLM inference engines and serving frameworks (e.g., vLLM, SGLang, TensorRT-LLM) in production

  • Experience with distributed/multi-node fine-tuning and training pipelines (SFT, DPO/RLHF, LoRA), including data preparation and evaluation

  • GPU performance work — multi-node/distributed inference, KV-cache/memory optimization, quantization (FP8/INT8/AWQ/GPTQ), or cold-start/throughput tuning

  • Experience with Kubernetes, cloud infrastructure (AWS/GCP), GPUs, serverless/elastic GPU platforms (e.g., Modal), or high-throughput batch systems

  • Experience with LLM gateways, model routing, vendor abstraction, or cost attribution

  • Experience building developer platforms, internal platforms, or self-serve infrastructure

  • Experience building and deploying AI agents or MCP servers in production

  • Experience with eval systems, LLM observability, tracing, RAG, search, or vector databases

Diversity, Equity and Inclusion

At Deliveroo, we know that a great workplace reflects the world around us and that true diversity and inclusion make us stronger, more creative, and better at what we do. We’re committed to fostering an environment where everyone can do their best work and feel they belong.

We believe in equality of opportunity and welcome candidates from all backgrounds regardless of age, gender, ethnicity, disability, sexual orientation, gender identity, socio-economic background, religion, or belief.

If you have a disability or long-term health condition and need support to apply for one of our roles, or require any reasonable adjustments during the recruitment process, you’ll have the opportunity to let us know once you’ve submitted your application. We’ll share details on how to request support so we can ensure you have a fair and equitable experience.

If you’re excited about making a real impact in a fast-moving marketplace and growing your career alongside ambitious, supportive teams, we’d love to hear from you!

Redirects to Deliveroo's application page.

Other roles

More at Deliveroo.

View all 24 roles