Software Engineer, Machine Learning Infrastructure
Machine Learning Infrastructure Engineer · Mid · Full Time
Opens Deliveroo's application page
Role
What you'll do.
Join Deliveroo's GenAI Platform team to build production-grade infrastructure for generative AI across DoorDash, Wolt, and Deliveroo. This role focuses on scaling open-weight LLM and VLM infrastructure, including real-time GPU serving, high-throughput batch inference, and distributed fine-tuning pipelines while optimizing cost and latency. You'll design systems that power AI agents, automation, and personalization while working across model serving frameworks, GPU autoscaling, backend services, and observability in a fast-moving technical environment.
Responsibilities
- Design and build production GPU serving infrastructure: Architect and implement real-time GPU serving endpoints, high-throughput batch inference pipelines, and autoscaling systems for open-weight LLMs and VLMs. Optimize for cost, latency, and throughput while handling multi-model deployments across inference engines like vLLM and SGLang.
- Develop and optimize model fine-tuning infrastructure: Build distributed fine-tuning and training pipelines supporting SFT, DPO, RLHF, and LoRA techniques on autoscaling GPUs. Handle data preparation, evaluation frameworks, and production-ready training orchestration that supports rapid experimentation.
- Implement GPU autoscaling and resource optimization: Engineer intelligent GPU autoscaling systems, optimize GPU utilization rates, and implement cost attribution mechanisms. Focus on KV-cache optimization, quantization strategies (FP8/INT8/AWQ/GPTQ), and multi-node distributed inference patterns.
- Own platform gateway and observability systems: Develop and maintain LLM Gateway and Agent Gateway components, implement comprehensive observability including tracing and monitoring, build evals infrastructure and guardrails, and establish production reliability standards including SLOs and incident playbooks.
- Partner across engineering organizations: Collaborate closely with ML engineers, product engineers, data scientists, and platform teams across DoorDash, Wolt, and Deliveroo to translate business requirements into scalable platform primitives and infrastructure abstractions.
- Push cost and performance frontiers: Continuously optimize inference and fine-tuning systems to deliver cost and latency wins, turning days-long batch jobs into hours while reducing inference costs by multiples. Evaluate emerging model architectures, vendor offerings, and optimization techniques.
- Establish production excellence standards: Build platforms that support rapid experimentation while maintaining production standards for reliability, performance, monitoring, and operational excellence. Develop playbooks, runbooks, and automated remediation capabilities.
Qualifications
What we look for.
Technical
Backend engineering in Python
Strong proficiency in Python for building scalable backend services, distributed systems, and data pipelines. Experience designing APIs, handling concurrency, and optimizing performance for production workloads.
Distributed systems architecture
Deep understanding of distributed systems concepts including consensus, partitioning, replication, and orchestration. Experience designing systems that scale horizontally across multiple nodes and handle network failures gracefully.
Production LLM serving and inference
Hands-on production experience with LLM inference, model serving frameworks (vLLM, SGLang, TensorRT-LLM), and optimization techniques including batching, autoscaling, KV-cache management, and quantization methods.
GPU computing and CUDA fundamentals
Understanding of GPU architecture, memory hierarchies, CUDA programming basics, and GPU-specific optimization strategies for machine learning workloads including distributed multi-GPU training and inference.
Production observability and debugging
Proficiency in production monitoring, logging, tracing, and debugging complex distributed systems. Experience with metrics collection, anomaly detection, and performance profiling at scale.
Open-weight model fine-tuning
Production experience with fine-tuning open-weight models including supervised fine-tuning (SFT), DPO, RLHF, and LoRA techniques. Understanding of training dynamics, convergence, and evaluation methodologies.
Cloud infrastructure and containerization
Experience with Kubernetes orchestration, cloud platforms (AWS/GCP), containerization strategies, and serverless/elastic GPU platforms. Knowledge of infrastructure-as-code and deployment automation.
Education
Bachelor's degree in Computer Science
BSc in Computer Science, Computer Engineering, or equivalent discipline providing foundational knowledge in algorithms, data structures, and systems design.
Master's degree or higher (preferred)
MSc or PhD in Computer Science, Machine Learning, or related field demonstrating advanced research capabilities and deep theoretical understanding of distributed systems or machine learning infrastructure.
Experience
3+ years software engineering industry experience
Minimum three years of professional software engineering experience building production systems, preferably in backend infrastructure, platform engineering, or machine learning systems.
Production-scale data infrastructure
Track record of building and operating production services, APIs, data pipelines, or ML infrastructure serving significant scale. Demonstrated ability to handle reliability, performance optimization, and operational excellence.
Production system operations experience
Real-world experience operating systems in production including incident response, performance optimization, cost optimization, debugging complex failures, and establishing reliability standards.
AI coding tools proficiency
Proficiency using modern AI coding assistants (Claude, Codex, Cursor) throughout the full software development lifecycle including design, code generation, testing, monitoring, and deployment.
Skills
Required
Python backend development
Expert-level Python programming for building scalable backend services, production APIs, and data processing pipelines with focus on performance and reliability.
Distributed systems design
Ability to architect and implement distributed systems with consideration for scalability, fault tolerance, consistency models, and operational complexity.
Production LLM infrastructure
Practical experience building or operating LLM serving infrastructure in production environments, including inference optimization, model routing, and resource management.
GPU and CUDA fundamentals
Working knowledge of GPU architecture, CUDA programming basics, and how to optimize machine learning computations for GPU execution and memory efficiency.
Cloud infrastructure (AWS/GCP)
Hands-on experience deploying and managing infrastructure on AWS or GCP including compute services, networking, storage, and cost optimization.
Kubernetes and container orchestration
Proficiency with Kubernetes for container orchestration, service deployment, resource management, and production operations.
Systems observability
Experience implementing monitoring, logging, tracing, and alerting for production systems. Ability to debug performance issues and establish SLOs.
Fine-tuning techniques for LLMs
Practical knowledge of supervised fine-tuning (SFT), direct preference optimization (DPO), reinforcement learning from human feedback (RLHF), and LoRA-based adaptation methods.
Preferred
vLLM and SGLang framework experience
Nice to haveProduction experience with vLLM, SGLang, or TensorRT-LLM inference engines for optimizing LLM serving performance, batching, and multi-GPU inference.
LLM quantization techniques
Nice to haveExperience with quantization methods including FP8, INT8, AWQ, GPTQ, and other techniques for reducing model size and inference latency while maintaining accuracy.
Multi-node distributed training
Nice to haveExperience implementing or operating distributed training systems using frameworks like PyTorch Distributed, DeepSpeed, or Ray for large-scale fine-tuning and training.
Vector databases and semantic search
Nice to haveFamiliarity with vector databases (Pinecone, Weaviate, Milvus), embedding systems, and RAG architectures for retrieval-augmented generation applications.
LLM evaluation and evals infrastructure
Nice to haveExperience building evaluation frameworks, running LLM benchmarks, and establishing metrics for model quality assessment.
AI agents and MCP servers
Nice to haveProduction experience building AI agents, orchestration frameworks, or Model Context Protocol (MCP) servers that coordinate multiple tools and models.
API gateway and routing systems
Nice to haveExperience building or operating API gateways, load balancers, or intelligent routing systems that handle vendor abstraction and dynamic routing decisions.
Cost attribution and metering
Nice to haveExperience implementing cost tracking, usage metering, and attribution systems for multi-tenant infrastructure serving business units.
Serverless and elastic GPU platforms
Nice to haveExperience with serverless computing platforms or elastic GPU services (Modal, Anyscale, Lambda Labs) for dynamic workload scaling.
Performance profiling and optimization
Nice to haveAdvanced skills in profiling tools, identifying bottlenecks in GPU code, and implementing targeted optimizations for latency and throughput.
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·GBP 140,000 – 200,000
Equity·Stock options
Benefits
Equity and stock options
Participate in Deliveroo's growth through competitive equity grants allowing you to build long-term wealth alongside the company.
Comprehensive health and wellness benefits
Health insurance coverage including medical, dental, and vision plans with options for you and your family members.
Flexible working arrangements
Remote work flexibility and flexible scheduling to support work-life balance while collaborating across distributed teams.
Professional development budget
Annual learning and development budget for conferences, courses, certifications, and training programs to grow your technical expertise.
Generous time off policy
Competitive paid time off including vacation days, sick leave, and parental leave to ensure adequate rest and personal time.
Pension and retirement planning
Employer-contributed pension scheme supporting long-term financial security and retirement planning.
Mental health and wellbeing support
Access to mental health resources, counseling services, and wellness programs supporting holistic employee wellbeing.
Parental leave and family support
Comprehensive parental leave policies for various family situations and childcare support options.
Commuter benefits and transportation
Support for transportation costs including commuter passes and transit benefits for employees commuting to offices.
Diversity and inclusion initiatives
Active commitment to fostering an inclusive workplace with employee resource groups, diversity programs, and inclusive hiring practices.
Process
Interview steps.
- 01
Initial screening call
30-minute phone screening with a recruiter covering your background, experience with ML infrastructure, and initial technical alignment with the role requirements.
- 02
Technical screening interview
60-90 minute technical conversation with an ML Infrastructure engineer focusing on your production experience with distributed systems, LLM serving, and GPU infrastructure. Expect discussion of architectural decisions and optimization techniques.
- 03
System design deep dive
90-minute interview centered on designing large-scale ML infrastructure systems. You'll discuss tradeoffs in model serving architectures, GPU autoscaling strategies, and how you'd approach cost/performance optimization challenges.
- 04
Production experience panel
Interview with multiple engineers on the GenAI Platform team exploring your hands-on experience operating production systems, incident response capabilities, and how you approach reliability and observability.
- 05
Cross-functional collaboration discussion
Conversation with stakeholders from product, data science, or adjacent platform teams to assess your ability to translate business requirements into platform abstractions and work effectively across organizations.
- 06
Final conversation with hiring manager
Discussion with the engineering manager covering team dynamics, long-term career goals, and fit for Deliveroo's culture of technical excellence and operational ownership.
Full posting
Original listing.
Software Engineer, Machine Learning Infrastructure - Generative AI
About the Team
Deliveroo's GenAI Platform team sits within Machine Learning Platform and builds the shared infrastructure that helps DoorDash, Wolt, and Deliveroo teams safely bring GenAI-powered products, agents, automation, and personalization to production. Our mission is to increase the velocity of business impact from GenAI. A central pillar of that work is running frontier open-weight LLMs and VLMs (such as GLM, Qwen, Kimi, and DeepSeek) ourselves — real-time GPU serving, high-throughput batch inference, and fine-tuning on autoscaling GPUs — delivering large cost and latency wins (for example, a billion embeddings produced roughly 20× cheaper and visual models served roughly 72% cheaper). We also own core platform surfaces including the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution.
About the Role
You will join a small, high-leverage team building production infrastructure for Generative AI at Deliveroo and DoorDash, with a primary focus on our open-weights model platform spanning inference and fine-tuning: real-time GPU serving, high-throughput batch inference, and model fine-tuning. You’ll work across model serving and inference engines, fine-tuning and training pipelines, GPU autoscaling and utilization, batch pipelines, backend services, and observability. This role is ideal for an engineer who enjoys pushing the cost/performance frontier of GPU inference and fine-tuning in a fast-moving technical area where product needs, model capabilities, vendor ecosystems, and cost/performance tradeoffs are evolving quickly.
You’re excited about this opportunity because you will…
Build the infrastructure that helps Deliveroo teams move GenAI ideas from prototype to production, increasing the velocity of business impact from AI across the company.
Work on our open-weights serving stack — real-time GPU endpoints, high-throughput batch inference, and fine-tuning (SFT/DPO/LoRA) — alongside the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution.
Design scalable, high-performance systems for model serving, batch inference, GPU autoscaling, and fine-tuning that power real customer and internal automation use cases
Push the cost and latency frontier of GPU inference — turning batch jobs that took days into hours and cutting inference cost by multiples — while giving product teams a clean choice across open-weight and closed-source models with reliability, fallback, observability, and cost controls built in.
Build platforms that support rapid experimentation while meeting production standards for latency, scale, monitoring, SLOs, playbooks, and operational excellence.
Partner closely with ML engineers, product engineers, data scientists, and platform teams across DoorDash, Wolt, and Deliveroo to turn emerging GenAI capabilities into durable platform primitives.
Shape the future of the centralized GenAI platform — including emerging directions such as reinforcement learning (RLHF/RLVR), agent optimization, and other post-training and agentic techniques — enabling the next generation of AI-powered products, agents, automation, and personalization.
We’re excited about you because you have…
BSc, MSc, or PhD in Computer Science or equivalent
3+ years of industry experience in software engineering
Strong backend engineering fundamentals, especially in Python and distributed systems.
Experience building production services, APIs, data pipelines, or ML infrastructure at scale.
Experience operating systems in production, including observability, debugging, reliability, incident response, and performance/cost optimization.
Hands-on experience with LLM inference and/or fine-tuning of open-weight models in production — serving (latency, throughput, batching, autoscaling, GPU utilization) and/or fine-tuning (SFT/DPO/LoRA).
Ability to work across ambiguous, fast-moving technical areas and turn customer use cases into reusable platform capabilities
Proficiency in using AI coding tools (e.g., Claude Code, Codex, Cursor) in the full software development lifecycle, including designing, generating code, testing, monitoring and releasing software
Nice To Haves
Experience with LLM inference engines and serving frameworks (e.g., vLLM, SGLang, TensorRT-LLM) in production
Experience with distributed/multi-node fine-tuning and training pipelines (SFT, DPO/RLHF, LoRA), including data preparation and evaluation
GPU performance work — multi-node/distributed inference, KV-cache/memory optimization, quantization (FP8/INT8/AWQ/GPTQ), or cold-start/throughput tuning
Experience with Kubernetes, cloud infrastructure (AWS/GCP), GPUs, serverless/elastic GPU platforms (e.g., Modal), or high-throughput batch systems
Experience with LLM gateways, model routing, vendor abstraction, or cost attribution
Experience building developer platforms, internal platforms, or self-serve infrastructure
Experience building and deploying AI agents or MCP servers in production
Experience with eval systems, LLM observability, tracing, RAG, search, or vector databases
Diversity, Equity and Inclusion
At Deliveroo, we know that a great workplace reflects the world around us and that true diversity and inclusion make us stronger, more creative, and better at what we do. We’re committed to fostering an environment where everyone can do their best work and feel they belong.
We believe in equality of opportunity and welcome candidates from all backgrounds regardless of age, gender, ethnicity, disability, sexual orientation, gender identity, socio-economic background, religion, or belief.
If you have a disability or long-term health condition and need support to apply for one of our roles, or require any reasonable adjustments during the recruitment process, you’ll have the opportunity to let us know once you’ve submitted your application. We’ll share details on how to request support so we can ensure you have a fair and equitable experience.
If you’re excited about making a real impact in a fast-moving marketplace and growing your career alongside ambitious, supportive teams, we’d love to hear from you!
Redirects to Deliveroo's application page.
Other roles
More at Deliveroo.
Software Engineer, Backend (BI Platform)
Mid
BI Engineer
Mid
Senior Software Engineer
Senior
Senior Software Engineer - Full Stack (Merchants)
Senior
Staff Software Engineer
Staff