Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Backend Engineer · Senior · Full Time
Opens Perplexity AI's application page
Role
What you'll do.
Perplexity AI seeks a Member of Technical Staff to design and own a self-serve GPU cluster infrastructure platform supporting hundreds of millions of monthly AI queries. This role requires deep expertise in Kubernetes orchestration, multi-cloud GPU fleet management, and distributed systems to build a unified platform that abstracts complexity for inference engineers and researchers while balancing the competing demands of long-running training jobs and production inference services across NVIDIA hardware deployed at scale.
Responsibilities
- Design and Build Self-Serve GPU Compute Platform: Architect and implement a unified, self-serve compute platform that enables inference engineers and machine learning researchers to launch distributed training jobs and operate production inference services without requiring deep knowledge of GPU provisioning, cluster configuration, or cloud provider-specific infrastructure. This platform must abstract multi-cloud complexity into a consistent user experience while maintaining fine-grained control for power users.
- Own GPU Fleet Operations and Lifecycle Management: Take full ownership of provisioning, lifecycle management, reliability, and capacity integration across multiple cloud providers including CoreWeave, AWS, and GCP. Ensure consistent compute resource availability and performance across all deployment locations while managing hardware inventory, driver updates, firmware management, and decommissioning workflows.
- Implement Intelligent GPU Scheduling and Resource Optimization: Develop sophisticated scheduling and placement algorithms that discover available GPU capacity across providers, efficiently pack workloads onto hardware, and optimize resource utilization under real-world constraints including GPU scarcity, network bandwidth limitations, and multi-tenancy requirements. Build logic that maximizes cost efficiency while meeting service level objectives.
- Support Dual Workload Orchestration (Training and Inference): Manage fundamentally different infrastructure requirements for long-running distributed deep learning training jobs and low-latency production inference services running simultaneously on shared GPU clusters. Implement workload isolation, priority schemes, and resource guarantees that prevent inference latency degradation while maintaining efficient training job throughput.
- Architect Kubernetes GPU Orchestration Framework: Design and implement custom Kubernetes operators, Custom Resource Definitions (CRDs), and federation mechanisms to enable multi-cluster GPU orchestration across providers. Ensure consistent platform behavior, API interfaces, and operational patterns across all Kubernetes clusters while managing provider-specific nuances and networking configurations.
- Build Fault Tolerance and Resilience Systems: Engineer sophisticated fault tolerance mechanisms including automatic failover, node health detection, and graceful degradation. Implement cluster autoscaling policies, node eviction handling, and workload migration strategies that maintain service availability through provider outages, hardware failures, and capacity fluctuations without manual intervention.
- Establish Observability and Monitoring Infrastructure: Build comprehensive observability systems for GPU cluster health, utilization metrics, and workload performance. Implement monitoring dashboards, alerting strategies, and debugging tools that provide visibility into resource allocation decisions, network performance, GPU utilization patterns, and cost attribution across cloud providers and workload types.
- Define Technical Direction and Platform Architecture: Partner with inference engineering teams and cloud infrastructure engineers to translate operational constraints and business requirements into a coherent long-term platform architecture. Own technical roadmap decisions, evaluate emerging GPU orchestration technologies, and establish best practices for GPU cluster management across the organization.
Qualifications
What we look for.
Technical
Advanced Kubernetes Expertise
Demonstrated mastery of Kubernetes beyond basic kubectl operations, including design and implementation of custom operators, Custom Resource Definitions (CRDs), admission webhooks, multi-cluster federation, cluster networking, persistent volume management, and RBAC security models. Experience operating Kubernetes in production at significant scale with multiple concurrent workloads.
GPU Cluster Management and NVIDIA Stack
Hands-on experience provisioning, configuring, and operating large-scale GPU clusters using NVIDIA hardware. Deep knowledge of CUDA runtime, NVIDIA Container Runtime, GPU memory management, PCIe topology, NVIDIA Management Library (NVML), and NVIDIA Fabric Manager for peer-to-peer GPU communication. Understanding of GPU utilization patterns, thermal management, and hardware health monitoring.
High-Speed Interconnect Networking
Production experience with high-performance GPU networking technologies including InfiniBand, RoCE (RDMA over Converged Ethernet), or equivalent high-speed interconnects. Knowledge of network topology optimization, RDMA configuration, network performance tuning for collective communication operations (All-Reduce, All-Gather), and latency-sensitive workload networking.
Multi-Cloud Orchestration
Proven experience orchestrating compute workloads across multiple cloud providers such as CoreWeave, AWS (EC2, ECS, EKS), GCP (GKE, Compute Engine), or Azure. Understanding of provider-specific GPU offerings, network configurations, pricing models, quota management, service limits, and how to abstract provider differences into unified APIs.
Distributed Systems Design and Implementation
Strong foundation in distributed systems principles including resource scheduling algorithms, load balancing strategies, consensus mechanisms, fault tolerance patterns, and handling of distributed system failures. Ability to design systems that maintain consistency, availability, or partition tolerance trade-offs appropriate to workload requirements.
Systems Programming Languages
Production-level proficiency in Go, Rust, or C++ for writing infrastructure and systems-level code. Experience with low-level memory management, concurrent programming patterns, performance optimization, and debugging of systems-level software. Ability to write code suitable for long-running services, operators, and resource-constrained environments.
Training and Inference Workload Architecture
Comprehensive understanding of how distributed deep learning training and production ML inference services present different and often conflicting infrastructure demands. Knowledge of GPU memory requirements, communication patterns, fault recovery strategies, and scaling characteristics for both training frameworks (PyTorch, TensorFlow) and inference serving architectures.
Education
Bachelor's Degree in Computer Science or Related Field
Formal education in Computer Science, Computer Engineering, or equivalent discipline providing foundational knowledge in algorithms, data structures, operating systems, and computer architecture. Alternatively, demonstrable equivalent practical experience and self-directed learning in these areas.
Experience
GPU Infrastructure Operations at Scale
3+ years managing production GPU infrastructure supporting machine learning workloads, with hands-on responsibility for provisioning, capacity planning, incident response, and operational excellence. Experience handling real-world challenges including hardware failures, scaling to serve hundreds of concurrent workloads, and cost optimization.
Kubernetes Platform Engineering
3+ years of platform or infrastructure engineering focused on Kubernetes, with direct experience designing and operating Kubernetes clusters in production. Preferably including building internal platforms or developer-facing abstractions on top of Kubernetes for reducing operational complexity.
End-to-End System Ownership
Demonstrated ability to take ownership of complex technical problems without detailed specifications or clear path forward. Track record of independently identifying requirements, designing solutions, implementing systems, validating assumptions through testing, and iterating based on feedback.
Multi-Cloud or Hybrid Infrastructure
2+ years experience managing infrastructure spanning multiple cloud providers or hybrid environments, dealing with billing reconciliation, vendor-specific API differences, quota management, and architectural patterns for portability.
Skills
Required
Kubernetes
Advanced proficiency with Kubernetes including operators, CRDs, multi-cluster management, networking (CNI plugins), storage (CSI), and RBAC. Experience writing production Kubernetes manifests and controllers.
Go Programming
Production-level Go development for systems software, APIs, and long-running services. Experience with concurrency patterns, goroutines, channels, and writing maintainable systems code.
GPU Cluster Management
Deep operational knowledge of NVIDIA GPU clusters including CUDA, NVIDIA drivers, GPU orchestration, and cluster topology optimization for performance and reliability.
Distributed Systems Design
Strong understanding of scheduling algorithms, resource allocation, consensus mechanisms, fault tolerance patterns, and system resilience under load and failure scenarios.
Cloud Platforms
Hands-on experience with at least two major cloud providers (AWS, GCP, Azure, CoreWeave) including networking, instance management, quota handling, and cost optimization.
Infrastructure-as-Code and Configuration Management
Proficiency with infrastructure automation tools such as Terraform, Helm, Ansible, or similar for managing complex infrastructure deployments across multiple environments.
Linux System Administration
Deep Linux kernel knowledge, process management, networking stack, storage systems, and container runtime internals relevant to GPU workload optimization and debugging.
Preferred
Rust Systems Programming
Nice to haveExperience writing systems-level code in Rust, particularly for performance-critical components, memory safety requirements, or concurrent systems.
Inference Serving Frameworks
Nice to haveFamiliarity with production ML inference serving platforms including vLLM, SGLang, TensorRT-LLM, or Triton Inference Server and their infrastructure requirements.
HPC Schedulers
Nice to haveExperience with high-performance computing job schedulers such as Slurm, PBS, or similar, including job submission, queue management, and resource allocation policies.
CUDA Kernel Development
Nice to haveExperience writing or optimizing CUDA kernels or familiarity with GPU kernel optimization using Triton or similar languages for specialized computation patterns.
High-Speed Interconnect Configuration
Nice to haveProduction experience configuring and troubleshooting InfiniBand, RoCE, or RDMA networks for GPU-to-GPU communication, collective operations optimization, and network performance tuning.
ML Observability and Monitoring
Nice to haveExperience building observability systems for ML workloads using Prometheus, Grafana, Weights & Biases, or similar platforms for monitoring cluster health and workload performance.
PyTorch or TensorFlow Distributed Training
Nice to haveUnderstanding of distributed training frameworks, distributed data parallel and model parallel training patterns, gradient compression, and communication optimization techniques.
Network Performance Optimization
Nice to haveExperience diagnosing and optimizing network performance for communication-heavy workloads, including collective communication operations, topology awareness, and latency minimization.
Compensation
Pay and benefits.
Base·USD 250,000 – 485,000
Full posting
Original listing.
Perplexity serves hundreds of millions of queries a month, and every one of them fans out into multiple AI inference requests running in real time. Behind that sits a large GPU fleet spread across several cloud providers. Today, our inference engineers and researchers build models while also managing networking, securing capacity, and operating the underlying GPU clusters, responsibilities we want a dedicated platform team to own. Your job is to take ownership of that infrastructure and hide its complexity behind a unified, self-serve platform for running training and inference workloads.
Responsibilities
Build a self-serve compute platform. Design and own the systems that let inference engineers and researchers launch training jobs and operate inference services without managing GPU provisioning, cluster configuration, or provider-specific infrastructure.
Operate the GPU fleet. Own provisioning, lifecycle management, reliability, and capacity integration across providers, giving teams a consistent way to use compute regardless of where it runs.
Solve for GPU scarcity. Build the scheduling and placement logic that finds available capacity across providers, packs it efficiently, and gets the right workload onto the right hardware under real constraints.
Support two very different workloads. Keep long-running distributed training jobs healthy while simultaneously guaranteeing the availability and latency of production inference services on the same fleet.
Own the Kubernetes for GPU orchestration. Write the operators and CRDs, and manage many clusters across providers so the platform behaves the same everywhere we run.
Make failure boring. Build the fault tolerance, autoscaling, and observability that keep the fleet utilized and let workloads survive node loss, provider hiccups, and capacity shifts without human intervention.
Set technical direction across teams. Partner with inference and cloud infrastructure engineers to turn operational constraints into a coherent platform architecture and roadmap.
Qualifications
We expect you to have real depth in most of these:
Deep Kubernetes experience — custom operators, CRDs, and multi-cluster federation, not just running kubectl apply.
You've managed GPU clusters at scale: NVIDIA hardware, CUDA, and the networking that makes them fast (InfiniBand or RoCE).
You've orchestrated compute across multiple clouds (CoreWeave, AWS, GCP, or similar) and understand how different each one really is.
Strong distributed systems fundamentals: scheduling, resource allocation, and fault tolerance under load.
You write infrastructure and systems-level code in Go, Rust or C++.
You've supported both long-running training jobs and high-availability inference services, and you know why they pull infrastructure in opposite directions.
You own problems end-to-end and do well when the path forward isn't laid out for you.
Additional experience we value
Inference serving stacks: vLLM, SGLang, or TensorRT-LLM.
Slurm or other HPC schedulers.
GPU kernel work in CUDA or Triton — not required, but notable.
High-speed interconnects: InfiniBand, RoCE, or RDMA in production.
Observability for ML workloads: Prometheus, Grafana, or Weights & Biases.
Redirects to Perplexity AI's application page.
Other roles
More at Perplexity AI.
Member of Technical Staff (Android Engineer, Computer Growth)
Senior
Member of Technical Staff (iOS Engineer, Computer Growth)
Senior
Member of Technical Staff (Software Engineer, Enterprise Adoption)
Senior
Member of Technical Staff (Software Engineer, Agentic Enterprise)
Senior
Member of Technical Staff (Software Engineer, Connector Platform)
Senior