Cohere

Engineering Manager, GPU Infrastructure

Cohere4 days ago
Location

United States

Type

Full Time

Salary

USD 185,000 – 280,000

Level

Manager

Role

Engineering Manager

Posted

Jul 21, 2026

Full TimeManager

The role

Summary

Lead the GPU Clusters team at Cohere, a leading enterprise AI company, managing highly skilled engineers responsible for building and operating superclusters that power frontier large language models. This Engineering Manager role sits at the intersection of cutting-edge hardware, distributed systems, and AI research, requiring deep technical expertise in Kubernetes, GPU infrastructure, and machine learning environments combined with strong team leadership and mentorship capabilities. You'll drive technical strategy, cross-functional collaboration with AI researchers and cloud providers, and operational excellence while fostering a culture of technical innovation in a remote-first, globally distributed organization.

What you'll do

Team Leadership & Technical Mentorship: Lead and mentor a team of specialized GPU infrastructure engineers, establishing a culture of technical excellence, continuous improvement, and professional growth. Conduct regular one-on-one meetings with team members, provide technical guidance on complex distributed systems challenges, manage performance reviews, oversee career development planning, and lead the hiring process to build a high-performing team aligned with Cohere's mission to enable frontier AI model deployment.
Technical Roadmap Development & Execution: Define and execute the comprehensive technical roadmap for GPU cluster deployment, optimization, and scaling across Cohere's infrastructure. Oversee implementation of topology-aware scheduling systems, hardware fault detection mechanisms, and performance optimization initiatives. Drive decisions on new technologies and architectural approaches while balancing innovation with operational stability and cost efficiency.
GPU Cluster Deployment & Infrastructure Operations: Ensure reliable, scalable, and secure GPU infrastructure across all environments, overseeing deployment strategies and operational procedures. Collaborate with cloud providers to validate and deploy emerging GPU architectures, manage infrastructure provisioning, and maintain comprehensive monitoring and observability systems for GPU utilization, performance metrics, and reliability tracking.
Infrastructure-as-Code & Automation Implementation: Drive adoption of infrastructure-as-code practices using tools like Terraform and configuration management systems. Establish automation frameworks for cluster provisioning, deployment pipelines, and operational procedures. Implement monitoring solutions using Prometheus, Grafana, or similar tools to track infrastructure health and performance, reducing manual operational overhead and improving system reliability.
Cross-Functional Collaboration & Stakeholder Management: Partner with AI researchers and the Foundations team to understand emerging infrastructure requirements and translate technical needs into robust, scalable solutions. Coordinate with Capacity EPM on resource delivery timelines and planning, interface with Legal and Security teams on compliance and security requirements, and collaborate with other infrastructure teams to align on shared goals and technical dependencies.
Cost Optimization & Vendor Management: Drive cost optimization initiatives while maintaining performance standards and system reliability. Manage vendor relationships with cloud providers and hardware suppliers, negotiate contracts, and oversee capacity planning to optimize GPU infrastructure spending. Implement monitoring and analysis to track cost drivers and identify efficiency improvements across the cluster infrastructure.
Documentation & Knowledge Management: Ensure comprehensive, up-to-date technical documentation and operational runbooks are maintained and accessible to all stakeholders. Establish knowledge-sharing practices within the team and across infrastructure organizations to enable smooth onboarding, disaster recovery procedures, and institutional continuity of critical infrastructure knowledge.

What we look for

Technical

Deep GPU & HPC Infrastructure ExpertiseDemonstrated mastery of GPU/TPU cluster architecture, high-performance computing environments, and distributed training frameworks including JAX, PyTorch, and TensorFlow. Experience designing and managing infrastructure that supports large-scale machine learning model training and deployment, with understanding of GPU memory management, compute optimization, and multi-GPU synchronization patterns.
Kubernetes at Scale for AI WorkloadsProven production experience deploying, managing, and troubleshooting Kubernetes clusters at enterprise scale, specifically for machine learning and AI workloads in multi-cloud environments. Proficiency with cloud-native deployment patterns, workload scheduling, resource management, and handling stateful applications across multiple cloud providers.
Infrastructure Monitoring & ObservabilityHands-on experience implementing comprehensive monitoring solutions using Prometheus, Grafana, and related observability tools. Capability to design monitoring architectures that track GPU utilization, application performance, infrastructure health metrics, and enable data-driven operational decisions and performance optimization.
Infrastructure-as-Code & Automation ToolsProficiency with infrastructure-as-code tools including Terraform, CloudFormation, or equivalent systems, combined with experience in ArgoCD, CI/CD pipelines, and configuration management. Ability to design and implement reproducible, version-controlled infrastructure automation that enables rapid scaling and disaster recovery.
Distributed Systems & Performance EngineeringStrong understanding of distributed systems principles, networking concepts, and performance optimization techniques relevant to large-scale compute clusters. Experience with system-level troubleshooting, bottleneck identification, and implementing solutions that optimize throughput, latency, and resource utilization in complex distributed environments.
Cloud Platform ArchitectureDeep experience with major cloud platforms (AWS, GCP, Azure) and their infrastructure services, particularly those supporting GPU instances and high-performance computing. Understanding of cloud-native architecture patterns, cost optimization strategies, and multi-cloud deployment considerations.

Education

Bachelor's Degree in Computer Science or Related FieldEducational foundation in Computer Science, Electrical Engineering, Physics, or closely related discipline, providing theoretical grounding in distributed systems, computer architecture, and algorithmic principles applicable to infrastructure engineering.
Continuous Learning in Cloud & Infrastructure TechnologiesCommitment to ongoing professional development through certifications (Kubernetes CKA/CKAD, cloud platform certifications), courses, and technical community engagement to maintain expertise in rapidly evolving infrastructure technologies and best practices.

Experience

5+ Years Managing Engineering TeamsSubstantial experience leading and mentoring engineering teams through multiple project cycles, demonstrating skill in hiring, performance management, career development, and fostering high-performing cultures. Experience managing teams through scaling challenges and organizational growth.
8+ Years GPU/ML Infrastructure EngineeringExtensive hands-on experience building and operating GPU infrastructure for machine learning applications, including experience with multi-cloud deployments, large-scale training infrastructure, and optimization of compute resources for AI workloads. Track record of solving complex infrastructure challenges unique to machine learning environments.
3+ Years Leadership in Distributed Systems or InfrastructureLeadership experience directing infrastructure teams or projects focused on distributed systems, cloud-native architectures, or high-performance computing environments. Demonstrated ability to translate business requirements into technical strategies and execute large-scale infrastructure initiatives.
Collaboration with AI Researchers & ML EngineersProven track record of working effectively with machine learning researchers, data scientists, and ML engineers to understand complex technical requirements and translate them into infrastructure solutions that enable cutting-edge AI research and model training.

Skills

Required skills

Kubernetes Management at ScaleProduction-level expertise deploying and managing Kubernetes clusters for enterprise AI workloads, including cluster scaling, workload scheduling optimization, and troubleshooting distributed container orchestration at scale.
GPU Infrastructure ArchitectureDeep technical knowledge of GPU cluster design, topology-aware scheduling, hardware configuration optimization, and strategies for maximizing compute utilization and throughput in multi-GPU environments.
Distributed Training FrameworksWorking knowledge of PyTorch, TensorFlow, JAX, and other distributed training frameworks used in machine learning infrastructure, including understanding of their infrastructure requirements and optimization opportunities.
Infrastructure Monitoring & ObservabilityProficiency with Prometheus, Grafana, and related monitoring tools to implement comprehensive infrastructure observability, enabling performance tracking and data-driven optimization decisions.
Infrastructure-as-Code (IaC)Hands-on expertise with Terraform, ArgoCD, and related infrastructure-as-code tools for reproducible, version-controlled infrastructure management and deployment automation.
Team Leadership & MentorshipDemonstrated capability to lead technical teams, mentor engineers through challenging problems, provide career development guidance, and establish high-performing team cultures focused on technical excellence.
Cross-Functional CommunicationAbility to translate complex technical concepts for diverse audiences including researchers, business stakeholders, and non-technical partners, facilitating alignment across organizational boundaries.
Data-Driven Decision MakingStrong analytical and problem-solving abilities with demonstrated capability to make informed decisions under pressure using data, metrics, and system observability to guide infrastructure optimization choices.

Nice to have

Cloud Provider Relationships & NegotiationsExperience managing relationships with cloud infrastructure providers, negotiating contracts for cloud services and GPU quotas, and coordinating with vendors on new hardware architectures and deployments.
Multi-Cloud Deployment ExperienceHands-on experience designing and operating infrastructure spanning multiple cloud providers (AWS, GCP, Azure), including strategies for workload distribution, cost optimization, and avoiding vendor lock-in.
Cost Optimization & Capacity PlanningTrack record of implementing cost optimization initiatives for infrastructure while maintaining performance and reliability standards, including capacity forecasting and resource planning for growing compute demands.
Security & Compliance in InfrastructureFamiliarity with security best practices, compliance requirements, and audit procedures relevant to infrastructure management, including containerized environments and data protection considerations.
ML/HPC Infrastructure ResearchExperience with or understanding of emerging infrastructure approaches for machine learning and high-performance computing, including familiarity with academic research and industry innovations in GPU cluster optimization.
Distributed Systems OptimizationExperience identifying and resolving performance bottlenecks in distributed systems, implementing optimizations for throughput and latency in complex compute environments, and designing systems for extreme scale.
Remote Team ManagementProven success leading geographically distributed engineering teams in remote-first environments, establishing communication practices and team cohesion across time zones and locations.
Hardware Fault Detection & RecoveryExperience implementing or managing systems for hardware fault detection, automatic failover procedures, and recovery mechanisms in large-scale infrastructure environments.

Compensation & benefits

Salary

USD 185,000 – 280,000 (annual)


Apply for this position

You'll be redirected to the company's application page