Software Engineer, Data Infrastructure
Backend Engineer · Mid · Full Time
Opens Cohere's application page
Role
What you'll do.
Join Cohere's Data Infrastructure team to design and operate distributed storage systems at petabyte scale that power large-scale AI model training. This role focuses on building unified storage layers using Kubernetes, managing multi-terabyte data movement, and optimizing throughput and latency for GPU clusters, requiring strong fundamentals in distributed systems, storage architecture, and cloud infrastructure. You'll work at the intersection of infrastructure engineering and machine learning operations, solving complex challenges in data consistency, networking, and cross-region data replication.
Responsibilities
- Design and Build Distributed Storage Architecture: Design, architect, and implement the core distributed storage system that feeds model training and evaluation pipelines. This involves creating a unified storage abstraction layer capable of handling petabytes of training data and model checkpoints while maintaining performance, consistency, and durability guarantees.
- Operate Storage Systems at Scale on Kubernetes: Deploy, manage, and operate stateful systems on Kubernetes clusters at petabyte scale. This includes working with Persistent Volumes, Container Storage Interface (CSI) drivers, and StatefulSets to ensure reliable data availability and access patterns for thousands of GPUs across multiple training clusters.
- Optimize Data Movement and I/O Performance: Work through complex networking, I/O bottleneck resolution, and consistency challenges when moving large datasets and model checkpoints across regions and storage backends. Continuously optimize for GPU idle time minimization and time-to-insight metrics, directly impacting model training efficiency.
- Collaborate with ML and Training Infrastructure Teams: Partner with researchers and training infrastructure engineers to understand actual data read/write patterns from model training jobs. Translate workload characteristics into concrete throughput, latency, and durability requirements that drive system design decisions.
- Implement Data Lifecycle and Consistency Management: Develop and maintain systems for data replication, consistency models, caching strategies, and data lifecycle policies. Ensure data durability across multiple regions and cloud backends while optimizing cost and performance tradeoffs.
Qualifications
What we look for.
Technical
Distributed Storage Systems Fundamentals
Demonstrated expertise in storage architecture principles including data replication strategies (synchronous vs. asynchronous), consistency models (strong vs. eventual), caching mechanisms, and data lifecycle management across multiple storage tiers.
Kubernetes and Stateful Systems
Hands-on production experience running stateful workloads on Kubernetes, including Persistent Volumes, Container Storage Interface (CSI) drivers, StatefulSets, and understanding of pod networking, storage provisioning, and resource management.
Cloud Object Storage and Filesystems
Practical experience with cloud object storage systems (particularly Amazon S3) and POSIX-style filesystems. Understanding of object storage semantics, consistency guarantees, performance characteristics, and when to use each storage paradigm.
Programming in Python or Go
Strong coding fundamentals with demonstrated proficiency in either Python or Go. Ability to write efficient, maintainable systems code and willingness to learn the other language as needed for the role.
Distributed Systems and Networking
Solid understanding of distributed systems concepts including eventual consistency, data replication, network protocols, latency optimization, and debugging tools for analyzing I/O and network performance at scale.
Education
Bachelor's Degree in Computer Science or Related Field
Formal education in Computer Science, Computer Engineering, or closely related discipline providing foundational knowledge in systems, algorithms, and data structures. Equivalent practical experience in infrastructure engineering is valued equally.
Systems Programming Knowledge
Understanding of operating system concepts, file I/O, memory management, and low-level system interactions. This can be demonstrated through coursework, projects, or professional experience working with systems-level code.
Experience
Production Systems Engineering
3+ years of experience building and operating production infrastructure systems. Experience operating systems under real-world constraints, managing reliability, debugging complex issues, and understanding performance characteristics under load.
Cloud Infrastructure or Data Systems
Background working with large-scale data systems, cloud platforms (AWS, GCP, Azure), or infrastructure engineering. Experience managing systems that handle significant data volumes, concurrent operations, and strict performance requirements.
Storage or Infrastructure Architecture
Experience designing or contributing to storage layer architecture, data pipelines, or infrastructure systems. Familiarity with tradeoffs between different storage approaches and how to optimize for specific workload patterns.
Skills
Required
Python or Go Programming
Proficiency in Python for system scripting and data pipeline work, or Go for high-performance systems programming. Strong ability to write clean, efficient code with proper error handling and testing.
Kubernetes Administration and Operations
Ability to design, deploy, and troubleshoot Kubernetes clusters with emphasis on stateful workloads. Understanding of networking, storage provisioning, resource management, and operational best practices.
Storage Systems Architecture
Understanding of distributed storage design patterns, replication strategies, consistency models (strong vs. eventual), caching algorithms, and data durability guarantees. Knowledge of tradeoffs between consistency, availability, and partition tolerance.
Cloud Storage Services
Practical experience with cloud object storage (S3) and understanding of cloud storage concepts including buckets, versioning, access control, lifecycle policies, and performance optimization.
Debugging and Performance Analysis
Strong problem-solving skills with ability to debug distributed systems issues. Experience using profiling tools, log analysis, and performance monitoring to identify and resolve bottlenecks in I/O, networking, and CPU-bound operations.
Preferred
Parallel and HPC Filesystems
Nice to haveExperience with high-performance computing filesystems such as Weka, VAST, or Lustre. Understanding of how these systems achieve high throughput and how they're optimized for parallel data access patterns.
Machine Learning Training Infrastructure
Nice to haveFamiliarity with data loading and checkpointing patterns used in large-scale model training. Understanding of how distributed training frameworks (PyTorch, TensorFlow) interact with storage systems and data loading bottlenecks.
Large-Scale ML Model Knowledge
Nice to haveInterest in and basic understanding of how large language models are trained and evaluated. Knowledge of training data characteristics, checkpoint formats, evaluation data generation, and how infrastructure decisions impact training efficiency.
Multi-Region Data Management
Nice to haveExperience managing data replication and consistency across geographic regions. Understanding of latency implications, cost optimization, and disaster recovery patterns in multi-region deployments.
CSI Driver Development
Nice to haveExperience developing or customizing Container Storage Interface drivers for Kubernetes. Understanding of volume lifecycle management and how CSI drivers bridge Kubernetes and external storage systems.
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 150,000 – 220,000
Equity·Stock options
Benefits
Weekly Lunch Stipend
$75 USD weekly (or equivalent in local currency) for lunch support, reducing meal planning burden for in-office and remote employees.
Comprehensive Health and Dental Coverage
Full health and dental benefits with a separate dedicated budget specifically allocated for mental health services and wellness support.
Retirement Savings Programs
RRSP matching for Canadian employees, 401(k) plans for US-based employees, and Pension Scheme access for UK employees, ensuring robust retirement security.
100% Parental Leave Top-up
Full salary continuation for up to 6 months of parental leave for either parent, supporting work-life balance and family planning.
Annual Enrichment Benefits
Designated budgets for arts and culture experiences, fitness and wellness activities, quality time initiatives, and workspace improvement credit to enhance both professional and personal growth.
Professional Development and Learning
Education and learning stipend specifically for attending conferences, taking courses, and engaging executive coaching to support continuous skill development.
Generous Paid Vacation
6 weeks of paid vacation time (30 working days annually), providing substantial time for rest, travel, and personal pursuits.
Global Office Access and Travel
Budget for traveling to other Cohere offices if working remotely, plus attendance at an annual company offsite for team building and company-wide connection.
Home Office Setup Stipend
$500 one-time credit for establishing a proper home office workspace with ergonomic furniture and equipment.
Remote-Friendly Work Environment
Flexible work arrangements with offices in Toronto, London, New York City, San Francisco, Montreal, Paris, Berlin, and Seoul, plus co-working benefits for those outside office locations.
In-Office Perks
For those working in physical offices: daily lunch program, ample snacks, and regular community and social events fostering team connection.
Co-working Membership
For remote employees not near an office: co-working space membership enabling collaboration with other professionals in your local city.
Process
Interview steps.
- 01
Initial Application Review
Your application is reviewed against core qualifications using both AI-enabled screening and human recruiter evaluation. Both resume content and demonstrated experience with storage systems, Kubernetes, and distributed infrastructure are assessed.
- 02
Technical Phone Screening
Initial conversation with a recruiter or engineer focusing on background, current work, experience with distributed systems and storage architecture, and general technical fit for the role.
- 03
Technical Interview - System Design
In-depth technical discussion around distributed storage system design. Expect questions about data replication strategies, consistency models, caching approaches, and how you would architect systems for specific constraints (throughput, latency, durability).
- 04
Technical Interview - Implementation and Kubernetes
Conversation focused on practical implementation details, your experience operating Kubernetes at scale, managing Persistent Volumes, CSI drivers, and StatefulSets. May include coding discussion or whiteboarding system components.
- 05
Infrastructure and Operations Discussion
Detailed discussion of past experience operating production systems, debugging complex infrastructure issues, performance optimization, and how you approach operational excellence and reliability.
- 06
Team and Culture Fit Interview
Conversation with potential team members or managers discussing collaboration style, approach to problem-solving in unfamiliar domains, interest in large-scale AI systems, and alignment with Cohere's values and mission.
- 07
Offer and Negotiation
Upon successful interviews, receive offer details including compensation, equity, benefits, and specific role expectations. Competitive negotiation is welcome for experienced candidates meeting all core requirements.
Full posting
Original listing.
Who are we?
Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.
We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that.
We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft.
We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us!
Why this role?
The Data Infrastructure team at Cohere is responsible for the storage and data movement layer underlying every model training run. We're building the unified storage layer that feeds our training workloads. It needs to serve petabytes of training data and model checkpoints fast enough to keep thousands of GPUs busy across several training clusters. In this role, you’d have an opportunity to build this system from the ground up. You’d be a key contributor, working on a problem few teams have had to solve at this scale.
In this role, you will:
Design, build, and operate the distributed storage system that feeds model training and evaluation.
Run this system multiple on Kubernetes clusters at petabyte scale.
Work with researchers and training-infra teams on how jobs actually read and write data, and turn that into throughput, latency, and durability requirements
Work through the networking, I/O, and consistency problems of moving large datasets and checkpoints across regions and backends, with GPU idle time and time-to-insight as the measures of success
You may be a good fit if you have:
Strong storage fundamentals, including replication, consistency, caching, and data lifecycle management.
Strong coding ability. We work in Python and Go; experience in either is enough, but you should be willing to pick up the other
Experience running stateful systems on Kubernetes, including Persistent Volumes, CSI drivers, and StatefulSets.
Hands-on experience with cloud object storage such as S3 as well as POSIX-style filesystems.
It’s a bonus if you have:
Experience with parallel or HPC filesystems such as Weka, VAST, or Lustre.
Familiarity with the data-loading and checkpointing patterns used in large-scale model training.
An interest in how large models are trained and evaluated, including the trajectory and eval data those runs generate, and a desire to understand the workloads the platform is serving.
Full-Time Employees at Cohere enjoy these Perks:
A weekly lunch stipend of $75/£75 or equivalent in your local currency for lunch.
Full health and dental benefits, including a separate budget for mental health.
RRSP matching, 401K, Pension Scheme.
100% Parental Leave top-up for up to 6 months, for either parent.
Annual enrichment benefits:
Arts & culture, fitness/wellness, quality time, and a workspace improvement credit.
Education & learning stipend for conferences, courses, and coaching.
6 weeks of paid vacation (30 working days!)
Budget for traveling to other offices if you are remote, plus an annual company offsite.
How and Where We Work:
Cohere is remote-friendly, but we also have offices in Toronto, London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul with more opening soon.
For those in the office: a daily lunch program, plenty of snacks, and regular community and social events.
For those not near an office: a co-working benefit so you can work alongside others in your city.
Everyone receives a $500 home office stipend to set up your workspace properly.
If any of the above doesn’t line up exactly with your experience, we still encourage you to apply.
We strive to create an inclusive work environment for all; we welcome applicants from all backgrounds and are committed to providing equal opportunities. Should you require any accommodations during the recruitment process, please submit an Accommodations Request Form, and we will work together to meet your needs.
We may use AI-enabled tools to screen and assess applicants against the criteria for this position. This helps our recruiters identify potentially qualified candidates, but it doesn't limit the applications our recruiters may review or consider.
Beware of Scams: Cohere will never ask for payment or third-party services (e.g., CV writing) as part of our hiring process. All legitimate roles are listed on the Cohere careers page and LinkedIn only, with all communications from Cohere employees coming from an @cohere.com or @cw.cohere email alias. If jobs are viewed on other sites then please verify these through our official careers page.
Redirects to Cohere's application page.
Other roles
More at Cohere.
Product Security Engineer, North Security
Senior
Forward Deployed Engineer, Infrastructure Specialist (Middle East)
Mid
Forward Deployed Engineer, Agentic Platform (West Coast)
Senior
Senior Mobile Engineer (iOS or Android)
Senior
Forward Deployed Engineer, Agentic Platform (UK Public Sector)
Senior