Senior/Staff Software Engineer, Developer Experience

Staff Software Engineer · Staff · Full Time · Remote

London · RemoteUSD 215k – 325k2mo ago
Apply for this role

Opens Cohere's application page

Role

What you'll do.

Senior/Staff Software Engineer specializing in Developer Experience at Cohere, a leading enterprise AI company. This role involves designing and implementing robust automation infrastructure, CI/CD pipelines, and testing frameworks that empower engineering teams to ship with confidence. You'll be responsible for building scalable systems for test automation, environment management, and performance benchmarking across diverse configurations using cutting-edge technologies like GitHub Actions, Kubernetes, and infrastructure-as-code tools.

Responsibilities

  • Design and Implement Automation Pipelines: Architect and build comprehensive automation pipelines supporting multi-environment testing with varying feature flags and realistic customer data profiles. Design systems that handle diverse configuration combinations and enable parallel execution across cloud infrastructure.
  • Develop Testing Agents and Intelligence: Create intelligent testing agents that simulate realistic user behavior patterns to validate different configuration combinations. Build frameworks that can automatically detect regressions and flag anomalies across diverse usage scenarios in production-like environments.
  • GitHub Workflows and CI/CD Orchestration: Build and maintain sophisticated GitHub workflows and actions to automate testing, deployment, and validation processes. Establish patterns and templates that enable other engineering teams to self-serve their testing and deployment needs with confidence.
  • Infrastructure-as-Code and Containerization: Develop infrastructure-as-code templates and configurations for reproducible test environments. Implement containerization strategies using Docker and Kubernetes, manage Helm charts for deployment consistency, and implement ArgoCD workflows for continuous deployment across environments.
  • Performance and Reliability Frameworks: Create comprehensive benchmarking frameworks to measure performance and reliability across different configurations and cloud environments. Monitor and improve test coverage metrics, establish performance baselines, and generate detailed analytical reports for stakeholder communication.
  • Establish Testing Best Practices: Define and evangelize testing methodologies and best practices across all engineering teams. Build testing infrastructure that enables individual engineers to own quality themselves, removing barriers and reducing the complexity of writing and executing comprehensive test suites.
  • Scalable Infrastructure Development: Build scalable infrastructure supporting parallel test execution across diverse configurations and accommodate growing customer base demands. Optimize resource utilization, implement efficient caching strategies, and design systems capable of handling exponential growth in test volume.
  • Cross-Team Collaboration and Troubleshooting: Partner with product and engineering teams to understand testing requirements and translate them into automated solutions. Troubleshoot and resolve complex testing infrastructure issues, serve as escalation point for infrastructure challenges, and mentor teams on best practices.

Qualifications

What we look for.

Technical

  • CI/CD Pipeline Architecture

    Deep expertise designing and maintaining CI/CD pipelines that support complex automated testing workflows, deployment orchestration, and multi-environment management at scale.

  • Containerization and Orchestration

    Advanced proficiency with Docker for containerization and Kubernetes for orchestration, including experience managing containerized test environments, pod networking, and resource optimization.

  • Infrastructure-as-Code Principles

    Strong understanding of IaC principles and hands-on experience with Helm charts for Kubernetes configuration management, enabling reproducible and version-controlled infrastructure.

  • Testing Frameworks and Methodologies

    Deep knowledge of comprehensive testing methodologies including unit, integration, end-to-end, and performance testing. Experience building and extending testing frameworks that support complex validation scenarios.

  • Cloud Platform Architecture

    Demonstrated expertise with at least one major cloud platform (AWS, GCP, or Azure), including networking, storage, compute resources, and cost optimization strategies for testing infrastructure.

Education

  • Computer Science or Related Field

    Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, or equivalent practical experience demonstrating mastery of core software engineering principles and system design.

Experience

  • Automation and Testing Infrastructure

    Minimum 5+ years of dedicated software engineering experience with primary focus on automation and testing infrastructure. Proven track record of designing and implementing systems that have improved test coverage, reliability, and developer productivity at scale.

  • Platform and Developer Tools Engineering

    Background in platform engineering or developer tools with demonstrated ability to create internal tools and infrastructure that empower engineering teams. Experience building systems that abstract complexity and enable self-service capabilities.

  • Performance and Benchmarking

    Hands-on experience implementing performance testing methodologies, benchmarking frameworks, and profiling tools. Demonstrated ability to measure, analyze, and communicate performance characteristics across diverse environments.

Skills

Required

  • Python

    Expert-level proficiency with Python for building automation scripts, testing frameworks, and infrastructure tooling. Ability to design clean, maintainable code with strong testing practices.

  • TypeScript

    Expert-level proficiency with TypeScript for backend services, automation scripts, and GitHub Actions. Strong understanding of type systems and ability to build robust, maintainable code at scale.

  • GitHub Actions and Workflows

    Extensive hands-on experience designing and implementing sophisticated GitHub Actions workflows. Deep understanding of workflow syntax, matrix strategies, caching, secrets management, and integration with external services.

  • Docker and Kubernetes

    Advanced proficiency building and optimizing Docker images, managing container registries, and orchestrating applications with Kubernetes. Experience with StatefulSets, DaemonSets, ConfigMaps, and resource management.

  • Helm Charts

    Strong working knowledge of Helm for Kubernetes package management. Experience creating and maintaining production-grade Helm charts, managing dependencies, and implementing templating strategies.

  • ArgoCD and Continuous Deployment

    Practical experience implementing and maintaining ArgoCD for declarative continuous deployment. Understanding of GitOps principles and ability to manage multi-environment deployments through version-controlled configurations.

  • Test Automation Frameworks

    Proficiency building and maintaining test automation frameworks. Experience with test suite organization, parallel execution strategies, flakiness detection, and test result aggregation.

  • System Design and Architecture

    Strong system design skills with ability to architect complex automation systems, anticipate scalability challenges, and make thoughtful technology trade-offs for infrastructure solutions.

Preferred

  • LLM Production Experience

    Nice to have

    Hands-on experience working with Large Language Models in production environments, including understanding model deployment, inference optimization, and quality testing for AI-powered systems.

  • Terraform or Pulumi

    Nice to have

    Experience with infrastructure-as-code tools like Terraform or Pulumi for provisioning and managing cloud infrastructure in a version-controlled, reproducible manner.

  • Performance Testing Tools

    Nice to have

    Familiarity with performance testing frameworks and tools such as k6, JMeter, Locust, or similar platforms for load testing and performance benchmarking.

  • Monitoring and Observability

    Nice to have

    Knowledge of monitoring, logging, and observability tools like Prometheus, Grafana, DataDog, or ELK stack for instrumenting infrastructure and debugging production issues.

  • Test Framework Development

    Nice to have

    Background in developing custom testing frameworks or extending existing frameworks to support specialized testing requirements and methodologies.

  • API Design and Integration

    Nice to have

    Experience designing and consuming APIs, building integration testing frameworks, and orchestrating complex testing workflows across microservices architectures.

Tech stack

Languages

PythonTypeScriptYAMLBash/Shell

Frameworks

KubernetesHelmGitHub ActionsArgoCDDocker

Databases

PostgreSQLRedis

Tools

GitTerraformPulumiAWS/GCP/AzurePrometheusGrafana

Other

CI/CD ArchitectureTest Automation StrategyPerformance BenchmarkingMulti-Environment Configuration Management

Compensation

Pay and benefits.

Base·USD 215,000 – 325,000

Equity·Stock options

Benefits

  • Weekly Lunch Stipend

    Weekly lunch allowance of $75 USD or equivalent in your local currency, supporting convenient meal solutions during workdays

  • Comprehensive Health Coverage

    Full health and dental benefits with a separate dedicated budget for mental health support, ensuring comprehensive wellbeing coverage

  • Retirement Benefits

    RRSP matching for Canadian employees, 401K for US employees, and Pension Scheme for UK employees with company contributions

  • Parental Leave

    100% parental leave top-up for up to 6 months for either parent, supporting work-life balance and family priorities

  • Annual Enrichment Benefits

    Annual budgets for arts and culture, fitness and wellness initiatives, quality time activities, and workspace improvement credits

  • Professional Development Stipend

    Education and learning budget for attending conferences, taking courses, and engaging with professional coaches to support continuous growth

  • Generous Vacation

    6 weeks of paid vacation (30 working days annually), providing substantial time for rest and personal pursuits

  • Office Network and Remote Benefits

    Access to Cohere offices globally (Toronto, London, NYC, San Francisco, Montreal, Paris, Berlin, Seoul) with daily lunch programs and community events; co-working stipend for remote employees

  • Home Office Setup

    Annual $500 home office stipend to properly equip your workspace for optimal productivity and comfort

  • Travel and Company Offsite

    Budget for traveling to other Cohere offices when working remotely, plus participation in annual company offsites for team building and connection

Process

Interview steps.

  1. 01

    Initial Application Review

    Your application will be reviewed by Cohere's recruiting team. The company may use AI-enabled screening tools to assess candidates against the technical and experience criteria, though this does not limit manual review of applications.

  2. 02

    Recruiter Screening Call

    If selected, you'll have a conversation with a Cohere recruiter to discuss your background, experience with testing infrastructure and automation, and alignment with the role's requirements.

  3. 03

    Technical Screening

    Technical assessment focusing on your expertise in automation, CI/CD systems, and infrastructure design. You may be asked to discuss specific projects, architectural decisions, and your approach to building scalable systems.

  4. 04

    System Design Interview

    In-depth discussion of system design capabilities, including how you would architect testing infrastructure, handle scalability challenges, and implement automation solutions for complex requirements.

  5. 05

    Engineering Team Interviews

    Conversations with members of the engineering team and platform infrastructure group to discuss collaboration style, problem-solving approaches, and how you would contribute to the team's goals.

  6. 06

    Leadership Discussion

    Meeting with engineering leadership to discuss vision for the role, your experience leading infrastructure initiatives, and how you would shape testing culture and best practices across teams.

  7. 07

    Final Offer Stage

    Upon successful completion of interviews, Cohere will provide a comprehensive offer including base compensation, equity consideration, and benefits package tailored to your location.

Full posting

Original listing.

Who are we?

Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.

We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that.

We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft.

We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us!

About the Role

We're seeking a Senior/Staff Engineer to build and maintain the automation infrastructure that powers the development cycles of our North platform. This engineer will design and implement robust automation systems that enable engineers to efficiently test and validate changes across diverse environments and configurations. This role sits at the intersection of infrastructure and standards. You'll build the systems, frameworks, and culture that allow the rest of engineering to own quality themselves; improving and extending our testing platform by creating the infrastructure that allows engineers to write and execute tests, and enable every engineering team to ship with more confidence.

Key Responsibilities

  • Design and implement automation pipelines that support comprehensive testing across multiple environments with varying feature flags and realistic customer data profiles

  • Create intelligent testing agents that simulate real user behavior to validate different configuration combinations

  • Develop and maintain GitHub workflows and actions to automate testing, deployment, and validation processes

  • Manage and optimize Helm charts for deployment consistency across environments

  • Implement and maintain ArgoCD workflows for continuous deployment and environment management

  • Establish best practices for testing methodologies and ensure adoption across engineering teams

  • Build scalable infrastructure that supports parallel test execution across diverse configurations

  • Develop infrastructure-as-code templates and configurations for reproducible test environments

  • Implement containerization strategies for test environments and dependencies

  • Create benchmarking frameworks to measure performance and reliability across different configurations

  • Monitor and improve test coverage and reliability metrics

  • Collaborate with product and engineering teams to understand testing requirements and translate them into automated solutions

  • Troubleshoot and resolve complex testing infrastructure issues

Required Qualifications

  • 5+ years of software engineering experience with a focus on automation and testing infrastructure

  • Expert proficiency in Python and TypeScript

  • Extensive experience with GitHub workflows and actions

  • Deep understanding of testing methodologies and best practices

  • Experience building and maintaining CI/CD pipelines

  • Containerization experience (Docker, Kubernetes)

  • Benchmarking experience and performance testing methodologies

  • Cloud platform experience (AWS, GCP, or Azure)

  • Background in developer tools or platform engineering

  • Ability to design and implement complex automation systems

  • Strong problem-solving skills and attention to detail

Preferred Qualifications

  • Experience working with LLMs in production environments

  • Familiarity with infrastructure-as-code principles

  • Experience with container orchestration and management

  • Knowledge of performance testing tools and frameworks

  • Experience with monitoring and observability tools

  • Background in test framework development

  • Strong working knowledge of Helm charts and ArgoCD

  • Infrastructure-as-code experience (Terraform, Pulumi, or similar)

What You'll Build

You'll enhance our testing platform to allow engineers to:

  • Spin up environments with specific feature flag combinations using infrastructure-as-code

  • Load test configurations with realistic customer data volumes in containerized environments

  • Run comprehensive test suites across multiple environment configurations with automated benchmarking

  • Generate detailed performance and reliability reports across different cloud environments

  • Automatically detect and flag regressions in diverse usage scenarios

  • Scale testing infrastructure to accommodate our growing customer base

This is a critical role for ensuring the reliability and scalability of our North platform as we continue to grow our customer base and expand our feature set. You'll have the opportunity to shape the future of our development infrastructure and make a significant impact on product quality. (edited)

Compensation

  • For candidates based in California, New York and Washington States, the compensation range is: $215,000 – $325,000

  • For candidates based elsewhere in the US, the compensation range is: $180,000 – $275,000

  • For candidates based in Canada, the compensation range is: $260,000 – $385,000.

Full-Time Employees at Cohere enjoy these Perks:

  • A weekly lunch stipend of $75/£75 or equivalent in your local currency for lunch.

  • Full health and dental benefits, including a separate budget for mental health.

  • RRSP matching, 401K, Pension Scheme.

  • 100% Parental Leave top-up for up to 6 months, for either parent.

  • Annual enrichment benefits:

    Arts & culture, fitness/wellness, quality time, and a workspace improvement credit.

    Education & learning stipend for conferences, courses, and coaching.

  • 6 weeks of paid vacation (30 working days!)

  • Budget for traveling to other offices if you are remote, plus an annual company offsite.

How and Where We Work:

  • Cohere is remote-friendly, but we also have offices in Toronto, London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul with more opening soon.

  • For those in the office: a daily lunch program, plenty of snacks, and regular community and social events.

  • For those not near an office: a co-working benefit so you can work alongside others in your city.

  • Everyone receives a $500 home office stipend to set up your workspace properly.

If any of the above doesn’t line up exactly with your experience, we still encourage you to apply.


We strive to create an inclusive work environment for all; we welcome applicants from all backgrounds and are committed to providing equal opportunities. Should you require any accommodations during the recruitment process, please submit an Accommodations Request Form, and we will work together to meet your needs.

We may use AI-enabled tools to screen and assess applicants against the criteria for this position. This helps our recruiters identify potentially qualified candidates, but it doesn't limit the applications our recruiters may review or consider.

Beware of Scams: Cohere will never ask for payment or third-party services (e.g., CV writing) as part of our hiring process. All legitimate roles are listed on the Cohere careers page and LinkedIn only, with all communications from Cohere employees coming from an @cohere.com or @cw.cohere email alias. If jobs are viewed on other sites then please verify these through our official careers page.

Redirects to Cohere's application page.

Other roles

More at Cohere.

View all 31 roles