# Senior Software Engineer - Streaming AI (Remote - Ontario / British Columbia)
**Company:** [Confluent](https://scaleengineer.com/companies/confluent)
As a Senior Software Engineer on Confluent's AI streaming team, you will architect and operate large-scale, high-performance infrastructure enabling AI inference directly on real-time data streams. This role focuses on building distributed systems that eliminate the need for external ML stacks, requiring expertise in cloud infrastructure, systems reliability, and scalable serving layers across AWS, Azure, and GCP. You'll tackle complex challenges in networking, compute optimization, and security while collaborating with core platform teams to deliver production-grade AI capabilities at scale.
**Role:** Senior Software Engineer
**Seniority:** Senior
**Locations:** Remote, Ontario, Canada
**Remote:** yes
**Salary:** 144200–169400 CAD
[Apply](https://jobs.ashbyhq.com/confluent/ce64fe49-184c-44b1-9abf-dc5ce0dc9381)
Canonical: https://scaleengineer.com/jobs/confluent/senior-software-engineer-streaming-ai-remote-ontario-british-columbia
---
## Responsibilities

- Infrastructure Architecture and Design: Design, develop, and operate large-scale, high-performance infrastructure that powers Confluent Cloud's AI streaming capabilities. Take ownership of end-to-end product architecture from user-facing APIs down to production serving layers, ensuring systems are optimized for reliability, scalability, and cost efficiency in handling complex inference workloads.
- Distributed Systems Development: Build foundational software addressing distributed systems challenges including consensus algorithms, failover strategies, resource allocation, and state management. Work on systems that enable reliable AI agent execution on streaming data without requiring data movement to external platforms.
- Multi-Cloud Platform Optimization: Troubleshoot, improve, and optimize system reliability, observability, and performance across AWS, Azure, and GCP environments. Ensure consistent performance characteristics and cost efficiency across cloud providers while maintaining high availability standards.
- Cross-Team Collaboration and Integration: Collaborate with Confluent's broader platform teams whose foundational systems you build upon. Coordinate infrastructure enhancements for real-time data streaming use cases, participate in architectural reviews, and drive integration between AI serving infrastructure and core data plane systems.
- Production Operations and Reliability: Own the operational aspects of production AI serving infrastructure, establishing observability standards, defining reliability metrics, and implementing automated solutions for system monitoring. Drive continuous improvements in system resilience and incident response procedures for production environments.

## Requirements

### education

- {"name":"Computer Science or Related Degree","description":"BS, MS, or PhD in computer science, electrical engineering, mathematics, or related field preferred. Equivalent professional experience demonstrating mastery of core computer science concepts is acceptable."}

### technical

- {"name":"Distributed Systems Fundamentals","description":"Strong foundational knowledge of distributed systems principles including consistency models, fault tolerance, consensus algorithms, and distributed tracing. Understanding of networking protocols, load balancing, and inter-service communication patterns."}
- {"name":"Statically Typed Programming Languages","description":"Proficiency in Java, Scala, C++, Go, or equivalent statically typed languages used for building production infrastructure. Ability to write clean, maintainable code optimized for performance-critical systems."}
- {"name":"Cloud Infrastructure and Kubernetes","description":"Working knowledge of cloud infrastructure concepts including containerization, orchestration platforms, networking architectures, and infrastructure-as-code practices for managing complex deployments."}
- {"name":"Systems Performance and Optimization","description":"Understanding of performance optimization techniques including memory management, CPU efficiency, I/O optimization, and profiling tools. Ability to identify and resolve bottlenecks in high-throughput systems."}

### experience

- {"name":"Backend Systems Development","description":"2-5 years of industry experience designing, building, and supporting backend systems in production environments. Demonstrated track record of shipping complex distributed systems that handle significant scale and traffic."}
- {"name":"Large-Scale Infrastructure Operations","description":"Proven experience building and operating large-scale, high-availability systems with expertise in maintaining system reliability, performance monitoring, and incident response in production environments."}
- {"name":"Cloud Platform Experience","description":"Hands-on experience with at least one major cloud provider (AWS, Azure, or GCP), including understanding of cloud services architecture, deployment patterns, and cost optimization strategies."}

## Skills

### required

- {"name":"Distributed Systems Design","description":"Ability to architect reliable, scalable systems handling consensus, failover, state management, and data consistency across multiple nodes and geographical regions."}
- {"name":"Backend Development in Statically Typed Languages","description":"Production-level proficiency in Java, Scala, C++, Go, or similar languages with deep understanding of memory management, concurrency patterns, and performance characteristics."}
- {"name":"Cloud Platform Architecture","description":"Practical expertise with AWS, Azure, or GCP including services like compute instances, databases, networking, load balancing, and monitoring. Understanding of multi-cloud deployment strategies."}
- {"name":"Systems Reliability and Observability","description":"Experience building observable systems with comprehensive logging, metrics, distributed tracing, and alerting. Expertise in on-call practices, incident response, and postmortem-driven improvements."}
- {"name":"High-Performance Systems Engineering","description":"Proven ability to optimize systems for throughput, latency, and resource efficiency. Skilled in performance profiling, bottleneck identification, and implementing architectural improvements for scale."}

### preferred

- {"name":"Model Serving Infrastructure","description":"Experience building or operating platforms for serving machine learning models at scale, including understanding of inference optimization, batching, and resource management for ML workloads."}
- {"name":"LLM and AI Agent Infrastructure","description":"Exposure to systems supporting large language models or autonomous agents, including challenges around prompt caching, context management, or multi-turn interactions at scale."}
- {"name":"Streaming Data Systems","description":"Experience with streaming platforms like Apache Kafka, AWS Kinesis, or similar systems. Understanding of stream processing patterns, state management, and exactly-once semantics in distributed contexts."}
- {"name":"Apache Kafka Ecosystem","description":"Familiarity with Kafka architecture, including broker operations, topic management, client implementations, or contributing to Kafka-based systems. Understanding of Kafka internals and operational patterns."}
- {"name":"Security in Distributed Systems","description":"Understanding of security considerations in cloud infrastructure including authentication, authorization, encryption in transit and at rest, and compliance requirements for production environments."}

## Tech stack

### tools

- {"name":"Kubernetes","description":"Container orchestration platform for managing containerized workloads at scale across multiple cloud providers with automatic scaling and self-healing capabilities."}
- {"name":"Docker","description":"Containerization platform enabling consistent deployment of applications across development, testing, and production environments in cloud infrastructure."}
- {"name":"Terraform","description":"Infrastructure-as-code tool for provisioning and managing cloud resources declaratively across AWS, Azure, and GCP with version control support."}
- {"name":"Prometheus","description":"Time-series database and monitoring system for collecting metrics from distributed systems with powerful querying and alerting capabilities."}
- {"name":"Grafana","description":"Data visualization and dashboarding platform for monitoring and observability, enabling real-time insights into system performance and health."}
- {"name":"Jenkins","description":"Continuous integration and deployment automation tool for building, testing, and releasing infrastructure components with repeatable deployment pipelines."}

### others

- {"name":"AWS (Amazon Web Services)","description":"Cloud computing platform providing compute, networking, storage, and machine learning services. Experience with EC2, S3, VPC, RDS, Lambda, and container services."}
- {"name":"Microsoft Azure","description":"Enterprise cloud platform offering compute, database, storage, and AI services. Familiarity with Virtual Machines, Azure Kubernetes Service, Cosmos DB, and networking services."}
- {"name":"Google Cloud Platform (GCP)","description":"Data and analytics-focused cloud platform providing Compute Engine, Cloud Storage, BigQuery, Pub/Sub, and Kubernetes Engine for distributed infrastructure."}
- {"name":"REST APIs and HTTP Protocols","description":"Expertise in designing and implementing RESTful API architectures, HTTP protocol optimization, and async communication patterns for distributed systems."}
- {"name":"Consensus Algorithms","description":"Understanding of consensus mechanisms including Raft, Paxos, and quorum-based approaches for achieving consistency in distributed systems."}

### databases

- {"name":"PostgreSQL","description":"Production-grade relational database for storing configuration, state, and metadata in Confluent Cloud infrastructure with ACID guarantees and reliability."}
- {"name":"Apache Kafka","description":"Core distributed streaming platform underlying Confluent's architecture for handling real-time data with durability, replication, and fault tolerance guarantees."}
- {"name":"Elasticsearch","description":"Distributed search and analytics engine used for logs, metrics, and observability data in large-scale production environments."}

### languages

- {"name":"Java","description":"Primary backend language for building scalable, production-grade infrastructure at Confluent with excellent ecosystem for distributed systems and high-performance computing."}
- {"name":"Scala","description":"Functional programming language on the JVM used for systems requiring strong type safety and concurrent processing capabilities in infrastructure development."}
- {"name":"Go","description":"Modern systems programming language preferred for microservices, cloud-native tooling, and high-performance networking applications in cloud infrastructure."}
- {"name":"C++","description":"Systems programming language enabling maximum performance optimization for latency-sensitive and throughput-intensive infrastructure components."}

### frameworks

- {"name":"Spring Framework","description":"Widely-used Java framework for building robust backend services with dependency injection, transaction management, and ecosystem support for cloud deployments."}
- {"name":"gRPC","description":"High-performance RPC framework using Protocol Buffers for efficient inter-service communication in distributed systems with strong typing and multiplexing support."}
- {"name":"Netty","description":"Asynchronous event-driven networking framework for Java used in building high-performance, non-blocking I/O applications for streaming systems."}

## Benefits

### benefits

- {"name":"Comprehensive Health Coverage","description":"Medical, dental, and vision insurance plans covering employees and their families with options for various coverage levels and low out-of-pocket costs."}
- {"name":"Retirement Planning","description":"Competitive 401(k) matching program (US) or equivalent RRSP matching for Canadian employees, supporting long-term financial security and wealth building."}
- {"name":"Flexible Work Arrangements","description":"Remote-first role with flexibility to work from Ontario or British Columbia. Opportunity to maintain work-life balance while accessing Confluent's collaborative culture globally."}
- {"name":"Professional Development","description":"Learning stipends, internal training programs, technical certifications, and conference attendance support to advance expertise in distributed systems and cloud technologies."}
- {"name":"Stock Options and Equity","description":"Participation in company equity programs providing ownership stake and long-term wealth creation opportunities as Confluent grows."}
- {"name":"Paid Time Off","description":"Generous vacation policy, paid sick leave, and company holidays ensuring adequate time for rest, recovery, and personal pursuits throughout the year."}
- {"name":"Life and Disability Insurance","description":"Life insurance coverage and short/long-term disability protection providing financial security for you and your family in unexpected circumstances."}
- {"name":"Wellness Programs","description":"Mental health resources, fitness stipends, wellness initiatives, and employee assistance programs supporting holistic health and wellbeing."}
- {"name":"Parental Leave","description":"Competitive parental leave policies supporting work-life balance for growing families with paid time off and flexible return-to-work options."}

## Compensation

- **max:** 240000
- **min:** 165000
- **currency:** CAD
- **stockOptions:** true

## Interview process

### steps

- {"name":"Initial Screening Call","description":"Conversation with recruiter to discuss career goals, experience with distributed systems and cloud infrastructure, and alignment with Confluent's AI streaming mission. This 30-minute call establishes fit and clarifies role expectations."}
- {"name":"Technical Architecture Discussion","description":"45-60 minute conversation with senior engineering team member exploring your experience designing distributed systems, handling scale challenges, and operating production infrastructure. Expect questions about past architectural decisions and trade-offs."}
- {"name":"Coding and Systems Design","description":"Technical assessment involving practical problem-solving around distributed systems design, potentially including a take-home assignment or live coding session focused on infrastructure patterns rather than leetcode-style problems."}
- {"name":"System Design Deep Dive","description":"In-depth 60-90 minute technical interview with 2-3 engineers discussing large-scale system design, cloud infrastructure challenges, and how you would architect solutions for the specific problems the AI team faces at Confluent."}
- {"name":"Cross-Functional Collaboration","description":"Conversation with engineers from platform teams to assess collaboration style, communication skills, and ability to work effectively across team boundaries. Discussion of cross-team challenges and integration patterns."}
- {"name":"Leadership and Values Alignment","description":"Final round with hiring manager or director covering leadership approach, growth mindset, how you navigate ambiguity, and alignment with Confluent's culture of ownership, transparency, and collaborative problem-solving."}

## Full description
We’re not just building better tech. We’re rewriting how data moves and what the world can do with it. With Confluent, data doesn’t sit still. Our platform puts information in motion, streaming in near real-time so companies can react faster, build smarter, and deliver experiences as dynamic as the world around them.

It takes a certain kind of person to join this team. Those who ask hard questions, give honest feedback, and show up for each other. No egos, no solo acts. Just smart, curious humans pushing toward something bigger, together.

One Confluent. One Team. One Data Streaming Platform.

## **About the Team:**

We're a small, focused engineering team building the AI capabilities of Confluent Cloud. Our job is to make it possible to run machine learning and AI agents directly on real-time data — without customers having to stitch together a separate stack to do it. We own our products end to end, from the user-facing API down to the serving layer that runs inference in production, and we work closely with the broader platform teams whose systems we build on top of. It's a high-ownership, high-autonomy environment: small enough that what you build ships and matters, broad enough that the problems are genuinely hard.  
  
## **About the Role:**

As a Software Engineer on the AI team, you will take ownership of the infrastructure that enables our "in-place, at-scale" AI value proposition. You aren't just building a feature; you are architecting the systems that allow customers to run complex inference and AI agents directly on streaming data, eliminating the need to move data to external stacks. This is a high-impact role where your work on scalable, cost-efficient serving layers directly drives Confluent's growth and redefines the possibilities of real-time data.

We are looking for engineers who thrive on the technical complexity of large-scale distributed systems. You will tackle deep infrastructure challenges across networking, compute, and security to ensure our AI capabilities are as reliable as the core data plane itself. If you are passionate about building the foundational systems that power the next generation of AI in the cloud, this is the place to do it.

## **What You Will Do:**

* Design, develop, and operate large-scale, high-performance infrastructure that powers Confluent Cloud.
* Build foundational software to improve reliability, scalability, and efficiency across cloud environments.
* Work on distributed systems challenges such as consensus algorithms, failover strategies, and resource allocation.
* Collaborate with teams across Confluent to optimize and enhance infrastructure for real-time data streaming use cases.
* Troubleshoot and improve system reliability, observability, and performance across multiple cloud providers (AWS, Azure, GCP).

## **What You Will Bring:**

* 2-5 years of industry experience designing, building, and supporting backend systems in production.
* Strong fundamentals in distributed systems, cloud infrastructure, and networking.
* Experience in building and operating large-scale, high-availability systems.
* Good understanding of cloud platforms (AWS, Azure, or GCP) and their services.
* Proficiency in Java, Scala, C++, Go, or other statically typed languages.
* A self-starter with strong problem-solving skills and the ability to work in a fast-paced environment.
* BS, MS, or PhD in computer science or a related field, or equivalent work experience.

## **What Gives You an Edge:**

* Exposure to model serving, LLM/agent infrastructure, or streaming data systems.
* Note - You don't need a background in ML research or model training — this role is about building and operating the platform that serves AI reliably at scale, not inventing the models.

## **Ready to build what's next? Let’s get in motion.**

### 

# **Come As You Are**

Belonging isn’t a perk here. It’s the baseline. We work across time zones and backgrounds, knowing the best ideas come from different perspectives. And we make space for everyone to lead, grow, and challenge what’s possible.

We’re proud to be an equal opportunity workplace. Employment decisions are based on job-related criteria, without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, veteran status, or any other classification protected by law.

# **Privacy Statement**

Confluent is an IBM subsidiary which has been acquired by IBM and will be integrated into the IBM organization. By proceeding with this application, you understand that Confluent will share your personal information with other IBM affiliates involved in your recruitment process, wherever these are located. More Information on how IBM protects your personal information, including the safeguards in case of cross-border data transfer, are available [here](http://ibm.com/careers/us-en/privacy-policy/).
