Senior Software Engineer - Streaming AI (Remote - Ontario / British Columbia)

Senior Software Engineer · Senior · Full Time · Remote

Remote, Ontario, Canada · RemoteCAD 144k – 169k1mo ago
Apply for this role

Opens Confluent's application page

Role

What you'll do.

As a Senior Software Engineer on Confluent's AI streaming team, you will architect and operate large-scale, high-performance infrastructure enabling AI inference directly on real-time data streams. This role focuses on building distributed systems that eliminate the need for external ML stacks, requiring expertise in cloud infrastructure, systems reliability, and scalable serving layers across AWS, Azure, and GCP. You'll tackle complex challenges in networking, compute optimization, and security while collaborating with core platform teams to deliver production-grade AI capabilities at scale.

Responsibilities

  • Infrastructure Architecture and Design: Design, develop, and operate large-scale, high-performance infrastructure that powers Confluent Cloud's AI streaming capabilities. Take ownership of end-to-end product architecture from user-facing APIs down to production serving layers, ensuring systems are optimized for reliability, scalability, and cost efficiency in handling complex inference workloads.
  • Distributed Systems Development: Build foundational software addressing distributed systems challenges including consensus algorithms, failover strategies, resource allocation, and state management. Work on systems that enable reliable AI agent execution on streaming data without requiring data movement to external platforms.
  • Multi-Cloud Platform Optimization: Troubleshoot, improve, and optimize system reliability, observability, and performance across AWS, Azure, and GCP environments. Ensure consistent performance characteristics and cost efficiency across cloud providers while maintaining high availability standards.
  • Cross-Team Collaboration and Integration: Collaborate with Confluent's broader platform teams whose foundational systems you build upon. Coordinate infrastructure enhancements for real-time data streaming use cases, participate in architectural reviews, and drive integration between AI serving infrastructure and core data plane systems.
  • Production Operations and Reliability: Own the operational aspects of production AI serving infrastructure, establishing observability standards, defining reliability metrics, and implementing automated solutions for system monitoring. Drive continuous improvements in system resilience and incident response procedures for production environments.

Qualifications

What we look for.

Technical

  • Distributed Systems Fundamentals

    Strong foundational knowledge of distributed systems principles including consistency models, fault tolerance, consensus algorithms, and distributed tracing. Understanding of networking protocols, load balancing, and inter-service communication patterns.

  • Statically Typed Programming Languages

    Proficiency in Java, Scala, C++, Go, or equivalent statically typed languages used for building production infrastructure. Ability to write clean, maintainable code optimized for performance-critical systems.

  • Cloud Infrastructure and Kubernetes

    Working knowledge of cloud infrastructure concepts including containerization, orchestration platforms, networking architectures, and infrastructure-as-code practices for managing complex deployments.

  • Systems Performance and Optimization

    Understanding of performance optimization techniques including memory management, CPU efficiency, I/O optimization, and profiling tools. Ability to identify and resolve bottlenecks in high-throughput systems.

Education

  • Computer Science or Related Degree

    BS, MS, or PhD in computer science, electrical engineering, mathematics, or related field preferred. Equivalent professional experience demonstrating mastery of core computer science concepts is acceptable.

Experience

  • Backend Systems Development

    2-5 years of industry experience designing, building, and supporting backend systems in production environments. Demonstrated track record of shipping complex distributed systems that handle significant scale and traffic.

  • Large-Scale Infrastructure Operations

    Proven experience building and operating large-scale, high-availability systems with expertise in maintaining system reliability, performance monitoring, and incident response in production environments.

  • Cloud Platform Experience

    Hands-on experience with at least one major cloud provider (AWS, Azure, or GCP), including understanding of cloud services architecture, deployment patterns, and cost optimization strategies.

Skills

Required

  • Distributed Systems Design

    Ability to architect reliable, scalable systems handling consensus, failover, state management, and data consistency across multiple nodes and geographical regions.

  • Backend Development in Statically Typed Languages

    Production-level proficiency in Java, Scala, C++, Go, or similar languages with deep understanding of memory management, concurrency patterns, and performance characteristics.

  • Cloud Platform Architecture

    Practical expertise with AWS, Azure, or GCP including services like compute instances, databases, networking, load balancing, and monitoring. Understanding of multi-cloud deployment strategies.

  • Systems Reliability and Observability

    Experience building observable systems with comprehensive logging, metrics, distributed tracing, and alerting. Expertise in on-call practices, incident response, and postmortem-driven improvements.

  • High-Performance Systems Engineering

    Proven ability to optimize systems for throughput, latency, and resource efficiency. Skilled in performance profiling, bottleneck identification, and implementing architectural improvements for scale.

Preferred

  • Model Serving Infrastructure

    Nice to have

    Experience building or operating platforms for serving machine learning models at scale, including understanding of inference optimization, batching, and resource management for ML workloads.

  • LLM and AI Agent Infrastructure

    Nice to have

    Exposure to systems supporting large language models or autonomous agents, including challenges around prompt caching, context management, or multi-turn interactions at scale.

  • Streaming Data Systems

    Nice to have

    Experience with streaming platforms like Apache Kafka, AWS Kinesis, or similar systems. Understanding of stream processing patterns, state management, and exactly-once semantics in distributed contexts.

  • Apache Kafka Ecosystem

    Nice to have

    Familiarity with Kafka architecture, including broker operations, topic management, client implementations, or contributing to Kafka-based systems. Understanding of Kafka internals and operational patterns.

  • Security in Distributed Systems

    Nice to have

    Understanding of security considerations in cloud infrastructure including authentication, authorization, encryption in transit and at rest, and compliance requirements for production environments.

Tech stack

Languages

JavaScalaGoC++

Frameworks

Spring FrameworkgRPCNetty

Databases

PostgreSQLApache KafkaElasticsearch

Tools

KubernetesDockerTerraformPrometheusGrafanaJenkins

Other

AWS (Amazon Web Services)Microsoft AzureGoogle Cloud Platform (GCP)REST APIs and HTTP ProtocolsConsensus Algorithms

Compensation

Pay and benefits.

Base·CAD 144,200 – 169,400

Equity·Stock options

Benefits

  • Comprehensive Health Coverage

    Medical, dental, and vision insurance plans covering employees and their families with options for various coverage levels and low out-of-pocket costs.

  • Retirement Planning

    Competitive 401(k) matching program (US) or equivalent RRSP matching for Canadian employees, supporting long-term financial security and wealth building.

  • Flexible Work Arrangements

    Remote-first role with flexibility to work from Ontario or British Columbia. Opportunity to maintain work-life balance while accessing Confluent's collaborative culture globally.

  • Professional Development

    Learning stipends, internal training programs, technical certifications, and conference attendance support to advance expertise in distributed systems and cloud technologies.

  • Stock Options and Equity

    Participation in company equity programs providing ownership stake and long-term wealth creation opportunities as Confluent grows.

  • Paid Time Off

    Generous vacation policy, paid sick leave, and company holidays ensuring adequate time for rest, recovery, and personal pursuits throughout the year.

  • Life and Disability Insurance

    Life insurance coverage and short/long-term disability protection providing financial security for you and your family in unexpected circumstances.

  • Wellness Programs

    Mental health resources, fitness stipends, wellness initiatives, and employee assistance programs supporting holistic health and wellbeing.

  • Parental Leave

    Competitive parental leave policies supporting work-life balance for growing families with paid time off and flexible return-to-work options.

Process

Interview steps.

  1. 01

    Initial Screening Call

    Conversation with recruiter to discuss career goals, experience with distributed systems and cloud infrastructure, and alignment with Confluent's AI streaming mission. This 30-minute call establishes fit and clarifies role expectations.

  2. 02

    Technical Architecture Discussion

    45-60 minute conversation with senior engineering team member exploring your experience designing distributed systems, handling scale challenges, and operating production infrastructure. Expect questions about past architectural decisions and trade-offs.

  3. 03

    Coding and Systems Design

    Technical assessment involving practical problem-solving around distributed systems design, potentially including a take-home assignment or live coding session focused on infrastructure patterns rather than leetcode-style problems.

  4. 04

    System Design Deep Dive

    In-depth 60-90 minute technical interview with 2-3 engineers discussing large-scale system design, cloud infrastructure challenges, and how you would architect solutions for the specific problems the AI team faces at Confluent.

  5. 05

    Cross-Functional Collaboration

    Conversation with engineers from platform teams to assess collaboration style, communication skills, and ability to work effectively across team boundaries. Discussion of cross-team challenges and integration patterns.

  6. 06

    Leadership and Values Alignment

    Final round with hiring manager or director covering leadership approach, growth mindset, how you navigate ambiguity, and alignment with Confluent's culture of ownership, transparency, and collaborative problem-solving.

Full posting

Original listing.

We’re not just building better tech. We’re rewriting how data moves and what the world can do with it. With Confluent, data doesn’t sit still. Our platform puts information in motion, streaming in near real-time so companies can react faster, build smarter, and deliver experiences as dynamic as the world around them.

It takes a certain kind of person to join this team. Those who ask hard questions, give honest feedback, and show up for each other. No egos, no solo acts. Just smart, curious humans pushing toward something bigger, together.

One Confluent. One Team. One Data Streaming Platform.

About the Team:

We're a small, focused engineering team building the AI capabilities of Confluent Cloud. Our job is to make it possible to run machine learning and AI agents directly on real-time data — without customers having to stitch together a separate stack to do it. We own our products end to end, from the user-facing API down to the serving layer that runs inference in production, and we work closely with the broader platform teams whose systems we build on top of. It's a high-ownership, high-autonomy environment: small enough that what you build ships and matters, broad enough that the problems are genuinely hard.

About the Role:

As a Software Engineer on the AI team, you will take ownership of the infrastructure that enables our "in-place, at-scale" AI value proposition. You aren't just building a feature; you are architecting the systems that allow customers to run complex inference and AI agents directly on streaming data, eliminating the need to move data to external stacks. This is a high-impact role where your work on scalable, cost-efficient serving layers directly drives Confluent's growth and redefines the possibilities of real-time data.

We are looking for engineers who thrive on the technical complexity of large-scale distributed systems. You will tackle deep infrastructure challenges across networking, compute, and security to ensure our AI capabilities are as reliable as the core data plane itself. If you are passionate about building the foundational systems that power the next generation of AI in the cloud, this is the place to do it.

What You Will Do:

  • Design, develop, and operate large-scale, high-performance infrastructure that powers Confluent Cloud.

  • Build foundational software to improve reliability, scalability, and efficiency across cloud environments.

  • Work on distributed systems challenges such as consensus algorithms, failover strategies, and resource allocation.

  • Collaborate with teams across Confluent to optimize and enhance infrastructure for real-time data streaming use cases.

  • Troubleshoot and improve system reliability, observability, and performance across multiple cloud providers (AWS, Azure, GCP).

What You Will Bring:

  • 2-5 years of industry experience designing, building, and supporting backend systems in production.

  • Strong fundamentals in distributed systems, cloud infrastructure, and networking.

  • Experience in building and operating large-scale, high-availability systems.

  • Good understanding of cloud platforms (AWS, Azure, or GCP) and their services.

  • Proficiency in Java, Scala, C++, Go, or other statically typed languages.

  • A self-starter with strong problem-solving skills and the ability to work in a fast-paced environment.

  • BS, MS, or PhD in computer science or a related field, or equivalent work experience.

What Gives You an Edge:

  • Exposure to model serving, LLM/agent infrastructure, or streaming data systems.

  • Note - You don't need a background in ML research or model training — this role is about building and operating the platform that serves AI reliably at scale, not inventing the models.

Ready to build what's next? Let’s get in motion.

Come As You Are

Belonging isn’t a perk here. It’s the baseline. We work across time zones and backgrounds, knowing the best ideas come from different perspectives. And we make space for everyone to lead, grow, and challenge what’s possible.

We’re proud to be an equal opportunity workplace. Employment decisions are based on job-related criteria, without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, veteran status, or any other classification protected by law.

Privacy Statement

Confluent is an IBM subsidiary which has been acquired by IBM and will be integrated into the IBM organization. By proceeding with this application, you understand that Confluent will share your personal information with other IBM affiliates involved in your recruitment process, wherever these are located. More Information on how IBM protects your personal information, including the safeguards in case of cross-border data transfer, are available here.

Redirects to Confluent's application page.

Other roles

More at Confluent.

View all 13 roles