Senior Software Engineer - Streaming AI (Remote - Ontario / British Columbia)
Senior Software Engineer · Senior · Full Time · Remote
Opens Confluent's application page
Role
What you'll do.
As a Senior Software Engineer on Confluent's AI streaming team, you will architect and operate large-scale, high-performance infrastructure enabling AI inference directly on real-time data streams. This role focuses on building distributed systems that eliminate the need for external ML stacks, requiring expertise in cloud infrastructure, systems reliability, and scalable serving layers across AWS, Azure, and GCP. You'll tackle complex challenges in networking, compute optimization, and security while collaborating with core platform teams to deliver production-grade AI capabilities at scale.
Responsibilities
- Infrastructure Architecture and Design: Design, develop, and operate large-scale, high-performance infrastructure that powers Confluent Cloud's AI streaming capabilities. Take ownership of end-to-end product architecture from user-facing APIs down to production serving layers, ensuring systems are optimized for reliability, scalability, and cost efficiency in handling complex inference workloads.
- Distributed Systems Development: Build foundational software addressing distributed systems challenges including consensus algorithms, failover strategies, resource allocation, and state management. Work on systems that enable reliable AI agent execution on streaming data without requiring data movement to external platforms.
- Multi-Cloud Platform Optimization: Troubleshoot, improve, and optimize system reliability, observability, and performance across AWS, Azure, and GCP environments. Ensure consistent performance characteristics and cost efficiency across cloud providers while maintaining high availability standards.
- Cross-Team Collaboration and Integration: Collaborate with Confluent's broader platform teams whose foundational systems you build upon. Coordinate infrastructure enhancements for real-time data streaming use cases, participate in architectural reviews, and drive integration between AI serving infrastructure and core data plane systems.
- Production Operations and Reliability: Own the operational aspects of production AI serving infrastructure, establishing observability standards, defining reliability metrics, and implementing automated solutions for system monitoring. Drive continuous improvements in system resilience and incident response procedures for production environments.
Qualifications
What we look for.
Technical
Distributed Systems Fundamentals
Strong foundational knowledge of distributed systems principles including consistency models, fault tolerance, consensus algorithms, and distributed tracing. Understanding of networking protocols, load balancing, and inter-service communication patterns.
Statically Typed Programming Languages
Proficiency in Java, Scala, C++, Go, or equivalent statically typed languages used for building production infrastructure. Ability to write clean, maintainable code optimized for performance-critical systems.
Cloud Infrastructure and Kubernetes
Working knowledge of cloud infrastructure concepts including containerization, orchestration platforms, networking architectures, and infrastructure-as-code practices for managing complex deployments.
Systems Performance and Optimization
Understanding of performance optimization techniques including memory management, CPU efficiency, I/O optimization, and profiling tools. Ability to identify and resolve bottlenecks in high-throughput systems.
Education
Computer Science or Related Degree
BS, MS, or PhD in computer science, electrical engineering, mathematics, or related field preferred. Equivalent professional experience demonstrating mastery of core computer science concepts is acceptable.
Experience
Backend Systems Development
2-5 years of industry experience designing, building, and supporting backend systems in production environments. Demonstrated track record of shipping complex distributed systems that handle significant scale and traffic.
Large-Scale Infrastructure Operations
Proven experience building and operating large-scale, high-availability systems with expertise in maintaining system reliability, performance monitoring, and incident response in production environments.
Cloud Platform Experience
Hands-on experience with at least one major cloud provider (AWS, Azure, or GCP), including understanding of cloud services architecture, deployment patterns, and cost optimization strategies.
Skills
Required
Distributed Systems Design
Ability to architect reliable, scalable systems handling consensus, failover, state management, and data consistency across multiple nodes and geographical regions.
Backend Development in Statically Typed Languages
Production-level proficiency in Java, Scala, C++, Go, or similar languages with deep understanding of memory management, concurrency patterns, and performance characteristics.
Cloud Platform Architecture
Practical expertise with AWS, Azure, or GCP including services like compute instances, databases, networking, load balancing, and monitoring. Understanding of multi-cloud deployment strategies.
Systems Reliability and Observability
Experience building observable systems with comprehensive logging, metrics, distributed tracing, and alerting. Expertise in on-call practices, incident response, and postmortem-driven improvements.
High-Performance Systems Engineering
Proven ability to optimize systems for throughput, latency, and resource efficiency. Skilled in performance profiling, bottleneck identification, and implementing architectural improvements for scale.
Preferred
Model Serving Infrastructure
Nice to haveExperience building or operating platforms for serving machine learning models at scale, including understanding of inference optimization, batching, and resource management for ML workloads.
LLM and AI Agent Infrastructure
Nice to haveExposure to systems supporting large language models or autonomous agents, including challenges around prompt caching, context management, or multi-turn interactions at scale.
Streaming Data Systems
Nice to haveExperience with streaming platforms like Apache Kafka, AWS Kinesis, or similar systems. Understanding of stream processing patterns, state management, and exactly-once semantics in distributed contexts.
Apache Kafka Ecosystem
Nice to haveFamiliarity with Kafka architecture, including broker operations, topic management, client implementations, or contributing to Kafka-based systems. Understanding of Kafka internals and operational patterns.
Security in Distributed Systems
Nice to haveUnderstanding of security considerations in cloud infrastructure including authentication, authorization, encryption in transit and at rest, and compliance requirements for production environments.
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·CAD 144,200 – 169,400
Equity·Stock options
Benefits
Comprehensive Health Coverage
Medical, dental, and vision insurance plans covering employees and their families with options for various coverage levels and low out-of-pocket costs.
Retirement Planning
Competitive 401(k) matching program (US) or equivalent RRSP matching for Canadian employees, supporting long-term financial security and wealth building.
Flexible Work Arrangements
Remote-first role with flexibility to work from Ontario or British Columbia. Opportunity to maintain work-life balance while accessing Confluent's collaborative culture globally.
Professional Development
Learning stipends, internal training programs, technical certifications, and conference attendance support to advance expertise in distributed systems and cloud technologies.
Stock Options and Equity
Participation in company equity programs providing ownership stake and long-term wealth creation opportunities as Confluent grows.
Paid Time Off
Generous vacation policy, paid sick leave, and company holidays ensuring adequate time for rest, recovery, and personal pursuits throughout the year.
Life and Disability Insurance
Life insurance coverage and short/long-term disability protection providing financial security for you and your family in unexpected circumstances.
Wellness Programs
Mental health resources, fitness stipends, wellness initiatives, and employee assistance programs supporting holistic health and wellbeing.
Parental Leave
Competitive parental leave policies supporting work-life balance for growing families with paid time off and flexible return-to-work options.
Process
Interview steps.
- 01
Initial Screening Call
Conversation with recruiter to discuss career goals, experience with distributed systems and cloud infrastructure, and alignment with Confluent's AI streaming mission. This 30-minute call establishes fit and clarifies role expectations.
- 02
Technical Architecture Discussion
45-60 minute conversation with senior engineering team member exploring your experience designing distributed systems, handling scale challenges, and operating production infrastructure. Expect questions about past architectural decisions and trade-offs.
- 03
Coding and Systems Design
Technical assessment involving practical problem-solving around distributed systems design, potentially including a take-home assignment or live coding session focused on infrastructure patterns rather than leetcode-style problems.
- 04
System Design Deep Dive
In-depth 60-90 minute technical interview with 2-3 engineers discussing large-scale system design, cloud infrastructure challenges, and how you would architect solutions for the specific problems the AI team faces at Confluent.
- 05
Cross-Functional Collaboration
Conversation with engineers from platform teams to assess collaboration style, communication skills, and ability to work effectively across team boundaries. Discussion of cross-team challenges and integration patterns.
- 06
Leadership and Values Alignment
Final round with hiring manager or director covering leadership approach, growth mindset, how you navigate ambiguity, and alignment with Confluent's culture of ownership, transparency, and collaborative problem-solving.
Full posting
Original listing.
We’re not just building better tech. We’re rewriting how data moves and what the world can do with it. With Confluent, data doesn’t sit still. Our platform puts information in motion, streaming in near real-time so companies can react faster, build smarter, and deliver experiences as dynamic as the world around them.
It takes a certain kind of person to join this team. Those who ask hard questions, give honest feedback, and show up for each other. No egos, no solo acts. Just smart, curious humans pushing toward something bigger, together.
One Confluent. One Team. One Data Streaming Platform.
About the Team:
We're a small, focused engineering team building the AI capabilities of Confluent Cloud. Our job is to make it possible to run machine learning and AI agents directly on real-time data — without customers having to stitch together a separate stack to do it. We own our products end to end, from the user-facing API down to the serving layer that runs inference in production, and we work closely with the broader platform teams whose systems we build on top of. It's a high-ownership, high-autonomy environment: small enough that what you build ships and matters, broad enough that the problems are genuinely hard.
About the Role:
As a Software Engineer on the AI team, you will take ownership of the infrastructure that enables our "in-place, at-scale" AI value proposition. You aren't just building a feature; you are architecting the systems that allow customers to run complex inference and AI agents directly on streaming data, eliminating the need to move data to external stacks. This is a high-impact role where your work on scalable, cost-efficient serving layers directly drives Confluent's growth and redefines the possibilities of real-time data.
We are looking for engineers who thrive on the technical complexity of large-scale distributed systems. You will tackle deep infrastructure challenges across networking, compute, and security to ensure our AI capabilities are as reliable as the core data plane itself. If you are passionate about building the foundational systems that power the next generation of AI in the cloud, this is the place to do it.
What You Will Do:
Design, develop, and operate large-scale, high-performance infrastructure that powers Confluent Cloud.
Build foundational software to improve reliability, scalability, and efficiency across cloud environments.
Work on distributed systems challenges such as consensus algorithms, failover strategies, and resource allocation.
Collaborate with teams across Confluent to optimize and enhance infrastructure for real-time data streaming use cases.
Troubleshoot and improve system reliability, observability, and performance across multiple cloud providers (AWS, Azure, GCP).
What You Will Bring:
2-5 years of industry experience designing, building, and supporting backend systems in production.
Strong fundamentals in distributed systems, cloud infrastructure, and networking.
Experience in building and operating large-scale, high-availability systems.
Good understanding of cloud platforms (AWS, Azure, or GCP) and their services.
Proficiency in Java, Scala, C++, Go, or other statically typed languages.
A self-starter with strong problem-solving skills and the ability to work in a fast-paced environment.
BS, MS, or PhD in computer science or a related field, or equivalent work experience.
What Gives You an Edge:
Exposure to model serving, LLM/agent infrastructure, or streaming data systems.
Note - You don't need a background in ML research or model training — this role is about building and operating the platform that serves AI reliably at scale, not inventing the models.
Ready to build what's next? Let’s get in motion.
Come As You Are
Belonging isn’t a perk here. It’s the baseline. We work across time zones and backgrounds, knowing the best ideas come from different perspectives. And we make space for everyone to lead, grow, and challenge what’s possible.
We’re proud to be an equal opportunity workplace. Employment decisions are based on job-related criteria, without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, veteran status, or any other classification protected by law.
Privacy Statement
Confluent is an IBM subsidiary which has been acquired by IBM and will be integrated into the IBM organization. By proceeding with this application, you understand that Confluent will share your personal information with other IBM affiliates involved in your recruitment process, wherever these are located. More Information on how IBM protects your personal information, including the safeguards in case of cross-border data transfer, are available here.
Redirects to Confluent's application page.
Other roles
More at Confluent.
Senior Manager, Detection & Response (Security Engineering)
Manager
Staff Software Engineer
Staff
Distributed Systems Software Engineer - WarpSteam
Senior
Staff Software Engineer I
Staff
Senior Software Engineer
Senior