Distributed Systems Software Engineer - WarpSteam
Senior Backend Engineer · Senior · Full Time · Remote
Opens Confluent's application page
Role
What you'll do.
This Distributed Systems Software Engineer role at Confluent focuses on building WarpStream, an Apache Kafka-compatible data streaming platform running on object storage with zero local disks. As a Senior Engineer on a small, fully remote team, you'll architect and develop mission-critical backend services solving complex distributed systems problems across AWS, GCP, and Azure, with full-stack ownership spanning storage engines, multi-tenant control planes, and developer experiences.
Responsibilities
- Drive Complex Technical Projects End-to-End: Independently execute and deliver large-scale distributed systems projects from conception through production deployment. Own all phases of development including architecture decisions, technical design, implementation, testing, and operational support while managing stakeholder communication and timeline delivery.
- Design and Build Mission-Critical Backend Services: Architect, design, develop, and operationalize high-performance, scalable, and resilient microservices that form the core of the WarpStream platform. Focus on building systems that handle multi-tenant workloads across global infrastructure while maintaining strict availability and performance SLA commitments to customers.
- Optimize Distributed Storage Architecture: Work on the storage engine and control plane components that enable WarpStream's unique architecture of running Kafka-compatible systems directly on object storage without local disks. Improve scalability, performance, and reliability of the multi-tenant control plane serving customers across multiple cloud regions.
- Implement Advanced Streaming Features: Lead development of sophisticated data streaming features including transactions, active-active multi-region clusters, and enhanced protocol implementations. Contribute to the continuous evolution of the Apache Kafka-compatible API surface while maintaining backward compatibility and operational reliability.
- Troubleshoot and Debug Complex Systems: Diagnose and resolve technical issues within the multi-layered technical stack including microservices, containers, virtualization, and networking components. Apply deep systems knowledge to identify root causes and implement preventative measures across the distributed platform.
- Ensure Operational Excellence and Reliability: Own the operational readiness of critical production services with responsibility for on-call support and incident response. Maintain SLA commitments through proactive monitoring, alerting, capacity planning, and continuous reliability improvements. Participate in on-call rotations handling critical system incidents across distributed infrastructure.
- Develop Full-Stack Solutions: Contribute across all layers of the technology stack from low-level network protocol optimization and storage engine development to building developer-facing console experiences with HTML and JavaScript. Demonstrate versatility in solving problems at every level of abstraction.
- Provide Technical Leadership and Mentorship: Guide junior team members through technical problem-solving, code reviews, and architectural decisions. Foster a culture of quality, continuous learning, and operational excellence within the small distributed team while maintaining customer-centric perspectives.
Qualifications
What we look for.
Technical
Proficiency in Major Programming Languages
Strong programming skills in one or more production languages such as Go, Java, C/C++, or Python. Demonstrated ability to write clean, efficient code with deep understanding of algorithmic complexity and performance optimization.
Distributed Systems and Storage Systems Expertise
Deep technical understanding of distributed systems principles including consensus algorithms, replication strategies, fault tolerance, and eventual consistency models. Knowledge of storage engine design, object storage architecture, and multi-region system design.
Backend Systems Design and Architecture
Experience designing and implementing scalable backend systems handling millions of requests. Proficiency in microservices architecture, database design, caching strategies, and system scalability patterns. Knowledge of cloud-native architectures and containerized deployments.
Production Systems Operations
Proven experience building, deploying, and maintaining large-scale production services. Expertise in observability, monitoring, logging, and incident response. Understanding of SLA management, capacity planning, and operational reliability best practices.
Multi-Cloud Platform Experience
Working knowledge of major cloud service providers (AWS, GCP, Azure) including compute services, storage solutions, networking capabilities, and cross-cloud deployment strategies. Experience building systems with cloud-provider abstraction.
Software Engineering Practices
Mastery of professional software engineering disciplines including code review processes, testing strategies (unit, integration, and end-to-end testing), documentation standards, and version control workflows. Strong commitment to code quality and maintainability.
Education
Bachelor's Degree in Computer Science or Related Field
Formal education in Computer Science, Computer Engineering, Mathematics, or equivalent practical experience demonstrating theoretical computer science foundation and problem-solving capabilities.
Experience
5+ Years Backend Systems Development
Minimum five years of professional experience designing, building, scaling, and supporting complex backend systems in production environments. Track record of shipping reliable, high-performance systems that serve millions of users or handle massive data volumes.
Large-Scale System Delivery
Demonstrated success delivering large-scale, highly available systems with proven operational excellence. Experience with systems handling significant traffic, data volume, and uptime requirements. Portfolio of projects showing progression in scope and complexity.
On-Call and Incident Response
Active experience being on-call for critical production systems with hands-on involvement in debugging complex issues under pressure. Comfortable with incident response workflows, blameless postmortems, and implementing improvements based on operational learnings.
Technical Leadership
Evidence of driving technical decisions, mentoring junior engineers, and influencing architectural choices. Experience leading projects with ambiguous requirements, working cross-functionally with product and infrastructure teams.
Skills
Required
Go Programming Language
Production-level proficiency in Go with understanding of goroutines, channels, and concurrency patterns essential for building efficient distributed systems and microservices.
Distributed Systems Design
Deep knowledge of distributed systems patterns including consistency models, consensus protocols (Raft, Paxos), replication strategies, and failure handling mechanisms.
Kafka and Event Streaming
Strong understanding of Apache Kafka architecture, stream processing concepts, consumer group management, and Kafka protocol internals. Experience with Kafka-compatible systems and streaming data platforms.
Cloud Infrastructure (AWS/GCP/Azure)
Hands-on experience provisioning, deploying, and managing applications on major cloud providers. Proficiency with infrastructure-as-code tools and multi-cloud deployment strategies.
Microservices Architecture
Experience designing, building, and operating microservices-based systems. Knowledge of service communication patterns, API design, and distributed tracing.
Database and Storage Systems
Working knowledge of relational and NoSQL databases, object storage, distributed file systems, and storage engine optimization. Understanding of indexing, query optimization, and storage performance tuning.
System Design and Scalability
Ability to design systems for scale including load balancing, sharding, caching strategies, and performance optimization. Experience with capacity planning and handling systems at petabyte or higher scales.
Observability and Debugging
Expertise in building observable systems with comprehensive logging, metrics, and distributed tracing. Proficiency with monitoring platforms and tools for diagnosing complex issues in production systems.
Preferred
Java or C++ Programming Experience
Nice to haveProficiency in compiled languages for high-performance systems development. Valuable for understanding low-level performance characteristics and system-level programming requirements.
Storage Engine Development
Nice to haveDirect experience building or modifying storage engines, query execution engines, or database internals. Knowledge of write-ahead logs, B-trees, LSM trees, and storage optimization techniques.
Kubernetes and Container Orchestration
Nice to haveHands-on experience with Kubernetes deployments, container networking, storage classes, and cloud-native operational patterns in production environments.
Object Storage Expertise
Nice to haveSpecific experience optimizing systems for object storage backends (S3, GCS, Azure Blob) and building applications that leverage cloud storage for cost efficiency and scalability.
Active-Active Replication and Multi-Region Systems
Nice to haveExperience implementing cross-region replication, handling network partitions, and designing systems for geographic distribution with eventual consistency.
Protocol Implementation
Nice to haveExperience implementing or maintaining wire protocols, API compatibility layers, or protocol specifications. Valuable for work on Kafka protocol compatibility and extensions.
Full-Stack Development
Nice to haveComfort working across the full technology stack including backend services, frontend development with HTML/JavaScript, and infrastructure tooling. Flexibility in tackling problems at any layer.
Developer Experience and API Design
Nice to haveExperience building developer-friendly tools, SDKs, or APIs with attention to usability, documentation, and developer feedback. Valuable for console development and platform experience enhancement.
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 197,400 – 271,200
Equity·Stock options
Full posting
Original listing.
We’re not just building better tech. We’re rewriting how data moves and what the world can do with it. With Confluent, data doesn’t sit still. Our platform puts information in motion, streaming in near real-time so companies can react faster, build smarter, and deliver experiences as dynamic as the world around them.
It takes a certain kind of person to join this team. Those who ask hard questions, give honest feedback, and show up for each other. No egos, no solo acts. Just smart, curious humans pushing toward something bigger, together.
One Confluent. One Team. One Data Streaming Platform.
About the Role:
WarpStream is an Apache Kafka compatible data streaming platform built directly on top of object storage with zero local disks. It’s delivered to customers as a BYOC-style offering. The WarpStream team is responsible for this product offering from top to bottom, and as a Senior Engineer you’ll do everything from working on the storage engine running in the Agents, improving the scalability of our multi-tenant control plane, implementing new features like transactions, active-active multi-regions clusters, and even write some HTML/Javascript for our developer console. We’re a small team of < 10 engineers, so you’ll be expected to work across the entire stack from bashing bits at the network level to making sure our UI provides a great developer experience.
You will have an opportunity to solve complex distributed systems problems at scale. You will build services that can operate across different cloud service providers like AWS, GCP, and Azure. You’ll be on-call for all of it.
The WarpStream team is fully remote and geographically distributed, currently we have engineers in: Spain, France, Canada, Ireland, and the USA.
What You Will Do:
Independently drive execution of complex technical projects end to end.
Build mission-critical backend services that deliver value to our customers. You will play a crucial role in architecting, designing, developing and operationalizing high performance, scalable, reliable and resilient services.
Troubleshoot and debug technical issues inside a deep and complex technical stack that includes microservices, containers, and virtualization.
Ensure operational readiness of the services and meet the availability and performance SLA commitments to our customers.
What You Will Bring:
5+ years industry experience designing, building, scaling and supporting backend systems in production with a solid grasp on good software engineering practices such as code reviews, deep focus on quality, and documentation
Strong programming and algorithmic skills. Proficiency in a major programming language, e.g. Java, Go, C / C++, Python, etc
Deep curiosity and enthusiasm for distributed systems and storage systems
Strong focus on project delivery and communication skills
Experience in driving operational excellence for large production services
A strong sense of customer centricity, teamwork, technical leadership and mentorship and are excited about team and company success
What Gives You an Edge:
Proven track record of delivering large-scale, highly available, high quality systems
Hands-on technical expertise in large scale systems engineering or distributed systems
On-call experience handling critical systems
Ready to build what's next? Let’s get in motion.
Come As You Are
Belonging isn’t a perk here. It’s the baseline. We work across time zones and backgrounds, knowing the best ideas come from different perspectives. And we make space for everyone to lead, grow, and challenge what’s possible.
We’re proud to be an equal opportunity workplace. Employment decisions are based on job-related criteria, without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, veteran status, or any other classification protected by law.
Privacy Statement
Confluent is an IBM subsidiary which has been acquired by IBM and will be integrated into the IBM organization. By proceeding with this application, you understand that Confluent will share your personal information with other IBM affiliates involved in your recruitment process, wherever these are located. More Information on how IBM protects your personal information, including the safeguards in case of cross-border data transfer, are available here.
Redirects to Confluent's application page.
Other roles
More at Confluent.
Senior Manager, Detection & Response (Security Engineering)
Manager
Staff Software Engineer
Staff
Senior Software Engineer - Streaming AI (Remote - Ontario / British Columbia)
Senior
Staff Software Engineer I
Staff
Senior Software Engineer
Senior