Distributed Systems Software Engineer - WarpSteam

Senior Backend Engineer · Senior · Full Time · Remote

Remote, United States · RemoteUSD 197k – 271k1mo ago
Apply for this role

Opens Confluent's application page

Role

What you'll do.

This Distributed Systems Software Engineer role at Confluent focuses on building WarpStream, an Apache Kafka-compatible data streaming platform running on object storage with zero local disks. As a Senior Engineer on a small, fully remote team, you'll architect and develop mission-critical backend services solving complex distributed systems problems across AWS, GCP, and Azure, with full-stack ownership spanning storage engines, multi-tenant control planes, and developer experiences.

Responsibilities

  • Drive Complex Technical Projects End-to-End: Independently execute and deliver large-scale distributed systems projects from conception through production deployment. Own all phases of development including architecture decisions, technical design, implementation, testing, and operational support while managing stakeholder communication and timeline delivery.
  • Design and Build Mission-Critical Backend Services: Architect, design, develop, and operationalize high-performance, scalable, and resilient microservices that form the core of the WarpStream platform. Focus on building systems that handle multi-tenant workloads across global infrastructure while maintaining strict availability and performance SLA commitments to customers.
  • Optimize Distributed Storage Architecture: Work on the storage engine and control plane components that enable WarpStream's unique architecture of running Kafka-compatible systems directly on object storage without local disks. Improve scalability, performance, and reliability of the multi-tenant control plane serving customers across multiple cloud regions.
  • Implement Advanced Streaming Features: Lead development of sophisticated data streaming features including transactions, active-active multi-region clusters, and enhanced protocol implementations. Contribute to the continuous evolution of the Apache Kafka-compatible API surface while maintaining backward compatibility and operational reliability.
  • Troubleshoot and Debug Complex Systems: Diagnose and resolve technical issues within the multi-layered technical stack including microservices, containers, virtualization, and networking components. Apply deep systems knowledge to identify root causes and implement preventative measures across the distributed platform.
  • Ensure Operational Excellence and Reliability: Own the operational readiness of critical production services with responsibility for on-call support and incident response. Maintain SLA commitments through proactive monitoring, alerting, capacity planning, and continuous reliability improvements. Participate in on-call rotations handling critical system incidents across distributed infrastructure.
  • Develop Full-Stack Solutions: Contribute across all layers of the technology stack from low-level network protocol optimization and storage engine development to building developer-facing console experiences with HTML and JavaScript. Demonstrate versatility in solving problems at every level of abstraction.
  • Provide Technical Leadership and Mentorship: Guide junior team members through technical problem-solving, code reviews, and architectural decisions. Foster a culture of quality, continuous learning, and operational excellence within the small distributed team while maintaining customer-centric perspectives.

Qualifications

What we look for.

Technical

  • Proficiency in Major Programming Languages

    Strong programming skills in one or more production languages such as Go, Java, C/C++, or Python. Demonstrated ability to write clean, efficient code with deep understanding of algorithmic complexity and performance optimization.

  • Distributed Systems and Storage Systems Expertise

    Deep technical understanding of distributed systems principles including consensus algorithms, replication strategies, fault tolerance, and eventual consistency models. Knowledge of storage engine design, object storage architecture, and multi-region system design.

  • Backend Systems Design and Architecture

    Experience designing and implementing scalable backend systems handling millions of requests. Proficiency in microservices architecture, database design, caching strategies, and system scalability patterns. Knowledge of cloud-native architectures and containerized deployments.

  • Production Systems Operations

    Proven experience building, deploying, and maintaining large-scale production services. Expertise in observability, monitoring, logging, and incident response. Understanding of SLA management, capacity planning, and operational reliability best practices.

  • Multi-Cloud Platform Experience

    Working knowledge of major cloud service providers (AWS, GCP, Azure) including compute services, storage solutions, networking capabilities, and cross-cloud deployment strategies. Experience building systems with cloud-provider abstraction.

  • Software Engineering Practices

    Mastery of professional software engineering disciplines including code review processes, testing strategies (unit, integration, and end-to-end testing), documentation standards, and version control workflows. Strong commitment to code quality and maintainability.

Education

  • Bachelor's Degree in Computer Science or Related Field

    Formal education in Computer Science, Computer Engineering, Mathematics, or equivalent practical experience demonstrating theoretical computer science foundation and problem-solving capabilities.

Experience

  • 5+ Years Backend Systems Development

    Minimum five years of professional experience designing, building, scaling, and supporting complex backend systems in production environments. Track record of shipping reliable, high-performance systems that serve millions of users or handle massive data volumes.

  • Large-Scale System Delivery

    Demonstrated success delivering large-scale, highly available systems with proven operational excellence. Experience with systems handling significant traffic, data volume, and uptime requirements. Portfolio of projects showing progression in scope and complexity.

  • On-Call and Incident Response

    Active experience being on-call for critical production systems with hands-on involvement in debugging complex issues under pressure. Comfortable with incident response workflows, blameless postmortems, and implementing improvements based on operational learnings.

  • Technical Leadership

    Evidence of driving technical decisions, mentoring junior engineers, and influencing architectural choices. Experience leading projects with ambiguous requirements, working cross-functionally with product and infrastructure teams.

Skills

Required

  • Go Programming Language

    Production-level proficiency in Go with understanding of goroutines, channels, and concurrency patterns essential for building efficient distributed systems and microservices.

  • Distributed Systems Design

    Deep knowledge of distributed systems patterns including consistency models, consensus protocols (Raft, Paxos), replication strategies, and failure handling mechanisms.

  • Kafka and Event Streaming

    Strong understanding of Apache Kafka architecture, stream processing concepts, consumer group management, and Kafka protocol internals. Experience with Kafka-compatible systems and streaming data platforms.

  • Cloud Infrastructure (AWS/GCP/Azure)

    Hands-on experience provisioning, deploying, and managing applications on major cloud providers. Proficiency with infrastructure-as-code tools and multi-cloud deployment strategies.

  • Microservices Architecture

    Experience designing, building, and operating microservices-based systems. Knowledge of service communication patterns, API design, and distributed tracing.

  • Database and Storage Systems

    Working knowledge of relational and NoSQL databases, object storage, distributed file systems, and storage engine optimization. Understanding of indexing, query optimization, and storage performance tuning.

  • System Design and Scalability

    Ability to design systems for scale including load balancing, sharding, caching strategies, and performance optimization. Experience with capacity planning and handling systems at petabyte or higher scales.

  • Observability and Debugging

    Expertise in building observable systems with comprehensive logging, metrics, and distributed tracing. Proficiency with monitoring platforms and tools for diagnosing complex issues in production systems.

Preferred

  • Java or C++ Programming Experience

    Nice to have

    Proficiency in compiled languages for high-performance systems development. Valuable for understanding low-level performance characteristics and system-level programming requirements.

  • Storage Engine Development

    Nice to have

    Direct experience building or modifying storage engines, query execution engines, or database internals. Knowledge of write-ahead logs, B-trees, LSM trees, and storage optimization techniques.

  • Kubernetes and Container Orchestration

    Nice to have

    Hands-on experience with Kubernetes deployments, container networking, storage classes, and cloud-native operational patterns in production environments.

  • Object Storage Expertise

    Nice to have

    Specific experience optimizing systems for object storage backends (S3, GCS, Azure Blob) and building applications that leverage cloud storage for cost efficiency and scalability.

  • Active-Active Replication and Multi-Region Systems

    Nice to have

    Experience implementing cross-region replication, handling network partitions, and designing systems for geographic distribution with eventual consistency.

  • Protocol Implementation

    Nice to have

    Experience implementing or maintaining wire protocols, API compatibility layers, or protocol specifications. Valuable for work on Kafka protocol compatibility and extensions.

  • Full-Stack Development

    Nice to have

    Comfort working across the full technology stack including backend services, frontend development with HTML/JavaScript, and infrastructure tooling. Flexibility in tackling problems at any layer.

  • Developer Experience and API Design

    Nice to have

    Experience building developer-friendly tools, SDKs, or APIs with attention to usability, documentation, and developer feedback. Valuable for console development and platform experience enhancement.

Tech stack

Languages

GoJavaC/C++PythonJavaScript

Frameworks

Apache KafkagRPCProtocol Buffers

Databases

Object Storage (S3/GCS/Azure Blob)Distributed Databases

Tools

AWSGoogle Cloud Platform (GCP)Microsoft AzureKubernetesDockerGit and Version ControlTerraformPrometheus/GrafanaELK Stack or Similar Logging

Other

Distributed Consensus AlgorithmsWrite-Ahead Logging and Storage OptimizationMulti-Tenancy ArchitectureEvent-Driven ArchitectureNetwork Protocol DesignDisaster Recovery and High Availability

Compensation

Pay and benefits.

Base·USD 197,400 – 271,200

Equity·Stock options

Full posting

Original listing.

We’re not just building better tech. We’re rewriting how data moves and what the world can do with it. With Confluent, data doesn’t sit still. Our platform puts information in motion, streaming in near real-time so companies can react faster, build smarter, and deliver experiences as dynamic as the world around them.

It takes a certain kind of person to join this team. Those who ask hard questions, give honest feedback, and show up for each other. No egos, no solo acts. Just smart, curious humans pushing toward something bigger, together.

One Confluent. One Team. One Data Streaming Platform.

About the Role:

WarpStream is an Apache Kafka compatible data streaming platform built directly on top of object storage with zero local disks. It’s delivered to customers as a BYOC-style offering. The WarpStream team is responsible for this product offering from top to bottom, and as a Senior Engineer you’ll do everything from working on the storage engine running in the Agents, improving the scalability of our multi-tenant control plane, implementing new features like transactions, active-active multi-regions clusters, and even write some HTML/Javascript for our developer console. We’re a small team of < 10 engineers, so you’ll be expected to work across the entire stack from bashing bits at the network level to making sure our UI provides a great developer experience.

You will have an opportunity to solve complex distributed systems problems at scale. You will build services that can operate across different cloud service providers like AWS, GCP, and Azure. You’ll be on-call for all of it.

The WarpStream team is fully remote and geographically distributed, currently we have engineers in: Spain, France, Canada, Ireland, and the USA.

What You Will Do:

  • Independently drive execution of complex technical projects end to end.

  • Build mission-critical backend services that deliver value to our customers. You will play a crucial role in architecting, designing, developing and operationalizing high performance, scalable, reliable and resilient services.

  • Troubleshoot and debug technical issues inside a deep and complex technical stack that includes microservices, containers, and virtualization.

  • Ensure operational readiness of the services and meet the availability and performance SLA commitments to our customers.

What You Will Bring:

  • 5+ years industry experience designing, building, scaling and supporting backend systems in production with a solid grasp on good software engineering practices such as code reviews, deep focus on quality, and documentation

  • Strong programming and algorithmic skills. Proficiency in a major programming language, e.g. Java, Go, C / C++, Python, etc

  • Deep curiosity and enthusiasm for distributed systems and storage systems

  • Strong focus on project delivery and communication skills

  • Experience in driving operational excellence for large production services

  • A strong sense of customer centricity, teamwork, technical leadership and mentorship and are excited about team and company success

What Gives You an Edge:

  • Proven track record of delivering large-scale, highly available, high quality systems

  • Hands-on technical expertise in large scale systems engineering or distributed systems

  • On-call experience handling critical systems

Ready to build what's next? Let’s get in motion.

Come As You Are

Belonging isn’t a perk here. It’s the baseline. We work across time zones and backgrounds, knowing the best ideas come from different perspectives. And we make space for everyone to lead, grow, and challenge what’s possible.

We’re proud to be an equal opportunity workplace. Employment decisions are based on job-related criteria, without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, veteran status, or any other classification protected by law.

Privacy Statement

Confluent is an IBM subsidiary which has been acquired by IBM and will be integrated into the IBM organization. By proceeding with this application, you understand that Confluent will share your personal information with other IBM affiliates involved in your recruitment process, wherever these are located. More Information on how IBM protects your personal information, including the safeguards in case of cross-border data transfer, are available here.

Redirects to Confluent's application page.

Other roles

More at Confluent.

View all 12 roles