Confluent

Distributed Systems Software Engineer - WarpSteam

Confluent4 days ago
Location

Remote, United States

Workplace

Remote

Type

Full Time

Salary

USD 197,400 – 271,200

Level

Senior

Role

Senior Backend Engineer

Posted

Jul 21, 2026

Full TimeRemoteSenior

The role

Summary

This Distributed Systems Software Engineer role at Confluent focuses on building WarpStream, an Apache Kafka-compatible data streaming platform running on object storage with zero local disks. As a Senior Engineer on a small, fully remote team, you'll architect and develop mission-critical backend services solving complex distributed systems problems across AWS, GCP, and Azure, with full-stack ownership spanning storage engines, multi-tenant control planes, and developer experiences.

What you'll do

Drive Complex Technical Projects End-to-End: Independently execute and deliver large-scale distributed systems projects from conception through production deployment. Own all phases of development including architecture decisions, technical design, implementation, testing, and operational support while managing stakeholder communication and timeline delivery.
Design and Build Mission-Critical Backend Services: Architect, design, develop, and operationalize high-performance, scalable, and resilient microservices that form the core of the WarpStream platform. Focus on building systems that handle multi-tenant workloads across global infrastructure while maintaining strict availability and performance SLA commitments to customers.
Optimize Distributed Storage Architecture: Work on the storage engine and control plane components that enable WarpStream's unique architecture of running Kafka-compatible systems directly on object storage without local disks. Improve scalability, performance, and reliability of the multi-tenant control plane serving customers across multiple cloud regions.
Implement Advanced Streaming Features: Lead development of sophisticated data streaming features including transactions, active-active multi-region clusters, and enhanced protocol implementations. Contribute to the continuous evolution of the Apache Kafka-compatible API surface while maintaining backward compatibility and operational reliability.
Troubleshoot and Debug Complex Systems: Diagnose and resolve technical issues within the multi-layered technical stack including microservices, containers, virtualization, and networking components. Apply deep systems knowledge to identify root causes and implement preventative measures across the distributed platform.
Ensure Operational Excellence and Reliability: Own the operational readiness of critical production services with responsibility for on-call support and incident response. Maintain SLA commitments through proactive monitoring, alerting, capacity planning, and continuous reliability improvements. Participate in on-call rotations handling critical system incidents across distributed infrastructure.
Develop Full-Stack Solutions: Contribute across all layers of the technology stack from low-level network protocol optimization and storage engine development to building developer-facing console experiences with HTML and JavaScript. Demonstrate versatility in solving problems at every level of abstraction.
Provide Technical Leadership and Mentorship: Guide junior team members through technical problem-solving, code reviews, and architectural decisions. Foster a culture of quality, continuous learning, and operational excellence within the small distributed team while maintaining customer-centric perspectives.

What we look for

Technical

Proficiency in Major Programming LanguagesStrong programming skills in one or more production languages such as Go, Java, C/C++, or Python. Demonstrated ability to write clean, efficient code with deep understanding of algorithmic complexity and performance optimization.
Distributed Systems and Storage Systems ExpertiseDeep technical understanding of distributed systems principles including consensus algorithms, replication strategies, fault tolerance, and eventual consistency models. Knowledge of storage engine design, object storage architecture, and multi-region system design.
Backend Systems Design and ArchitectureExperience designing and implementing scalable backend systems handling millions of requests. Proficiency in microservices architecture, database design, caching strategies, and system scalability patterns. Knowledge of cloud-native architectures and containerized deployments.
Production Systems OperationsProven experience building, deploying, and maintaining large-scale production services. Expertise in observability, monitoring, logging, and incident response. Understanding of SLA management, capacity planning, and operational reliability best practices.
Multi-Cloud Platform ExperienceWorking knowledge of major cloud service providers (AWS, GCP, Azure) including compute services, storage solutions, networking capabilities, and cross-cloud deployment strategies. Experience building systems with cloud-provider abstraction.
Software Engineering PracticesMastery of professional software engineering disciplines including code review processes, testing strategies (unit, integration, and end-to-end testing), documentation standards, and version control workflows. Strong commitment to code quality and maintainability.

Education

Bachelor's Degree in Computer Science or Related FieldFormal education in Computer Science, Computer Engineering, Mathematics, or equivalent practical experience demonstrating theoretical computer science foundation and problem-solving capabilities.

Experience

5+ Years Backend Systems DevelopmentMinimum five years of professional experience designing, building, scaling, and supporting complex backend systems in production environments. Track record of shipping reliable, high-performance systems that serve millions of users or handle massive data volumes.
Large-Scale System DeliveryDemonstrated success delivering large-scale, highly available systems with proven operational excellence. Experience with systems handling significant traffic, data volume, and uptime requirements. Portfolio of projects showing progression in scope and complexity.
On-Call and Incident ResponseActive experience being on-call for critical production systems with hands-on involvement in debugging complex issues under pressure. Comfortable with incident response workflows, blameless postmortems, and implementing improvements based on operational learnings.
Technical LeadershipEvidence of driving technical decisions, mentoring junior engineers, and influencing architectural choices. Experience leading projects with ambiguous requirements, working cross-functionally with product and infrastructure teams.

Skills

Required skills

Go Programming LanguageProduction-level proficiency in Go with understanding of goroutines, channels, and concurrency patterns essential for building efficient distributed systems and microservices.
Distributed Systems DesignDeep knowledge of distributed systems patterns including consistency models, consensus protocols (Raft, Paxos), replication strategies, and failure handling mechanisms.
Kafka and Event StreamingStrong understanding of Apache Kafka architecture, stream processing concepts, consumer group management, and Kafka protocol internals. Experience with Kafka-compatible systems and streaming data platforms.
Cloud Infrastructure (AWS/GCP/Azure)Hands-on experience provisioning, deploying, and managing applications on major cloud providers. Proficiency with infrastructure-as-code tools and multi-cloud deployment strategies.
Microservices ArchitectureExperience designing, building, and operating microservices-based systems. Knowledge of service communication patterns, API design, and distributed tracing.
Database and Storage SystemsWorking knowledge of relational and NoSQL databases, object storage, distributed file systems, and storage engine optimization. Understanding of indexing, query optimization, and storage performance tuning.
System Design and ScalabilityAbility to design systems for scale including load balancing, sharding, caching strategies, and performance optimization. Experience with capacity planning and handling systems at petabyte or higher scales.
Observability and DebuggingExpertise in building observable systems with comprehensive logging, metrics, and distributed tracing. Proficiency with monitoring platforms and tools for diagnosing complex issues in production systems.

Nice to have

Java or C++ Programming ExperienceProficiency in compiled languages for high-performance systems development. Valuable for understanding low-level performance characteristics and system-level programming requirements.
Storage Engine DevelopmentDirect experience building or modifying storage engines, query execution engines, or database internals. Knowledge of write-ahead logs, B-trees, LSM trees, and storage optimization techniques.
Kubernetes and Container OrchestrationHands-on experience with Kubernetes deployments, container networking, storage classes, and cloud-native operational patterns in production environments.
Object Storage ExpertiseSpecific experience optimizing systems for object storage backends (S3, GCS, Azure Blob) and building applications that leverage cloud storage for cost efficiency and scalability.
Active-Active Replication and Multi-Region SystemsExperience implementing cross-region replication, handling network partitions, and designing systems for geographic distribution with eventual consistency.
Protocol ImplementationExperience implementing or maintaining wire protocols, API compatibility layers, or protocol specifications. Valuable for work on Kafka protocol compatibility and extensions.
Full-Stack DevelopmentComfort working across the full technology stack including backend services, frontend development with HTML/JavaScript, and infrastructure tooling. Flexibility in tackling problems at any layer.
Developer Experience and API DesignExperience building developer-friendly tools, SDKs, or APIs with attention to usability, documentation, and developer feedback. Valuable for console development and platform experience enhancement.

Compensation & benefits

Salary

USD 197,400 – 271,200 (annual)

Stock options

Available


Apply for this position

You'll be redirected to the company's application page


Confluent

Confluent

View all jobs

Confluent is an American data streaming platform company based on Apache Kafka.

Mountain View, California, United StatesFounded 2014confluent.io

Tech Stack

Languages
GoJavaC/C++PythonJavaScript
Frameworks
Apache KafkagRPCProtocol Buffers
Databases
Object Storage (S3/GCS/Azure Blob)Distributed Databases
Tools
AWSGoogle Cloud Platform (GCP)Microsoft AzureKubernetesDockerGit and Version ControlTerraformPrometheus/GrafanaELK Stack or Similar Logging
Other
Distributed Consensus AlgorithmsWrite-Ahead Logging and Storage OptimizationMulti-Tenancy ArchitectureEvent-Driven ArchitectureNetwork Protocol DesignDisaster Recovery and High Availability

Interview Guides

14 guides available for Confluent

Apply Now