Staff Software Engineer
Staff · Full Time · Remote
Opens Confluent's application page
Role
What you'll do.
Staff Software Engineer at Confluent will design and build backend services (Go, Java, Python) for real-time AI inference on streaming data within Confluent Cloud. This role requires 10+ years of distributed systems experience and demands end-to-end ownership of complex, cross-team technical initiatives spanning model lifecycle management, inference routing, and agent execution on production infrastructure.
Responsibilities
- Design and Build AI Inference Backend Services: Design and develop scalable backend services primarily in Go, Java, and Python that execute AI model inference and agent execution on real-time streaming data within the Confluent Cloud platform, ensuring low-latency performance and high throughput.
- Own End-to-End Feature Delivery: Take full ownership of significant product features from conception through production, including drafting comprehensive technical designs, aligning stakeholders across teams, driving design decisions to completion, and managing the complete delivery lifecycle.
- Make Cross-System Technical Decisions: Lead technical architecture decisions across multiple interconnected systems including model lifecycle management, inference request routing, agent orchestration, and inference serving layers, ensuring coherent system design and operational excellence.
- Ensure Production Quality and Reliability: Maintain high standards for code quality, comprehensive test coverage, clear documentation, operational observability, and safe deployment practices for production inference infrastructure serving live traffic, with emphasis on reliability and incident prevention.
- Mentor and Lead Technical Excellence: Elevate team capability through thorough code reviews, constructive design feedback, and personal accountability for complex cross-cutting technical work, building trust and establishing yourself as a go-to technical leader for ambiguous problems.
- Participate in On-Call Rotation and Operations: Maintain operational responsibility for services owned by the team through on-call participation, incident response, and proactive process improvements to ensure sustainable team operations and continuous service health.
Qualifications
What we look for.
Technical
Distributed Systems Architecture
Deep expertise designing, building, and operating distributed systems and cloud-native backend infrastructure in production environments at scale, including experience with system design tradeoffs, failure modes, and resilience patterns.
Kubernetes and Container Orchestration
Strong working knowledge of Kubernetes architecture, deployment patterns, and operational best practices, combined with deep understanding of containerization technologies and container networking fundamentals.
Distributed Systems Patterns
Proficiency with distributed systems design patterns including control loops, API server architecture, high-scale control plane design, eventual consistency models, and consensus algorithms in production systems.
Multi-Language Backend Development
Proficiency in at least one of Go, Java, or Python with demonstrated ability to work effectively across all three languages, understanding language-specific performance characteristics and ecosystem tooling.
Backend Infrastructure and DevOps
Expertise in building and operating production backend infrastructure, including logging, monitoring, alerting, tracing, infrastructure-as-code, CI/CD pipeline design, and deployment automation.
Education
Bachelor's Degree in Computer Science or Related Field
Bachelor's degree in Computer Science, Software Engineering, or related technical discipline, or equivalent professional experience demonstrating deep computer science fundamentals.
Experience
Senior Distributed Systems Engineering
10+ years of professional software engineering experience with at least 5-7 years focused on designing, implementing, and operating large-scale distributed systems in production environments.
Cross-Team Technical Leadership
Demonstrated track record of leading ambiguous, cross-functional technical initiatives that span multiple teams and systems, translating unclear requirements into actionable technical designs that gain stakeholder alignment.
Production Infrastructure Operations
Substantial experience owning production systems end-to-end, including operational responsibility through on-call rotations, incident response, postmortem analysis, and driving operational improvements.
System Design and Architecture
Experience making high-impact architectural decisions for complex backend systems, considering scalability, reliability, cost, and maintainability across multiple services and deployment environments.
Skills
Required
Go Programming
Proficiency in Go for building high-performance, concurrent systems with strong understanding of goroutines, channels, and Go's concurrency model.
Java Development
Strong Java expertise including modern frameworks, JVM performance tuning, garbage collection, memory management, and building scalable backend services.
Python Backend Development
Solid Python skills for backend service development, including async programming patterns, performance optimization, and integration with distributed systems.
Kubernetes
Deep knowledge of Kubernetes architecture, API objects, scheduling, resource management, networking, storage, and operational patterns for production deployments.
Distributed Systems Fundamentals
Core knowledge of CAP theorem, consistency models, fault tolerance, replication, data synchronization, and consensus mechanisms in distributed systems.
API Design and RESTful Services
Expertise designing clean, maintainable APIs for backend services, including request/response serialization, error handling, versioning strategies, and SDK development.
System Design and Architecture
Ability to design large-scale systems considering scalability, reliability, latency, throughput, and cost tradeoffs, and to articulate design decisions clearly to technical and non-technical stakeholders.
Technical Communication and Documentation
Excellent written and verbal communication skills including ability to write clear design documents, architecture decision records (ADRs), and technical specifications that align cross-functional teams.
Preferred
Model Serving Infrastructure
Nice to haveExperience building or operating platforms for serving machine learning models at scale, including batching strategies, inference optimization, and model lifecycle management.
LLM and AI Agent Infrastructure
Nice to haveExposure to large language model serving platforms, prompt management systems, or agent orchestration frameworks, understanding inference patterns and operational challenges.
Streaming Data Systems
Nice to haveExperience with streaming data platforms, event processing systems, or Kafka-based architectures, understanding real-time data pipelines and event-driven application patterns.
Apache Kafka
Nice to haveFamiliarity with Apache Kafka architecture, broker design, consumer group management, and distributed topic partitioning for high-throughput event streaming.
gRPC and Protocol Buffers
Nice to haveExperience implementing high-performance RPC systems using gRPC and Protocol Buffers for efficient service-to-service communication in distributed systems.
Cloud Platform Services
Nice to haveExperience building on major cloud platforms (AWS, GCP, Azure) with understanding of managed services, auto-scaling, cost optimization, and cloud-native architecture patterns.
Observability and Monitoring
Nice to haveStrong background in designing observable systems with distributed tracing, structured logging, metrics collection, and building dashboards for production system visibility.
Control Plane Design
Nice to haveExperience designing high-scale control planes that manage cluster state, coordinate distributed work, handle failures gracefully, and maintain consistency at scale.
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 235,700 – 277,000
Equity·Stock options
Benefits
Comprehensive Health Coverage
Medical, dental, and vision insurance options with competitive premiums, covering preventive care, specialist visits, and mental health services.
Retirement Planning
401(k) plan with employer match, helping you build long-term financial security with tax-advantaged savings options.
Paid Time Off
Generous vacation days, sick leave, and paid holidays enabling work-life balance and personal wellness.
Professional Development
Learning budget, conference attendance support, and access to online training platforms for continuous skill development and career growth.
Stock Options and Equity
Opportunity to participate in company equity through stock options or RSUs, aligning your success with Confluent's growth and providing wealth-building potential.
Flexible Work Arrangements
Remote-first culture supporting distributed teams across time zones with flexibility to work from home or office as needed.
Parental Leave
Generous parental leave policies supporting new parents during critical early months of childcare.
Wellness Programs
Fitness subsidies, mental health resources, wellness workshops, and employee assistance programs supporting holistic health.
Process
Interview steps.
- 01
Initial Phone Screen
Recruiter conducts 30-minute conversation to assess background, career trajectory, motivation, and alignment with Staff-level expectations. Expect discussion of past distributed systems work and cross-team technical leadership examples.
- 02
Technical Architecture Discussion
Engineer-led 1-hour conversation focused on system design thinking. You'll discuss approaches to large-scale problems similar to those in the role, such as designing inference routing systems or model lifecycle management. Bring examples of systems you've designed.
- 03
Distributed Systems Deep Dive
Senior engineer conducts 1.5-hour technical discussion covering Kubernetes internals, distributed consensus, fault tolerance, and operational patterns. Prepare to discuss production incidents, failure scenarios, and how you've debugged complex distributed system problems.
- 04
Cross-Functional Collaboration Assessment
Meeting with engineers from different teams to evaluate communication skills, ability to make sound technical decisions across boundaries, and how you approach building alignment on complex initiatives.
- 05
Leadership and Vision Discussion
Conversation with engineering manager or senior tech lead about your technical vision, how you mentor other engineers, your approach to end-to-end ownership, and your perspective on building systems at scale.
- 06
Final Executive Round
Brief conversation with director or VP-level executive covering career aspirations, culture fit, and strategic thinking about building AI infrastructure on streaming platforms.
Full posting
Original listing.
We’re not just building better tech. We’re rewriting how data moves and what the world can do with it. With Confluent, data doesn’t sit still. Our platform puts information in motion, streaming in near real-time so companies can react faster, build smarter, and deliver experiences as dynamic as the world around them.
It takes a certain kind of person to join this team. Those who ask hard questions, give honest feedback, and show up for each other. No egos, no solo acts. Just smart, curious humans pushing toward something bigger, together.
One Confluent. One Team. One Data Streaming Platform.
About the Role:
You'll help build Confluent Cloud's AI capabilities — the layer that lets customers bring AI and AI agents capabilities directly to their real-time data. Instead of moving data out to a separate system to run inference or build an agent, our customers do it in place, on streaming data, as part of the same platform they already use to move and process events at scale.
As an engineer, you'll own delivery of significant pieces of this product — not just writing code, but deciding how a capability should work across the services that make it up. The interesting problems here rarely live in one place: shipping something like inference-on-streaming-data or an AI agent that reacts to live events touches several systems at once — the user-facing API, the services that manage model and agent lifecycle, the control plane that schedules and runs the work, and the serving layer that actually executes inference. You'll be expected to reason across those boundaries, make sound design calls, and get engineers inside and outside the team aligned on the approach.
What You Will Do:
Design and build the backend services (primarily Go, Java, and Python) that run AI and model inference on real-time data.
Own features end to end — drafting the design, aligning stakeholders inside and outside the team, and driving the decision to a conclusion.
Make the technical calls on systems that span teams: model lifecycle, inference routing, and agent execution.
Own the quality of what you ship — code, test coverage, documentation, operability, and rollout safety. This is production infrastructure serving live inference, so reliability isn't an afterthought.
Make the engineers around you better through code review, design feedback, and being someone the team trusts with ambiguous, cross-cutting work.
Participate in on-call for the services your team owns, and help keep the team's processes and rituals healthy.
What You Will Bring:
10+ years of significant experience designing, building, and operating distributed systems or cloud-native backend infrastructure in production
.Strong working knowledge of Kubernetes and distributed-systems patterns (control loops, API servers, high-scale control planes), plus the fundamentals — containerization, networking, resource isolation.
Proficiency in at least one of Go, Java, or Python, and the willingness to work across all three.
A track record of leading cross-team technical work: turning ambiguous requirements into designs others can rally behind.
Excellent written and verbal communication — you can write a design doc that aligns people who don't report to you.
What Gives You an Edge:
Exposure to model serving, LLM/agent infrastructure, or streaming data systems.
You don't need a background in ML research or model training — this role is about building and operating the platform that serves AI reliably at scale, not inventing the models.
Ready to build what's next? Let’s get in motion.
Come As You Are
Belonging isn’t a perk here. It’s the baseline. We work across time zones and backgrounds, knowing the best ideas come from different perspectives. And we make space for everyone to lead, grow, and challenge what’s possible.
We’re proud to be an equal opportunity workplace. Employment decisions are based on job-related criteria, without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, veteran status, or any other classification protected by law.
Privacy Statement
Confluent is an IBM subsidiary which has been acquired by IBM and will be integrated into the IBM organization. By proceeding with this application, you understand that Confluent will share your personal information with other IBM affiliates involved in your recruitment process, wherever these are located. More Information on how IBM protects your personal information, including the safeguards in case of cross-border data transfer, are available here.
Redirects to Confluent's application page.
Other roles
More at Confluent.
Senior Manager, Detection & Response (Security Engineering)
Manager
Distributed Systems Software Engineer - WarpSteam
Senior
Senior Software Engineer - Streaming AI (Remote - Ontario / British Columbia)
Senior
Staff Software Engineer I
Staff
Senior Software Engineer
Senior