Staff Software Engineer

Staff · Full Time · Remote

Mountain View, California · RemoteUSD 236k – 277k1mo ago
Apply for this role

Opens Confluent's application page

Role

What you'll do.

Staff Software Engineer at Confluent will design and build backend services (Go, Java, Python) for real-time AI inference on streaming data within Confluent Cloud. This role requires 10+ years of distributed systems experience and demands end-to-end ownership of complex, cross-team technical initiatives spanning model lifecycle management, inference routing, and agent execution on production infrastructure.

Responsibilities

  • Design and Build AI Inference Backend Services: Design and develop scalable backend services primarily in Go, Java, and Python that execute AI model inference and agent execution on real-time streaming data within the Confluent Cloud platform, ensuring low-latency performance and high throughput.
  • Own End-to-End Feature Delivery: Take full ownership of significant product features from conception through production, including drafting comprehensive technical designs, aligning stakeholders across teams, driving design decisions to completion, and managing the complete delivery lifecycle.
  • Make Cross-System Technical Decisions: Lead technical architecture decisions across multiple interconnected systems including model lifecycle management, inference request routing, agent orchestration, and inference serving layers, ensuring coherent system design and operational excellence.
  • Ensure Production Quality and Reliability: Maintain high standards for code quality, comprehensive test coverage, clear documentation, operational observability, and safe deployment practices for production inference infrastructure serving live traffic, with emphasis on reliability and incident prevention.
  • Mentor and Lead Technical Excellence: Elevate team capability through thorough code reviews, constructive design feedback, and personal accountability for complex cross-cutting technical work, building trust and establishing yourself as a go-to technical leader for ambiguous problems.
  • Participate in On-Call Rotation and Operations: Maintain operational responsibility for services owned by the team through on-call participation, incident response, and proactive process improvements to ensure sustainable team operations and continuous service health.

Qualifications

What we look for.

Technical

  • Distributed Systems Architecture

    Deep expertise designing, building, and operating distributed systems and cloud-native backend infrastructure in production environments at scale, including experience with system design tradeoffs, failure modes, and resilience patterns.

  • Kubernetes and Container Orchestration

    Strong working knowledge of Kubernetes architecture, deployment patterns, and operational best practices, combined with deep understanding of containerization technologies and container networking fundamentals.

  • Distributed Systems Patterns

    Proficiency with distributed systems design patterns including control loops, API server architecture, high-scale control plane design, eventual consistency models, and consensus algorithms in production systems.

  • Multi-Language Backend Development

    Proficiency in at least one of Go, Java, or Python with demonstrated ability to work effectively across all three languages, understanding language-specific performance characteristics and ecosystem tooling.

  • Backend Infrastructure and DevOps

    Expertise in building and operating production backend infrastructure, including logging, monitoring, alerting, tracing, infrastructure-as-code, CI/CD pipeline design, and deployment automation.

Education

  • Bachelor's Degree in Computer Science or Related Field

    Bachelor's degree in Computer Science, Software Engineering, or related technical discipline, or equivalent professional experience demonstrating deep computer science fundamentals.

Experience

  • Senior Distributed Systems Engineering

    10+ years of professional software engineering experience with at least 5-7 years focused on designing, implementing, and operating large-scale distributed systems in production environments.

  • Cross-Team Technical Leadership

    Demonstrated track record of leading ambiguous, cross-functional technical initiatives that span multiple teams and systems, translating unclear requirements into actionable technical designs that gain stakeholder alignment.

  • Production Infrastructure Operations

    Substantial experience owning production systems end-to-end, including operational responsibility through on-call rotations, incident response, postmortem analysis, and driving operational improvements.

  • System Design and Architecture

    Experience making high-impact architectural decisions for complex backend systems, considering scalability, reliability, cost, and maintainability across multiple services and deployment environments.

Skills

Required

  • Go Programming

    Proficiency in Go for building high-performance, concurrent systems with strong understanding of goroutines, channels, and Go's concurrency model.

  • Java Development

    Strong Java expertise including modern frameworks, JVM performance tuning, garbage collection, memory management, and building scalable backend services.

  • Python Backend Development

    Solid Python skills for backend service development, including async programming patterns, performance optimization, and integration with distributed systems.

  • Kubernetes

    Deep knowledge of Kubernetes architecture, API objects, scheduling, resource management, networking, storage, and operational patterns for production deployments.

  • Distributed Systems Fundamentals

    Core knowledge of CAP theorem, consistency models, fault tolerance, replication, data synchronization, and consensus mechanisms in distributed systems.

  • API Design and RESTful Services

    Expertise designing clean, maintainable APIs for backend services, including request/response serialization, error handling, versioning strategies, and SDK development.

  • System Design and Architecture

    Ability to design large-scale systems considering scalability, reliability, latency, throughput, and cost tradeoffs, and to articulate design decisions clearly to technical and non-technical stakeholders.

  • Technical Communication and Documentation

    Excellent written and verbal communication skills including ability to write clear design documents, architecture decision records (ADRs), and technical specifications that align cross-functional teams.

Preferred

  • Model Serving Infrastructure

    Nice to have

    Experience building or operating platforms for serving machine learning models at scale, including batching strategies, inference optimization, and model lifecycle management.

  • LLM and AI Agent Infrastructure

    Nice to have

    Exposure to large language model serving platforms, prompt management systems, or agent orchestration frameworks, understanding inference patterns and operational challenges.

  • Streaming Data Systems

    Nice to have

    Experience with streaming data platforms, event processing systems, or Kafka-based architectures, understanding real-time data pipelines and event-driven application patterns.

  • Apache Kafka

    Nice to have

    Familiarity with Apache Kafka architecture, broker design, consumer group management, and distributed topic partitioning for high-throughput event streaming.

  • gRPC and Protocol Buffers

    Nice to have

    Experience implementing high-performance RPC systems using gRPC and Protocol Buffers for efficient service-to-service communication in distributed systems.

  • Cloud Platform Services

    Nice to have

    Experience building on major cloud platforms (AWS, GCP, Azure) with understanding of managed services, auto-scaling, cost optimization, and cloud-native architecture patterns.

  • Observability and Monitoring

    Nice to have

    Strong background in designing observable systems with distributed tracing, structured logging, metrics collection, and building dashboards for production system visibility.

  • Control Plane Design

    Nice to have

    Experience designing high-scale control planes that manage cluster state, coordinate distributed work, handle failures gracefully, and maintain consistency at scale.

Tech stack

Languages

GoJavaPython

Frameworks

gRPCSpring BootFastAPIGin or Echo

Databases

PostgreSQLRedisDistributed Event Logs

Tools

KubernetesDockerHelmPrometheus and GrafanaJaeger or ZipkinGit and GitHub

Other

Control Loop ArchitectureAPI Server PatternModel Serving PatternsObservability-Driven DevelopmentInfrastructure as Code

Compensation

Pay and benefits.

Base·USD 235,700 – 277,000

Equity·Stock options

Benefits

  • Comprehensive Health Coverage

    Medical, dental, and vision insurance options with competitive premiums, covering preventive care, specialist visits, and mental health services.

  • Retirement Planning

    401(k) plan with employer match, helping you build long-term financial security with tax-advantaged savings options.

  • Paid Time Off

    Generous vacation days, sick leave, and paid holidays enabling work-life balance and personal wellness.

  • Professional Development

    Learning budget, conference attendance support, and access to online training platforms for continuous skill development and career growth.

  • Stock Options and Equity

    Opportunity to participate in company equity through stock options or RSUs, aligning your success with Confluent's growth and providing wealth-building potential.

  • Flexible Work Arrangements

    Remote-first culture supporting distributed teams across time zones with flexibility to work from home or office as needed.

  • Parental Leave

    Generous parental leave policies supporting new parents during critical early months of childcare.

  • Wellness Programs

    Fitness subsidies, mental health resources, wellness workshops, and employee assistance programs supporting holistic health.

Process

Interview steps.

  1. 01

    Initial Phone Screen

    Recruiter conducts 30-minute conversation to assess background, career trajectory, motivation, and alignment with Staff-level expectations. Expect discussion of past distributed systems work and cross-team technical leadership examples.

  2. 02

    Technical Architecture Discussion

    Engineer-led 1-hour conversation focused on system design thinking. You'll discuss approaches to large-scale problems similar to those in the role, such as designing inference routing systems or model lifecycle management. Bring examples of systems you've designed.

  3. 03

    Distributed Systems Deep Dive

    Senior engineer conducts 1.5-hour technical discussion covering Kubernetes internals, distributed consensus, fault tolerance, and operational patterns. Prepare to discuss production incidents, failure scenarios, and how you've debugged complex distributed system problems.

  4. 04

    Cross-Functional Collaboration Assessment

    Meeting with engineers from different teams to evaluate communication skills, ability to make sound technical decisions across boundaries, and how you approach building alignment on complex initiatives.

  5. 05

    Leadership and Vision Discussion

    Conversation with engineering manager or senior tech lead about your technical vision, how you mentor other engineers, your approach to end-to-end ownership, and your perspective on building systems at scale.

  6. 06

    Final Executive Round

    Brief conversation with director or VP-level executive covering career aspirations, culture fit, and strategic thinking about building AI infrastructure on streaming platforms.

Full posting

Original listing.

We’re not just building better tech. We’re rewriting how data moves and what the world can do with it. With Confluent, data doesn’t sit still. Our platform puts information in motion, streaming in near real-time so companies can react faster, build smarter, and deliver experiences as dynamic as the world around them.

It takes a certain kind of person to join this team. Those who ask hard questions, give honest feedback, and show up for each other. No egos, no solo acts. Just smart, curious humans pushing toward something bigger, together.

One Confluent. One Team. One Data Streaming Platform.

About the Role:

You'll help build Confluent Cloud's AI capabilities — the layer that lets customers bring AI and AI agents capabilities directly to their real-time data. Instead of moving data out to a separate system to run inference or build an agent, our customers do it in place, on streaming data, as part of the same platform they already use to move and process events at scale.

As an engineer, you'll own delivery of significant pieces of this product — not just writing code, but deciding how a capability should work across the services that make it up. The interesting problems here rarely live in one place: shipping something like inference-on-streaming-data or an AI agent that reacts to live events touches several systems at once — the user-facing API, the services that manage model and agent lifecycle, the control plane that schedules and runs the work, and the serving layer that actually executes inference. You'll be expected to reason across those boundaries, make sound design calls, and get engineers inside and outside the team aligned on the approach.

What You Will Do:

  • Design and build the backend services (primarily Go, Java, and Python) that run AI and model inference on real-time data.

  • Own features end to end — drafting the design, aligning stakeholders inside and outside the team, and driving the decision to a conclusion.

  • Make the technical calls on systems that span teams: model lifecycle, inference routing, and agent execution.

  • Own the quality of what you ship — code, test coverage, documentation, operability, and rollout safety. This is production infrastructure serving live inference, so reliability isn't an afterthought.

  • Make the engineers around you better through code review, design feedback, and being someone the team trusts with ambiguous, cross-cutting work.

  • Participate in on-call for the services your team owns, and help keep the team's processes and rituals healthy.

What You Will Bring:

  • 10+ years of significant experience designing, building, and operating distributed systems or cloud-native backend infrastructure in production

  • .Strong working knowledge of Kubernetes and distributed-systems patterns (control loops, API servers, high-scale control planes), plus the fundamentals — containerization, networking, resource isolation.

  • Proficiency in at least one of Go, Java, or Python, and the willingness to work across all three.

  • A track record of leading cross-team technical work: turning ambiguous requirements into designs others can rally behind.

  • Excellent written and verbal communication — you can write a design doc that aligns people who don't report to you.

What Gives You an Edge:

  • Exposure to model serving, LLM/agent infrastructure, or streaming data systems.

    • You don't need a background in ML research or model training — this role is about building and operating the platform that serves AI reliably at scale, not inventing the models.

Ready to build what's next? Let’s get in motion.

Come As You Are

Belonging isn’t a perk here. It’s the baseline. We work across time zones and backgrounds, knowing the best ideas come from different perspectives. And we make space for everyone to lead, grow, and challenge what’s possible.

We’re proud to be an equal opportunity workplace. Employment decisions are based on job-related criteria, without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, veteran status, or any other classification protected by law.

Privacy Statement

Confluent is an IBM subsidiary which has been acquired by IBM and will be integrated into the IBM organization. By proceeding with this application, you understand that Confluent will share your personal information with other IBM affiliates involved in your recruitment process, wherever these are located. More Information on how IBM protects your personal information, including the safeguards in case of cross-border data transfer, are available here.

Redirects to Confluent's application page.

Other roles

More at Confluent.

View all 12 roles