Senior Software Engineer - Cortex AI - FDE

Backend Engineer · Senior · Full Time

US-CA-Menlo ParkUSD 200k – 270k1w ago
Apply for this role

Opens Snowflake's application page

Role

What you'll do.

As a Senior Software Engineer on Snowflake's Cortex AI team, you will architect and build the distributed backend infrastructure powering enterprise-grade agentic AI systems. This role focuses on designing high-performance orchestration engines, scalable context retrieval systems, and production-hardened microservices that enable Snowflake Intelligence, Cortex Agents, and Search to operate reliably at enterprise scale. You'll combine deep systems thinking with AI infrastructure expertise to solve complex distributed systems challenges across agent execution, retrieval-augmented generation (RAG), and LLM evaluation frameworks.

Responsibilities

  • Design and Build Agentic Runtime Orchestration: Architect scalable orchestration engines that execute complex agentic workflows with emphasis on low-latency tool execution, robust state management, and graceful failure handling across distributed infrastructure. Ensure the runtime supports multi-agent coordination, context propagation, and deterministic rollback capabilities for mission-critical enterprise deployments.
  • Scale Context Engineering Infrastructure: Engineer high-performance systems for retrieval-augmented generation (RAG) including vector database integration, scalable search indexing strategies, advanced query processing optimization, result ranking algorithms, semantic caching layers, and automated metadata extraction pipelines. Balance indexing performance with query latency and memory efficiency across diverse data types and customer workloads.
  • Develop Automated Evaluations Engine: Build production-grade infrastructure for running massive-scale golden set simulations, comprehensive error analysis pipelines, and hillclimbing experiments to systematically improve LLM and agent quality. Implement telemetry collection, metric aggregation, and automated quality regression detection to maintain system reliability standards.
  • Productionize AI Workflows into Enterprise Services: Collaborate closely with AI modeling teams to transform raw LLM capabilities into hardened, multi-tenant microservices featuring strict execution guardrails, comprehensive observability instrumentation, audit logging, and graceful degradation strategies. Ensure compliance with enterprise data residency and regulatory requirements.
  • Optimize System Performance and Cost Efficiency: Lead infrastructure strategy for intelligent model routing, prompt caching optimization, token consumption reduction, and compute resource allocation. Drive systematic performance profiling and optimization to ensure Snowflake's AI features deliver industry-leading efficiency metrics and cost-effectiveness for customers.
  • Perform Cross-Layer System Debugging and Root Cause Analysis: Diagnose complex production issues by tracing requests across distributed services, analyzing logs and telemetry data, and identifying root causes in unfamiliar codebases. Develop deep understanding of system interdependencies and create preventive monitoring strategies to minimize future incidents.
  • Drive Product Architecture Decisions: Assess customer problems to distinguish between bespoke requirements and platform gaps worthy of company-wide investment. Influence architectural choices across the backend to balance scalability, reliability, security, and developer velocity while maintaining long-term system maintainability.
  • Establish Quality Metrics and Evaluation Frameworks: Define meaningful quality metrics for LLM and agent systems, design evaluation methodologies, and implement systematic processes for quality improvement over time. Collaborate with product and data science teams to establish baselines, track regressions, and validate improvements through rigorous testing.

Qualifications

What we look for.

Technical

  • Go or Java for Systems

    Deep proficiency in Go or Java for building high-performance systems and backend infrastructure. Strong understanding of concurrency models, memory management, performance tuning, and production reliability patterns in your chosen language.

  • Python for AI Orchestration

    Solid proficiency in Python for AI orchestration, integration with ML frameworks, data pipeline development, and prototype experimentation. Familiarity with Python async patterns and production deployment considerations.

  • Distributed Database Internals

    Deep understanding of database architecture including indexing strategies, query optimization, transaction handling, replication, and sharding. Knowledge of both traditional SQL databases and modern distributed systems like FoundationDB.

  • Cloud-Native Architecture

    Production experience with Kubernetes orchestration, container management, service meshes, and cloud-native design patterns. Understanding of infrastructure-as-code, configuration management, and operational best practices.

  • Vector Databases and Embeddings

    Working knowledge of vector database technologies, embedding models, semantic search, and similarity algorithms. Understanding of vector index structures, approximate nearest neighbor search, and retrieval optimization.

  • LLM and Agent Platform Architecture

    Familiarity with LLM fundamentals, prompt engineering, agent frameworks, tool calling mechanisms, and orchestration patterns. Understanding of LLM limitations, token economics, and cost optimization strategies.

  • LLM Evaluation Frameworks

    Experience defining quality metrics for LLM and agent systems, building evaluation frameworks, implementing golden datasets, and using systematic evaluations to improve system quality and reliability over time.

  • Scalable Data Pipeline Architecture

    Expertise designing and implementing scalable ETL/data pipeline systems handling large data volumes, implementing fault tolerance, backpressure handling, and efficient data transformation at scale.

Education

  • Bachelor's Degree in Computer Science

    Bachelor's degree in Computer Science or equivalent technical field required. The degree should demonstrate foundational knowledge in algorithms, data structures, systems design, and software engineering principles essential for building distributed infrastructure.

Experience

  • Distributed Systems Architecture

    7+ years building production distributed systems with emphasis on high-throughput, low-latency data processing and service orchestration. Demonstrated experience designing fault-tolerant systems, implementing consensus protocols, managing distributed state, and operating systems at scale.

  • Backend Infrastructure for AI/ML Products

    Proven track record building backend infrastructure specifically for AI and ML products, including experience with model serving pipelines, feature engineering systems, data orchestration, and production AI workflows at meaningful scale.

  • Customer-Facing Technical Leadership

    Direct experience in customer-facing technical roles where you've diagnosed complex production issues, communicated technical failures to frustrated customers with credibility, and balanced internal technical analysis with customer-safe communication strategies.

  • Cross-Layer System Debugging

    Demonstrated mastery in debugging distributed systems by tracing requests across multiple services, correlating logs and telemetry data, root-causing failures in unfamiliar codebases without guesswork, and implementing systematic improvements based on findings.

  • Product Judgment and System Design

    Ability to evaluate whether individual customer problems represent bespoke needs or systemic platform gaps worthy of broader investment. Experience influencing product direction through technical insights and architectural recommendations.

Skills

Required

  • Distributed Systems Design

    Expert-level ability to architect distributed systems addressing scalability, consistency, fault tolerance, and operational complexity. Proven judgment in technology selection, tradeoff analysis, and long-term maintainability.

  • Backend API Development

    Advanced capability designing and implementing high-throughput, low-latency REST and gRPC APIs with robust error handling, comprehensive observability, and strict performance SLAs.

  • Production Systems Reliability

    Mastery in building reliable production systems including chaos engineering practices, comprehensive monitoring and alerting, incident response procedures, and postmortem-driven improvements.

  • Performance Profiling and Optimization

    Advanced skills in identifying and eliminating performance bottlenecks through profiling, benchmarking, and optimization. Experience optimizing at multiple layers from algorithmic improvements to infrastructure decisions.

  • Multi-Tenant System Architecture

    Experience designing systems that safely and efficiently serve multiple customers with proper isolation, fair resource allocation, and compliance with enterprise data residency requirements.

  • LLM Integration and Optimization

    Practical expertise integrating LLMs into production systems, managing token costs, implementing caching strategies, and optimizing for latency and throughput constraints.

  • Technical Communication

    Strong ability to communicate complex technical problems to diverse audiences including engineering teams, product leadership, and frustrated customers during production incidents.

Preferred

  • SQL Engine and Query Optimization

    Nice to have

    Deep familiarity with query optimization techniques, cost-based optimization, execution planning, and SQL engine internals. Experience optimizing complex analytical queries at scale.

  • Search Infrastructure Engineering

    Nice to have

    Direct experience building large-scale search infrastructure including inverted indices, ranking algorithms, relevance tuning, and search performance optimization.

  • Vector Database Implementation

    Nice to have

    Hands-on experience building or deeply optimizing vector databases, approximate nearest neighbor search algorithms, and semantic search systems.

  • Production Machine Learning Infrastructure

    Nice to have

    Experience building production ML infrastructure including model serving systems, feature stores, training orchestration, and ML system observability.

  • Kubernetes and Container Orchestration

    Nice to have

    Sophisticated experience designing and managing complex Kubernetes deployments, custom operators, resource optimization, and multi-cluster infrastructure.

  • Event Streaming Architectures

    Nice to have

    Experience with Apache Kafka, Pulsar, or similar event streaming platforms for building real-time data pipelines and event-driven system architectures.

  • Data Security and Compliance

    Nice to have

    Knowledge of encryption at rest and in transit, compliance frameworks relevant to enterprise data (SOC 2, HIPAA, GDPR), and secure multi-tenant data isolation patterns.

  • Infrastructure as Code

    Nice to have

    Proficiency with Terraform, CloudFormation, or similar IaC tools for managing cloud infrastructure, enabling reproducible deployments and infrastructure versioning.

Tech stack

Languages

GoJavaPythonSQL

Frameworks

gRPCREST APIsAgent FrameworksFastAPI or similar

Databases

FoundationDBVector DatabasesPostgreSQLSnowflake

Tools

KubernetesDockerTerraformPrometheus and GrafanaELK Stack or similarCI/CD Pipelines

Other

Distributed System ConceptsLLM and RAG FundamentalsProduction ObservabilityPerformance EngineeringEnterprise Data Security

Compensation

Pay and benefits.

Base·USD 200,000 – 270,000

Full posting

Original listing.

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done.

The Cortex Apps team is building the future of AI for enterprise data. This role focuses on the backend infrastructure that powers our flagship products like Snowflake Intelligence, Cortex Agents and Search making agentic AI fast, reliable, scalable and secure at the enterprise level.

You won’t just be using AI tools; you will be building the high-performance systems that orchestrate them. You’ll own and influence the architecture for agent execution environments, high-throughput context retrieval, or the ecosystem that allows our customers to iterate and launch agents in production.

What you will do in this role:

  • Architect Agentic Runtimes: Build and scale the orchestration engines that execute complex agentic workflows, ensuring low-latency tool execution and robust state management.

  • Scale Context Engineering Infra: Design high-performance systems for RAG (Retrieval-Augmented Generation), including vector database integration, scalable and efficient search indexing, query processing, and result ranking, semantic caching, and automated metadata extraction.

  • Build the "Evals Engine": Develop the automated infrastructure required to run massive-scale golden set simulations, error analysis pipelines, and "hillclimbing" experiments.

  • Productionize AI Workflows: Collaborate with the modeling team to take raw LLM capabilities and turn them into hardened, multi-tenant microservices with strict guardrails and observability.

  • Optimize Performance & Cost: Direct the infra strategy for model routing, prompt caching, and token optimization to ensure Snowflake’s AI features are the most efficient in the industry.

Requirements:

  • Education: Bachelor’s degree in Computer Science or a related technical field.

  • Experience: 7+ years of experience building distributed systems, high-throughput APIs, or backend infrastructure for AI/ML products.

  • Technical Stack: Deep proficiency in Go or Java (for systems) and Python (for AI orchestration).

  • Systems Thinking: Strong understanding of database internals, distributed state management, and cloud-native architecture (Kubernetes, FoundationDB, etc.).

  • Domain Expertise: Familiarity with the "plumbing" of AI: vector indices, agent platforms, and building scalable data pipelines.

  • Experience in a customer-facing technical role — you have explained a hard failure to a frustrated external audience and been believed, you can produce both the internal analysis and the customer-safe version.

  • Product instinct — you can judge whether one customer's problem is bespoke or a platform gap worth fixing for everyone. This is the core judgment call.

  • Cross-layer debugging — tracing a request across services and root-causing in unfamiliar code from logs and telemetry, not guesswork.

  • Eval frameworks for LLM/agent systems — defining quality metrics and using evals to improve quality systematically over time.

  • Comfort with ambiguity on open-ended, externally-driven problems.

(Bonus) Experience with:

  • Query optimization and SQL engine internals.

  • Designing multi-tenant systems that handle sensitive enterprise data at scale.

  • Developing search infrastructure for large-scale applications.

  • Direct experience with any of the subsystems outlined above.

Snowflake is growing fast, and we’re scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake.

How do you want to make your impact?

For jobs located in the United States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.com

Redirects to Snowflake's application page.

Other roles

More at Snowflake.

View all 78 roles