Senior Software Engineer - Cortex AI - FDE
Backend Engineer · Senior · Full Time
Opens Snowflake's application page
Role
What you'll do.
As a Senior Software Engineer on Snowflake's Cortex AI team, you will architect and build the distributed backend infrastructure powering enterprise-grade agentic AI systems. This role focuses on designing high-performance orchestration engines, scalable context retrieval systems, and production-hardened microservices that enable Snowflake Intelligence, Cortex Agents, and Search to operate reliably at enterprise scale. You'll combine deep systems thinking with AI infrastructure expertise to solve complex distributed systems challenges across agent execution, retrieval-augmented generation (RAG), and LLM evaluation frameworks.
Responsibilities
- Design and Build Agentic Runtime Orchestration: Architect scalable orchestration engines that execute complex agentic workflows with emphasis on low-latency tool execution, robust state management, and graceful failure handling across distributed infrastructure. Ensure the runtime supports multi-agent coordination, context propagation, and deterministic rollback capabilities for mission-critical enterprise deployments.
- Scale Context Engineering Infrastructure: Engineer high-performance systems for retrieval-augmented generation (RAG) including vector database integration, scalable search indexing strategies, advanced query processing optimization, result ranking algorithms, semantic caching layers, and automated metadata extraction pipelines. Balance indexing performance with query latency and memory efficiency across diverse data types and customer workloads.
- Develop Automated Evaluations Engine: Build production-grade infrastructure for running massive-scale golden set simulations, comprehensive error analysis pipelines, and hillclimbing experiments to systematically improve LLM and agent quality. Implement telemetry collection, metric aggregation, and automated quality regression detection to maintain system reliability standards.
- Productionize AI Workflows into Enterprise Services: Collaborate closely with AI modeling teams to transform raw LLM capabilities into hardened, multi-tenant microservices featuring strict execution guardrails, comprehensive observability instrumentation, audit logging, and graceful degradation strategies. Ensure compliance with enterprise data residency and regulatory requirements.
- Optimize System Performance and Cost Efficiency: Lead infrastructure strategy for intelligent model routing, prompt caching optimization, token consumption reduction, and compute resource allocation. Drive systematic performance profiling and optimization to ensure Snowflake's AI features deliver industry-leading efficiency metrics and cost-effectiveness for customers.
- Perform Cross-Layer System Debugging and Root Cause Analysis: Diagnose complex production issues by tracing requests across distributed services, analyzing logs and telemetry data, and identifying root causes in unfamiliar codebases. Develop deep understanding of system interdependencies and create preventive monitoring strategies to minimize future incidents.
- Drive Product Architecture Decisions: Assess customer problems to distinguish between bespoke requirements and platform gaps worthy of company-wide investment. Influence architectural choices across the backend to balance scalability, reliability, security, and developer velocity while maintaining long-term system maintainability.
- Establish Quality Metrics and Evaluation Frameworks: Define meaningful quality metrics for LLM and agent systems, design evaluation methodologies, and implement systematic processes for quality improvement over time. Collaborate with product and data science teams to establish baselines, track regressions, and validate improvements through rigorous testing.
Qualifications
What we look for.
Technical
Go or Java for Systems
Deep proficiency in Go or Java for building high-performance systems and backend infrastructure. Strong understanding of concurrency models, memory management, performance tuning, and production reliability patterns in your chosen language.
Python for AI Orchestration
Solid proficiency in Python for AI orchestration, integration with ML frameworks, data pipeline development, and prototype experimentation. Familiarity with Python async patterns and production deployment considerations.
Distributed Database Internals
Deep understanding of database architecture including indexing strategies, query optimization, transaction handling, replication, and sharding. Knowledge of both traditional SQL databases and modern distributed systems like FoundationDB.
Cloud-Native Architecture
Production experience with Kubernetes orchestration, container management, service meshes, and cloud-native design patterns. Understanding of infrastructure-as-code, configuration management, and operational best practices.
Vector Databases and Embeddings
Working knowledge of vector database technologies, embedding models, semantic search, and similarity algorithms. Understanding of vector index structures, approximate nearest neighbor search, and retrieval optimization.
LLM and Agent Platform Architecture
Familiarity with LLM fundamentals, prompt engineering, agent frameworks, tool calling mechanisms, and orchestration patterns. Understanding of LLM limitations, token economics, and cost optimization strategies.
LLM Evaluation Frameworks
Experience defining quality metrics for LLM and agent systems, building evaluation frameworks, implementing golden datasets, and using systematic evaluations to improve system quality and reliability over time.
Scalable Data Pipeline Architecture
Expertise designing and implementing scalable ETL/data pipeline systems handling large data volumes, implementing fault tolerance, backpressure handling, and efficient data transformation at scale.
Education
Bachelor's Degree in Computer Science
Bachelor's degree in Computer Science or equivalent technical field required. The degree should demonstrate foundational knowledge in algorithms, data structures, systems design, and software engineering principles essential for building distributed infrastructure.
Experience
Distributed Systems Architecture
7+ years building production distributed systems with emphasis on high-throughput, low-latency data processing and service orchestration. Demonstrated experience designing fault-tolerant systems, implementing consensus protocols, managing distributed state, and operating systems at scale.
Backend Infrastructure for AI/ML Products
Proven track record building backend infrastructure specifically for AI and ML products, including experience with model serving pipelines, feature engineering systems, data orchestration, and production AI workflows at meaningful scale.
Customer-Facing Technical Leadership
Direct experience in customer-facing technical roles where you've diagnosed complex production issues, communicated technical failures to frustrated customers with credibility, and balanced internal technical analysis with customer-safe communication strategies.
Cross-Layer System Debugging
Demonstrated mastery in debugging distributed systems by tracing requests across multiple services, correlating logs and telemetry data, root-causing failures in unfamiliar codebases without guesswork, and implementing systematic improvements based on findings.
Product Judgment and System Design
Ability to evaluate whether individual customer problems represent bespoke needs or systemic platform gaps worthy of broader investment. Experience influencing product direction through technical insights and architectural recommendations.
Skills
Required
Distributed Systems Design
Expert-level ability to architect distributed systems addressing scalability, consistency, fault tolerance, and operational complexity. Proven judgment in technology selection, tradeoff analysis, and long-term maintainability.
Backend API Development
Advanced capability designing and implementing high-throughput, low-latency REST and gRPC APIs with robust error handling, comprehensive observability, and strict performance SLAs.
Production Systems Reliability
Mastery in building reliable production systems including chaos engineering practices, comprehensive monitoring and alerting, incident response procedures, and postmortem-driven improvements.
Performance Profiling and Optimization
Advanced skills in identifying and eliminating performance bottlenecks through profiling, benchmarking, and optimization. Experience optimizing at multiple layers from algorithmic improvements to infrastructure decisions.
Multi-Tenant System Architecture
Experience designing systems that safely and efficiently serve multiple customers with proper isolation, fair resource allocation, and compliance with enterprise data residency requirements.
LLM Integration and Optimization
Practical expertise integrating LLMs into production systems, managing token costs, implementing caching strategies, and optimizing for latency and throughput constraints.
Technical Communication
Strong ability to communicate complex technical problems to diverse audiences including engineering teams, product leadership, and frustrated customers during production incidents.
Preferred
SQL Engine and Query Optimization
Nice to haveDeep familiarity with query optimization techniques, cost-based optimization, execution planning, and SQL engine internals. Experience optimizing complex analytical queries at scale.
Search Infrastructure Engineering
Nice to haveDirect experience building large-scale search infrastructure including inverted indices, ranking algorithms, relevance tuning, and search performance optimization.
Vector Database Implementation
Nice to haveHands-on experience building or deeply optimizing vector databases, approximate nearest neighbor search algorithms, and semantic search systems.
Production Machine Learning Infrastructure
Nice to haveExperience building production ML infrastructure including model serving systems, feature stores, training orchestration, and ML system observability.
Kubernetes and Container Orchestration
Nice to haveSophisticated experience designing and managing complex Kubernetes deployments, custom operators, resource optimization, and multi-cluster infrastructure.
Event Streaming Architectures
Nice to haveExperience with Apache Kafka, Pulsar, or similar event streaming platforms for building real-time data pipelines and event-driven system architectures.
Data Security and Compliance
Nice to haveKnowledge of encryption at rest and in transit, compliance frameworks relevant to enterprise data (SOC 2, HIPAA, GDPR), and secure multi-tenant data isolation patterns.
Infrastructure as Code
Nice to haveProficiency with Terraform, CloudFormation, or similar IaC tools for managing cloud infrastructure, enabling reproducible deployments and infrastructure versioning.
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 200,000 – 270,000
Full posting
Original listing.
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done.
The Cortex Apps team is building the future of AI for enterprise data. This role focuses on the backend infrastructure that powers our flagship products like Snowflake Intelligence, Cortex Agents and Search making agentic AI fast, reliable, scalable and secure at the enterprise level.
You won’t just be using AI tools; you will be building the high-performance systems that orchestrate them. You’ll own and influence the architecture for agent execution environments, high-throughput context retrieval, or the ecosystem that allows our customers to iterate and launch agents in production.
What you will do in this role:
Architect Agentic Runtimes: Build and scale the orchestration engines that execute complex agentic workflows, ensuring low-latency tool execution and robust state management.
Scale Context Engineering Infra: Design high-performance systems for RAG (Retrieval-Augmented Generation), including vector database integration, scalable and efficient search indexing, query processing, and result ranking, semantic caching, and automated metadata extraction.
Build the "Evals Engine": Develop the automated infrastructure required to run massive-scale golden set simulations, error analysis pipelines, and "hillclimbing" experiments.
Productionize AI Workflows: Collaborate with the modeling team to take raw LLM capabilities and turn them into hardened, multi-tenant microservices with strict guardrails and observability.
Optimize Performance & Cost: Direct the infra strategy for model routing, prompt caching, and token optimization to ensure Snowflake’s AI features are the most efficient in the industry.
Requirements:
Education: Bachelor’s degree in Computer Science or a related technical field.
Experience: 7+ years of experience building distributed systems, high-throughput APIs, or backend infrastructure for AI/ML products.
Technical Stack: Deep proficiency in Go or Java (for systems) and Python (for AI orchestration).
Systems Thinking: Strong understanding of database internals, distributed state management, and cloud-native architecture (Kubernetes, FoundationDB, etc.).
Domain Expertise: Familiarity with the "plumbing" of AI: vector indices, agent platforms, and building scalable data pipelines.
Experience in a customer-facing technical role — you have explained a hard failure to a frustrated external audience and been believed, you can produce both the internal analysis and the customer-safe version.
Product instinct — you can judge whether one customer's problem is bespoke or a platform gap worth fixing for everyone. This is the core judgment call.
Cross-layer debugging — tracing a request across services and root-causing in unfamiliar code from logs and telemetry, not guesswork.
Eval frameworks for LLM/agent systems — defining quality metrics and using evals to improve quality systematically over time.
Comfort with ambiguity on open-ended, externally-driven problems.
(Bonus) Experience with:
Query optimization and SQL engine internals.
Designing multi-tenant systems that handle sensitive enterprise data at scale.
Developing search infrastructure for large-scale applications.
Direct experience with any of the subsystems outlined above.
Snowflake is growing fast, and we’re scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake.
How do you want to make your impact?
For jobs located in the United States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.com
Redirects to Snowflake's application page.
Other roles
More at Snowflake.
Software Engineer AI Team
Mid
Staff/Principal AI Software Engineer - Snowflake CoWork
Principal
Sr Manager, Applied Field Engineering - AI/ML
Manager
Principal Data Platform Architect
Principal
Senior Software Engineer - NatSec
Senior