Senior Software Engineer, Cortex Quality
Senior Software Engineer · Senior · Full Time
Opens Snowflake's application page
Role
What you'll do.
As a Senior Software Engineer on the Cortex Quality team at Snowflake, you will own end-to-end quality, cost, and reliability metrics for Cortex Code, a production-grade AI coding agent used by thousands of enterprises. This measurement-first role combines systems thinking, AI/LLM expertise, and production debugging to transform research prototypes into dependable tools that data teams rely on daily, requiring 6+ years of software shipping experience with proficiency in Python, TypeScript, or Go and a rigorous approach to evaluation design and performance optimization.
Responsibilities
- Agent Quality Ownership and Evaluation Framework Design: Design, build, and maintain comprehensive evaluation harnesses that measure agent quality across real-world coding and data-engineering workflows. Develop rigorous metrics that account for sampling bias, ground-truth quality validation, leakage detection, and noise reduction. Diagnose failure modes from actual agent trajectories in production and establish measurement standards that prevent metric gaming while optimizing for genuine user outcomes.
- Production Reliability and Robustness Enhancement: Debug complex, unpredictable production failures by closing the loop from customer-reported broken runs to root cause identification, systematic fixes, and automated regression tests. Increase agent robustness through systematic debugging and establish frameworks to prevent recurring issues at Snowflake's enterprise scale.
- Cost and Latency Optimization Without Quality Regression: Implement advanced optimization techniques including prompt caching strategies, context compaction algorithms, tool-result offloading architectures, and intelligent model routing for sub-tasks. Balance efficiency improvements against quality metrics to maximize user value at scale, recognizing that cost reduction directly translates to improved user experience and product viability.
- Agent Prototype to Production Transition: Transform research-grade agent capabilities into production-ready systems by designing and refining agent behaviors for real workflows, establishing reliability standards, and implementing deployment strategies. Guide agents through the complete lifecycle from proof-of-concept to stable, mission-critical tools that enterprise data teams depend on.
- Real-World Data Pipeline and Feedback Loop Engineering: Architect and build data pipelines that capture production task insights, agent behavior patterns, and failure telemetry. Feed these insights back into evaluation frameworks and modeling workflows to create continuous feedback loops that enable data-driven agent improvements and inform frontier model onboarding decisions.
- Frontier Model Evaluation and Integration: Systematically measure new frontier models' capabilities, failure modes, latency characteristics, and cost profiles. Establish comparative benchmarks using rigorous evaluation methodology, determine optimal placement of models within the product architecture, and make data-driven recommendations for model selection and routing strategies.
- Cross-Functional Collaboration and Technical Communication: Partner with product, infrastructure, and research teams to shape user-facing agent behavior and establish quality standards for complex task completion. Translate technical quality metrics and cost analysis into clear, actionable insights for engineers, product managers, and leadership stakeholders. Lead collaborative decision-making on agent behavior changes and metric standards.
Qualifications
What we look for.
Technical
Production Software Shipping Experience with AI/LLM Features
Demonstrated track record of shipping complex software systems in production environments, with specific hands-on experience building, deploying, and maintaining AI/LLM-powered features at scale. Experience debugging production systems, optimizing performance under real-world constraints, and managing infrastructure reliability.
Multi-Language Programming Fluency
Strong proficiency in at least one of Python, TypeScript, or Go with demonstrated willingness and ability to work effectively across all three languages. Experience writing production-grade code, building systems that handle complex state and branching logic, and optimizing for performance and reliability.
Systems-Thinking and Measurement-First Methodology
Disciplined approach to designing evaluation frameworks that accurately reflect user outcomes without gaming metrics. Deep understanding of statistical sampling, ground-truth validation, bias detection, noise characterization, and metric correlation analysis. Ability to optimize for meaningful business and user outcomes rather than isolated performance indicators.
Complex System Debugging and Troubleshooting
Proven ability to debug unpredictable, real-world failures in distributed systems with multiple failure modes. Experience with production observability tools, log analysis, trace investigation, and systematic root cause analysis. Comfort with ambiguous problem statements and capability to design structured debugging approaches.
Data Pipeline Architecture and ETL Design
Experience designing and building scalable data pipelines that capture production insights, process large-scale datasets, and feed results back into modeling or decision-making systems. Familiarity with orchestration patterns, data quality management, and creating feedback loops in production systems.
Education
Bachelor's Degree in Computer Science, Engineering, Statistics, or Related Field
Required foundation in core computer science, engineering principles, or quantitative analysis. Demonstrates fundamental understanding of algorithms, data structures, systems design, and mathematical reasoning essential for evaluating complex agent behavior.
Master's Degree or Higher (Preferred)
Advanced degree in Computer Science, Machine Learning, Statistics, or related field preferred but not required. Indicates deeper expertise in specialized areas such as AI/ML systems, statistical methodology, or advanced systems architecture that aligns well with the role's measurement-first approach.
Experience
6+ Years of Production Software Shipping
Minimum six years of professional software engineering experience with a proven track record of shipping and maintaining production systems. This includes end-to-end ownership of feature delivery, production reliability, and system optimization in real-world environments serving actual users.
AI/LLM Feature Development and Deployment
Hands-on experience developing, testing, deploying, and supporting AI or LLM-powered features in production. Understanding of LLM capabilities, limitations, prompting strategies, evaluation challenges, and the operational complexities of shipping AI systems at scale to enterprise customers.
Complex System Ownership and Infrastructure Work
Experience building and owning complex systems including data pipelines, orchestration engines, or software with substantial state management, branching logic, and significant operational requirements. Demonstrates capability to take systems from conception through production maturity and ongoing optimization.
Skills
Required
Python Programming
Production-grade Python development for data processing, algorithm implementation, and system automation. Essential for building evaluation harnesses, data pipeline construction, and rapid prototyping of quality metrics.
LLM and Agentic AI Systems Understanding
Deep practical knowledge of large language models, agent architectures, prompt engineering, and the real-world behavior patterns of AI systems. Understanding of model capabilities, failure modes, latency characteristics, and cost implications of different model selection strategies.
Evaluation Framework Design
Ability to design rigorous evaluation harnesses that measure AI system quality without gaming metrics. Expertise in establishing ground truth, controlling for sampling bias, identifying metric leakage, and validating evaluation methodology against real user outcomes.
Production Systems Debugging
Systematic approach to identifying and resolving complex failures in production environments. Proficiency with observability tools, log analysis, distributed tracing, and the methodical debugging process required for systems with unpredictable behavior patterns.
Performance Optimization and Trade-off Analysis
Expertise in identifying optimization opportunities across cost, latency, and quality dimensions. Ability to implement optimization techniques including caching strategies, algorithmic improvements, and resource allocation optimization while maintaining system reliability.
Technical Communication and Stakeholder Management
Clear ability to translate complex technical metrics, experimental results, and system performance data into actionable insights for diverse audiences including engineers, product managers, and leadership. Strong written and verbal communication for documenting findings and building team alignment.
Preferred
Agentic Coding Tools Hands-On Experience
Nice to haveDeep practical familiarity with modern AI coding agents and their real-world behavior, performance characteristics, and limitations. First-hand understanding of how these tools fail, where they excel, and what drives user satisfaction or frustration in production usage.
Evaluation Harness and LLM Observability Development
Nice to havePrior experience building evaluation harnesses for language models, implementing observability systems for AI components, or developing safety and guardrail systems for production LLM deployment. Demonstrates existing expertise in the specific technical domain of the role.
Data Engineering and Analytics Background
Nice to haveProfessional experience in data engineering, data modeling, analytics architecture, retrieval-augmented generation (RAG) systems, or semantic layer development. Highly relevant for understanding data-centric coding agents and the workflows of Snowflake's target users.
TypeScript and Go Programming
Nice to haveProduction experience with TypeScript for backend systems or web services, and Go for systems programming and infrastructure tooling. Demonstrates polyglot development capability and experience across the full technology stack used by the team.
Large-Scale Data and Production Log Analysis
Nice to haveExperience working with large-scale datasets, production system logs, or enterprise telemetry systems. Proficiency with data exploration techniques, statistical analysis of real-world data, and extracting actionable insights from high-volume production data streams.
Model Selection and Routing Strategies
Nice to haveExperience implementing intelligent model routing, multi-model systems, or frameworks that dynamically select between different AI models based on task characteristics, cost constraints, or performance requirements. Understanding of when and why different models should be deployed for different use cases.
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 200,000 – 270,000
Equity·Stock options
Full posting
Original listing.
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done.
Cortex Code is Snowflake’s coding agent for building with data. It ships inside the platform that thousands of the world’s largest enterprises — including a large share of the Forbes Global 2000 — run their data on, which means the quality of this agent is felt by the data teams behind a meaningful slice of the global economy. We are taking coding agents from impressive demos to tools Data Science and Engineering teams depend on every day, and we hold them to a rigorous, public bar: see our data engineering agent benchmark.
About the Role
This is a measurement-first role that owns the quality and efficiency of Cortex Code end to end: how good the agent is, how much it costs to run, and how reliably it behaves in production. You will take agents from research capability to real, measurable user value — turning fuzzy “the agent feels worse” signals into hard metrics, running the experiments that move them, and shipping the changes that stick. You will work on a small, high-powered modeling and infrastructure team where your work reaches every developer building on Snowflake.
What you will do in this role
Take agents from prototype to production: design and refine agent behaviors for real coding and data-engineering workflows, and make them reliable enough to depend on.
Own agent quality end to end: build eval harnesses, diagnose failure modes from real agent trajectories, and run experiments that hillclimb the metrics that matter.
Drive down cost and latency without regressing quality: prompt caching, context compaction, tool-result offloading, and cheaper model routing for sub-tasks. At Snowflake scale, efficiency is user value.
Debug production failures and systematically increase robustness: close the loop from a customer’s broken run back to a fix and a regression test.
Build the data pipelines that feed real-world task insights back into evals and modeling.
Onboard and bake off frontier models: measure their strengths and failure modes, and decide where each belongs in the product.
Partner with product and infra: shape user-facing agent behavior, and set the metrics and standards that define a successfully completed complex task.
Requirements
Bachelor’s degree in Computer Science, Engineering, Statistics, or a related field. Master’s or higher preferred but not a requirement.
6+ years of experience shipping software in production, including AI/LLM features.
Fluency in at least one of Python, TypeScript, or Go, and willingness to work across all three
A measurement-first, systems-thinking instinct: you optimize for user outcomes over isolated metrics, and you can design an eval that is not fooling you (sampling, ground-truth quality, leakage, noise).
Comfort debugging complex, unpredictable, real-world failures.
Strong communication skills: you can make a quality or cost result legible to engineers, product, and leadership, and collaborate effectively in a team environment.
Nice to have
Deep hands-on experience with agentic coding tools and real intuition for model strengths, failure modes, and prompting limits.
Prior work on eval harnesses, LLM observability, or safety/guardrails in production.
Background in data engineering, data modeling, analytics, retrieval/RAG, or semantic layers, which is highly relevant for data-centric coding agents.
Experience working with large-scale datasets or production system logs.
You may be a particularly good fit if you
Are a power user of modern coding agents and want to turn that intuition into systematic measurement and improvement.
Have built and owned complex systems — pipelines, orchestration, or software with substantial state, branching logic, and operational requirements.
Thrive in high-intensity environments with short feedback loops and high standards for rigor.
Take problems to completion independently: you don’t stop at a prototype; you care about production reliability and clear metrics.
Are genuinely bothered by numbers that do not reconcile.
Snowflake is growing fast, and we’re scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake.
How do you want to make your impact?
For jobs located in the United States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.com
Redirects to Snowflake's application page.
Other roles
More at Snowflake.
Software Engineer AI Team
Mid
Staff/Principal AI Software Engineer - Snowflake CoWork
Principal
Sr Manager, Applied Field Engineering - AI/ML
Manager
Principal Data Platform Architect
Principal
Senior Software Engineer - NatSec
Senior