Senior Software Engineer - Cloud Efficiency
Senior Software Engineer · Senior · Full Time
Opens Snowflake's application page
Role
What you'll do.
Lead the evolution of cloud cost management at Snowflake by designing and building scalable observability solutions that transform billions of dollars in annual cloud spend into competitive advantage. This Senior Software Engineer role focuses on developing monitoring, attribution, and optimization products that empower thousands of engineers to continuously optimize per-unit costs across AWS, GCP, and Azure infrastructure. You'll tackle end-to-end challenges spanning data collection from heterogeneous cloud stacks, efficient storage architecture, real-time analytics dashboards, and AI-driven insights generation.
Responsibilities
- Design and Develop Scalable Cloud Efficiency Monitoring Solutions: Architect and build production-grade monitoring systems that aggregate cloud cost data alongside utilization, attribution, hardware performance, and architectural metrics across multiple cloud providers. Design efficient data collection pipelines that reliably ingest costs from AWS, GCP, and Azure while correlating with system performance metrics, establishing the foundation for actionable cost intelligence at scale.
- Build Automation and Tooling for System Intelligence: Develop advanced automation frameworks, alerting systems, and root cause analysis tools that surface critical efficiency insights. Create self-serve products that enable teams to quickly identify cost anomalies, performance bottlenecks, and optimization opportunities, reducing mean-time-to-resolution for cost-related incidents.
- Optimize Data Infrastructure for Scale and Performance: Lead optimization initiatives for data ingestion pipelines, storage architectures, and query performance across the cloud efficiency platform. Design systems capable of handling billions of dollars in annual cloud spend metrics while maintaining real-time query capabilities and sub-second dashboard response times for thousands of concurrent users.
- Drive Cross-Functional Collaboration and Integration: Partner with infrastructure, platform, finance, and product teams across Snowflake to deeply understand cloud efficiency requirements and monitoring gaps. Translate business objectives into technical solutions, ensuring seamless integration with existing observability infrastructure and establishing data contracts that serve multiple internal stakeholders.
- Ensure System Reliability and Production Excellence: Maintain high availability and reliability standards through participation in on-call rotations, incident management, and post-incident analysis. Implement comprehensive monitoring, alerting, and resilience patterns that ensure the cloud efficiency platform remains a trusted system of record for cost and performance insights enterprise-wide.
- Advance Industry Practices in Cloud Cost Observability: Contribute to open-source projects and establish industry best practices in cloud cost monitoring, distributed systems observability, and efficiency analytics. Share architectural innovations and solutions through technical documentation, internal presentations, and community engagement to position Snowflake as a thought leader in cloud economics.
Qualifications
What we look for.
Technical
Large-Scale Distributed Systems Design and Architecture
Proven expertise designing, implementing, and debugging distributed monitoring systems capable of handling massive throughput (multi-billion cost data points annually) while maintaining sub-second query latencies. Deep understanding of system tradeoffs including consistency, availability, partition tolerance, and practical experience with architectural patterns for observability platforms.
Cloud Cost Optimization and Economics
Hands-on experience optimizing systems on private clouds or major public cloud providers (AWS, GCP, Azure) with measurable ROI impact. Understanding of cloud pricing models, resource utilization patterns, cost allocation methodologies, and techniques for identifying and implementing cost optimization opportunities at infrastructure scale.
Performance and Efficiency Analysis
Advanced capability to conduct systemic analysis on complex systems to identify performance bottlenecks, cost drivers, and efficiency opportunities. Experience correlating multiple data dimensions (CPU, memory, I/O, network, cost) to surface actionable optimization insights and establish causal relationships between architectural decisions and cost outcomes.
Production Systems Engineering and Troubleshooting
Expert-level troubleshooting skills for complex production issues in distributed environments. Ability to rapidly diagnose root causes using system instrumentation, logs, metrics, and traces; implement targeted fixes; and establish monitoring to prevent recurrence. Comfortable with both reactive incident response and proactive reliability engineering.
Data Infrastructure and Observability Architecture
Strong background with observability platforms and tooling (OpenTelemetry, Prometheus, Grafana, or equivalent). Experience designing scalable telemetry collection, storage optimization for time-series data, and real-time query engines for analytics. Familiarity with monitoring best practices across multiple infrastructure paradigms.
Education
Bachelor's Degree in Computer Science, Engineering, or Related Technical Field
Foundational education in computer science, electrical engineering, mathematics, or equivalent technical discipline. Formal training in algorithms, data structures, systems design, and software engineering principles. Master's degree in Computer Science or related field preferred but not required for candidates with exceptional industry experience.
Experience
7+ Years Building Production Systems at Scale
Extensive industry experience designing, developing, and supporting large-scale systems in production environments. Track record of building systems that handle significant scale, reliability requirements, and complex operational challenges. Experience progressing from mid-level to senior engineer with increasing ownership of architectural decisions.
Cross-Functional Large Project Delivery
Demonstrated success delivering complex, multi-quarter initiatives requiring coordination across engineering teams, product, and business stakeholders. Ability to navigate dependencies, align on technical strategy, and drive projects to measurable business impact. Evidence of ownership beyond individual coding contributions.
Cloud Infrastructure and DevOps Experience
Background working with cloud infrastructure, containerization (Kubernetes, Docker), CI/CD pipelines, and infrastructure-as-code patterns. Practical exposure to operating systems at scale, capacity planning, and infrastructure optimization. Familiarity with cloud provider ecosystems and multi-cloud architectures.
Large Language Models and AI Integration (Preferred)
Prior hands-on experience leveraging LLMs for analysis, code generation, or intelligent system features. Understanding of how to integrate AI capabilities into observability platforms for automated insights, anomaly detection, or natural language query interfaces to accelerate cost analysis.
Skills
Required
Go Programming Language
Production expertise writing efficient, concurrent Go code for backend systems. Experience with goroutines, channels, and Go's ecosystem for building high-performance server applications and CLI tools. Comfort writing idiomatic Go for distributed systems.
Python
Strong proficiency in Python for data processing, automation scripting, and analytics. Experience with NumPy, Pandas, or similar libraries for working with large datasets. Ability to rapidly prototype solutions and integrate with data infrastructure.
Java
Solid understanding of Java for building scalable backend services. Experience with modern Java frameworks, concurrent programming patterns, and JVM internals optimization. Familiarity with Java ecosystem tools and libraries for systems programming.
Distributed Systems Fundamentals
Deep understanding of distributed computing principles including consensus algorithms, eventual consistency, CAP theorem, and failure modes. Experience architecting systems that maintain reliability despite partial failures and network partitions.
Problem-Solving and Troubleshooting Methodology
Systematic approach to diagnosing complex issues through hypothesis formation, targeted testing, and root cause analysis. Ability to navigate ambiguity in production incidents, prioritize debugging strategies, and communicate findings clearly to diverse technical audiences.
Technical Communication and Collaboration
Exceptional ability to articulate complex architectural concepts, tradeoffs, and design decisions to engineers at all levels and non-technical stakeholders. Strong listening skills to understand requirements deeply and collaborate effectively across teams.
Database and Data Storage Technologies
Hands-on experience with SQL and NoSQL databases, time-series databases, and data warehousing systems. Understanding of query optimization, indexing strategies, and schema design for analytical workloads. Familiarity with columnar formats and distributed query engines.
Preferred
OpenTelemetry and Observability Standards
Nice to haveExperience implementing or extending OpenTelemetry instrumentation across systems. Knowledge of OTEL architecture, semantic conventions, and ecosystem of exporters and processors for building extensible observability platforms.
Prometheus and Monitoring Tooling
Nice to haveHands-on experience with Prometheus, PromQL, and the broader CNCF monitoring ecosystem. Understanding of metrics instrumentation, cardinality management, and optimization techniques for large-scale metrics collection.
Kubernetes and Container Orchestration
Nice to havePractical experience deploying and operating systems on Kubernetes. Understanding of resource requests, limits, autoscaling, and Kubernetes internals. Experience troubleshooting container orchestration issues at scale.
Large Language Models and LLM Integration
Nice to haveHands-on experience with LLM APIs (GPT-4, Claude, etc.) or fine-tuned models. Understanding of prompt engineering, retrieval-augmented generation (RAG), and integration patterns for embedding AI into observability workflows.
Time-Series Data and Analytics
Nice to haveExperience working with time-series databases (InfluxDB, TimescaleDB, VictoriaMetrics) or data warehousing solutions optimized for analytical queries. Understanding of time-series compression, retention policies, and efficient aggregation patterns.
Rust Programming
Nice to haveFamiliarity with Rust for systems programming where performance and memory safety are critical. Experience leveraging Rust's concurrency model for high-throughput data processing pipelines.
AWS, GCP, and Azure Cloud Platforms
Nice to haveDeep practical experience across multiple public cloud providers. Knowledge of cost optimization features, resource management, and APIs for programmatic access to billing and utilization data across cloud stacks.
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 200,000 – 287,500
Full posting
Original listing.
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done.
Snowflake’s cloud spend is in billions of dollars per year. Hence, it is critical for us to govern and optimize our cloud spend, both for margins and long-term competitive advantage. Cloud Efficiency team’s charter is to build scalable products that enable governance, monitoring and optimization of cloud spend. Think of this as Observability for cloud costs and efficiency. The team’s vision is to “Transform cloud spend into a competitive advantage by empowering teams to continuously optimize the per-unit cost.”
In order to improve the overall cloud efficiency (i.e. cost per unit), it is critical to build monitoring products that collate costs with other factors such as utilization, attribution, hardware performance and architecture. Hence, there is an opportunity to build a unified, self-serve cloud efficiency product across Snowflake, that delivers actionable, real-time efficiency datasets through streamlined user experiences. This will enable thousands of engineers at Snowflake and will elevate cloud efficiency at Snowflake for long-term success.
When developing these solutions, we think about the problem end-to-end: how do we collect data from different stacks (e.g. costs from AWS, GCP, Azure and CPU, Memory, Utilization) across Snowflake reliably, how do we store it efficiently, how do we present this information to the user, what actions do we take on the insights from the data? What AI skills and agents we need to develop to enable quick and accurate cost analysis and identifying insights?
We are actively looking for a senior software engineer. If you love solving problems at scale, prefer to write scalable, reliable, and testable software, are an ace troubleshooter, and are deeply technical, then this is the role for you! Snowflake’s target of landing on the S&P 500 in combination with our rapidly growing portfolio in a highly innovative space demands more engineering maturity in the above. While these four domains describe the shape of our current goals, the engineer would be expected to drive the strategy and deliverables for clear company impact.
AS A SOFTWARE ENGINEER AT SNOWFLAKE YOU WILL:
Design, develop, and maintain scalable monitoring solutions for cloud cost and efficiency
Build tools and automation to enhance system monitoring, alerting, and root cause analysis.
Improve and optimize data ingestion, storage, and query efficiency for monitoring cloud efficiency data at scale.
Collaborate with teams across Snowflake to understand monitoring needs and implement solutions that improve operational visibility.
Contribute to open-source and industry best practices in monitoring and distributed systems monitoring.
Ensure high availability, reliability, and performance of monitoring systems by participating in on-call rotations and incident management.
WHAT WE LOOK FOR:
7+ years of industry experience designing, building and supporting large scale systems in production.
Experience with cost optimization and efficiency work on systems built on large private clouds or public cloud providers
Deep system and architectural analysis experience to identify actionable performance and efficiency insights
History of delivering large cross functional projects with measurable impact
Proficiency in programming languages such as Go, Python or Java.
Excellent problem-solving skills and ability to troubleshoot complex issues in a production environment.
Strong communication skills and ability to collaborate effectively in a team environment.
BS / MS in Computer Science, Engineering or related fields
Experience developing or using observability infrastructure such as OpenTelemetry, Prometheus etc is a plus.
Prior background working with LLMs is a plus
Snowflake is growing fast, and we’re scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake.
How do you want to make your impact?
For jobs located in the United States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.com
Redirects to Snowflake's application page.
Other roles
More at Snowflake.
Senior Software Engineer - Cloud Security
Senior
Senior AI/ML Architect, Applied Field Engineering
Senior
Staff Software Engineer - Query Processing (Snowtrail)
Staff
Staff Software Engineer - Snowflake Feature Store
Staff
Software Engineer - Dynamic Tables
Mid