Senior Software Engineer — LLM Post-Training Platform
Senior Software Engineer · Senior · Full Time
Opens Snowflake's application page
Role
What you'll do.
Snowflake is seeking a Senior Software Engineer to join their ML Platform team, focusing on the Cortex Training LLM post-training platform. The ideal candidate will work on scaling distributed systems for GPU compute, designing APIs, and productionizing advanced ML infrastructure to enable enterprise-scale AI workloads.
Responsibilities
- Full Stack Design: Design and build across the full stack, from public training APIs and SDK to the control plane and GPU data plane
- Distributed Systems Scaling: Scale multi-tenant GPU scheduling, placement, and capacity-aware routing across regional GPU pools with built-in fault tolerance
- Performance Optimization: Drive end-to-end performance at scale, maintaining fast training, inference, and RL loops while keeping GPUs saturated under heavy concurrent load
- Research Productionization: Partner with Snowflake Research to transform state-of-the-art training and inference techniques into reliable, scalable enterprise-ready components
Qualifications
What we look for.
Technical
ML Systems Experience
5+ years of experience building and shipping production Machine Learning systems
Distributed Systems Expertise
Strong foundation in designing scalable, fault-tolerant services and operating them on Kubernetes in production
GPU and LLM Infrastructure
Proficiency with tools like PyTorch, DeepSpeed/FSDP, Ray, CUDA/NCCL, vLLM, with ability to debug across data, infrastructure, and GPU layers
Education
Minimum Educational Requirement
BS in Computer Science or related field
Advanced Degree Preference
MS or PhD is a plus
Experience
System Reliability
Demonstrated ability to harden complex systems for reliability, throughput, and cost efficiency
LLM Post-Training
Hands-on experience with LLM post-training and modeling is highly desirable
Skills
Required
Distributed Computing
Expert-level understanding of distributed system design and implementation
GPU Computing
Advanced knowledge of GPU infrastructure and parallel computing techniques
Machine Learning Infrastructure
Deep expertise in building scalable ML platforms and services
Preferred
LLM Post-Training
Nice to haveAdvanced experience with large language model fine-tuning and adaptation techniques
Research Translation
Nice to haveAbility to convert cutting-edge research into production-ready engineering solutions
Tech stack
Languages
Frameworks
Tools
Other
Compensation
Pay and benefits.
Base·USD 200,000 – 287,500
Equity·Stock options
Benefits
Health Insurance
Comprehensive medical, dental, and vision coverage
Stock Options
Equity compensation to align employee interests with company growth
Professional Development
Ongoing learning opportunities, conference attendance, and skill development programs
401(k) Plan
Retirement savings plan with company matching
Process
Interview steps.
- 01
Initial Screening
Phone or video call with recruiting team to assess initial fit and background
- 02
Technical Interview
In-depth technical assessment focusing on distributed systems, ML infrastructure, and GPU computing expertise
- 03
System Design Challenge
Evaluate candidate's ability to design scalable ML platforms and solve complex distributed computing problems
- 04
Team Fit Interview
Discussion with potential team members to assess collaboration and alignment with Snowflake's innovative culture
- 05
Final Interview
Meeting with hiring manager to discuss role expectations and candidate's potential impact
Full posting
Original listing.
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done.
Senior Software Engineer — LLM Post-Training Platform
The Snowflake ML Platform team's mission is to let customers run their most demanding ML/AI workloads inside Snowflake. Cortex Training is our LLM post-training platform: it turns scarce, expensive GPU capacity into a simple, composable service, so customers can adapt open-weight foundation models to their own business problems while we handle the hard distributed-systems parts, including scheduling, orchestration, multi-node training and inference, fault tolerance, and throughput.
The platform already runs post-training at scale. Under the hood, it decouples GPU computation from the training loop and exposes it as primitive APIs that compose into everything from SFT to full RL workflows. You'll work alongside a team that ships fast & sweats reliability and the researchers behind DeepSpeed. We're looking for an engineer who thrives in the ML infrastructure layer and brings a solid understanding of LLMs and post-training to help us scale and grow it.
YOU WILL:
Design and build across the full stack — from the public training APIs and SDK through the control plane to the GPU data plane.
Scale the distributed systems that make GPU compute serverless — multi-tenant scheduling, placement, and capacity-aware routing across regional GPU pools, with fault tolerance built in.
Drive end-to-end performance at scale — keep the training, inference, and RL loops fast and the data plane responsive under heavy concurrent load, with GPUs kept saturated.
Productionize research building blocks — partner with Snowflake Research to turn state-of-the-art training and inference techniques into reliable, composable components customers can run at enterprise scale.
QUALIFICATIONS:
5+ years building and shipping production ML systems
Strong distributed systems and infrastructure foundation — designing scalable, fault-tolerant services and operating them on Kubernetes in production.
Familiarity with GPU and LLM infrastructure — e.g., PyTorch, DeepSpeed/FSDP, Ray, CUDA/NCCL, vLLM; able to debug across the data, infrastructure, and GPU layers.
Demonstrated ability to harden complex systems for reliability, throughput, and cost efficiency.
BS in Computer Science or a related field (MS/PhD a plus).
(Bonus) Hands-on LLM post-training / modeling experience — the strongest candidates pair deep infra skills with real post-training intuition.
Snowflake is growing fast, and we’re scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake.
How do you want to make your impact?
For jobs located in the United States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.com
Redirects to Snowflake's application page.
Other roles
More at Snowflake.
Director of Engineering, AI Enterprise
Director
Software Engineer- Postgres
Mid
Software Engineer - SnowCommand
Mid
Senior Fullstack Engineer - Marketplace Monetization
Senior
Account Engineer
Mid