Senior Machine Learning Engineer, Infrastructure

ML Engineer · Senior · Full Time

New YorkUSD 212k – 318k4d ago
Apply for this role

Opens Patreon's application page

Role

What you'll do.

As Senior Machine Learning Engineer, Infrastructure at Patreon, you'll architect and maintain high-throughput, low-latency live inference infrastructure for the Relevance team, which powers creator discovery and content ranking across a platform with 300,000+ creators and 25 million+ paid memberships. This role requires deep expertise in production ML infrastructure, feature store design, distributed systems, and Python backend engineering to build scalable systems that directly impact how fans discover creators globally.

Responsibilities

  • Architect and Scale Live Inference Infrastructure: Design, build, and maintain high-throughput, low-latency live inference infrastructure supporting real-time relevance systems that power creator discovery and content ranking across Patreon's platform serving millions of fans.
  • Own End-to-End Feature Store Lifecycle: Architect and manage complete feature store operations including data ingestion, transformation, and production serving, ensuring high availability, consistency between online and offline features, and reliable feature computation at scale.
  • Design Observability and Monitoring Systems: Implement comprehensive observability frameworks, monitoring dashboards, and validation systems to detect performance degradation, latency spikes, production data drift, and reliability issues in ML systems.
  • Enable Cross-Functional Collaboration: Partner with Product, Data Engineering, Trust & Safety, and other teams to translate product requirements into scalable infrastructure solutions and ensure alignment on ML platform capabilities and roadmap priorities.
  • Automate Deployment and Testing Pipelines: Build and maintain automated model deployment, canary testing, and reliability validation frameworks to accelerate developer velocity, reduce deployment friction, and ensure production system stability.
  • Debug and Optimize Complex Systems: Investigate and resolve performance bottlenecks, reliability issues, and latency problems in complex relevance systems by applying systematic debugging approaches and deep system understanding.

Qualifications

What we look for.

Technical

  • Production ML Infrastructure Design

    Demonstrated expertise building, deploying, and maintaining production-grade ML systems at scale, with proven experience architecting low-latency live inference pipelines and feature store systems serving millions of requests.

  • Distributed Systems Engineering

    Strong foundation in distributed systems architecture, understanding of consistency models, fault tolerance, scaling patterns, and ability to design systems handling high throughput and strict latency requirements.

  • Python Backend Development

    Proficiency writing robust, maintainable, production-quality Python code with experience in backend systems, performance optimization, and building infrastructure libraries used by multiple teams.

  • Performance Debugging and Optimization

    Systematic approach to debugging complex, high-throughput systems with ability to identify and resolve performance bottlenecks, profile code, analyze metrics, and optimize for latency and throughput.

  • Feature Engineering and MLOps

    Experience with feature store platforms, online/offline feature consistency, feature computation at scale, model serving infrastructure, and MLOps best practices.

Education

  • Computer Science or Related Field

    Bachelor's degree in Computer Science, Engineering, Mathematics, or equivalent professional experience demonstrating strong foundational knowledge in systems, algorithms, and software architecture.

Experience

  • Senior-Level ML Infrastructure Experience

    5+ years building ML systems in production environments with at least 3+ years focused on infrastructure, scaling, or backend ML systems responsible for high-availability production services.

  • Large-Scale System Architecture

    Proven track record architecting and owning 0-to-1 infrastructure systems that scale to production, remain reliable over time, and provide stable foundations for growing engineering teams.

  • Real-Time Systems Development

    Experience building low-latency, high-throughput systems with strict performance requirements, understanding tradeoffs between consistency, availability, and performance.

Skills

Required

  • Python

    Advanced proficiency in Python for backend systems, infrastructure libraries, and production code with focus on performance and maintainability.

  • Distributed Systems

    Deep understanding of distributed computing principles, consensus algorithms, fault tolerance, and architectural patterns for scaling systems.

  • Machine Learning Infrastructure

    Expertise in ML infrastructure components including model serving, feature stores, inference optimization, and production MLOps.

  • System Design

    Ability to design scalable, fault-tolerant systems with consideration for latency, throughput, consistency, and operational reliability.

  • Performance Optimization

    Skill in profiling systems, identifying bottlenecks, and optimizing for latency and throughput in production environments.

  • Technical Documentation

    Strong written communication skills for creating clear, comprehensive architecture documentation and infrastructure strategy guides.

Preferred

  • Feature Store Systems

    Nice to have

    Experience with feature store platforms like Tecton, Feast, or Databricks Feature Store for managing feature lifecycle at scale.

  • Kubernetes and Container Orchestration

    Nice to have

    Hands-on experience with Kubernetes, Docker, and container deployment for managing microservices and ML infrastructure.

  • Real-Time Stream Processing

    Nice to have

    Experience with streaming frameworks like Kafka, Flink, or Spark Streaming for real-time feature computation and data pipelines.

  • Model Serving Frameworks

    Nice to have

    Familiarity with model serving platforms like TensorFlow Serving, KServe, or similar systems for production inference serving.

  • Relevance Systems

    Nice to have

    Prior experience building ranking systems, recommendation engines, or search infrastructure that power user-facing discovery features.

  • Observability and Monitoring

    Nice to have

    Experience designing monitoring systems, implementing observability frameworks, and building dashboards for complex distributed systems.

  • Go or Rust

    Nice to have

    Experience with systems programming languages for building high-performance infrastructure components.

Tech stack

Languages

PythonGoSQL

Frameworks

Feature Store (Feast/Tecton/Databricks)Model Serving (TensorFlow Serving/KServe)FastAPI/Flask

Databases

PostgreSQLRedisApache CassandraDatastore/Firestore

Tools

KubernetesDockerApache KafkaPrometheus/GrafanaCI/CD (GitHub Actions/Jenkins)

Other

gRPCProtocol BuffersData Pipeline Orchestration

Compensation

Pay and benefits.

Base·USD 212,000 – 318,000

Equity·Stock options

Benefits

  • Competitive Equity Package

    Participate in Patreon's equity plan, offering ownership stake in a well-funded creator economy platform with significant market opportunity.

  • Comprehensive Healthcare Coverage

    Medical, dental, and vision insurance plans with company contribution to support your health and wellness.

  • Flexible Time Off

    Flexible paid time off policy allowing you to balance work and personal needs without strict accrual limits.

  • Recharge Days and Company Holidays

    Designated company-wide recharge days and holidays beyond standard PTO to promote team wellness and preventing burnout.

  • Learning and Development Stipend

    Annual budget for conferences, courses, certifications, and technical learning opportunities to advance your engineering skills and expertise.

  • Lifestyle Stipend

    Annual allowance for wellness activities, professional development, or quality-of-life improvements.

  • Commuter Benefits

    Pre-tax commuter benefits program for transportation costs to San Francisco or New York offices.

  • Patronage Program

    Monthly stipend to support creators on Patreon, aligning with company mission and enabling you to support content you love.

  • Parental Leave

    Generous parental leave policy supporting new parents with paid time away from work.

  • 401k Plan with Company Matching

    Retirement savings plan with company matching contribution to support long-term financial planning.

Process

Interview steps.

  1. 01

    Initial Screening Call

    Recruiter phone screen discussing background, ML infrastructure experience, interest in Patreon's mission, and alignment with role requirements.

  2. 02

    Technical Phone Interview

    Discussion of distributed systems design, ML infrastructure decisions, and previous experience building production systems. Expect questions on feature store architecture, inference pipeline design, and system scaling challenges.

  3. 03

    System Design Interview

    Collaborative whiteboarding or discussion of designing a low-latency inference system or feature store architecture. Evaluates ability to think through scale, reliability, consistency tradeoffs, and production considerations.

  4. 04

    Infrastructure Code Review and Discussion

    Review of infrastructure code (Python or similar) focusing on code quality, system thinking, performance considerations, and debugging approach. May include discussing production incidents or optimization challenges.

  5. 05

    Cross-Functional Collaboration Interview

    Meeting with Product or Data Engineering partner to assess communication skills, ability to translate requirements into infrastructure solutions, and cross-team collaboration experience.

  6. 06

    Final Round with Leadership

    Discussion with ML Infrastructure Lead or Engineering Manager covering strategic thinking, team dynamics, developer velocity mindset, and long-term vision for infrastructure systems at Patreon.

Full posting

Original listing.

Patreon is a media and community platform where over 300,000 creators give their biggest fans access to exclusive work and experiences. We offer creators a variety of ways to engage with their fans and build a lasting business including: paid memberships, free memberships, community chats, live video, and selling to fans directly with one-time purchases.

Ultimately our goal is simple: fund the creative class. And we're leaders in that space, with:

  • $10 billion+ generated by creators since Patreon's inception

  • 100 million+ free memberships for fans who may not be ready to pay just yet, and

  • 25 million+ paid memberships on Patreon today.

We're continuing to invest heavily in building the best creator platform with the best team in the creator economy and are looking for a Senior Machine Learning Engineer, Infrastructure to support our mission.

 

This role is based in San Francisco or New York as an in-office 3 days per week on a hybrid work model.

 

About the Team

You'll join the Relevance team, whose mission is to build the ML systems that power how fans discover creators and how content surfaces across Patreon. The team is responsible for search, feed ranking, and creator-fan matching. You'll work closely with a small, collaborative group of MLEs on shared infrastructure, code reviews, and roadmap alignment, while partnering cross-functionally with Product, Data Engineering, and Trust & Safety to deliver measurable impact across the platform.

About the Role

  • Architect, scale, and maintain high-throughput, low-latency live inference infrastructure to support our relevance systems.

  • Own the end-to-end feature store lifecycle—from ingestion and transformation to production serving, ensuring high availability and consistency between online and offline features.

  • Design and implement observability, monitoring, and validation frameworks to detect performance gaps, latency spikes, and production drift.

  • Collaborate with cross-functional partners, such as product, data engineering, and trust and safety, to translate product requirements into robust, scalable infrastructure solutions.

  • Automate model deployment and reliability testing to improve developer velocity and ensure system stability.

  • Debug complex relevance systems when monitoring identifies performance bottlenecks or reliability issues.

About You

  • You have deep experience building, deploying, and maintaining production-grade ML infrastructure at scale, specifically with low-latency live inference pipelines and feature store architectures.

  • You have a strong background in distributed systems and backend engineering, with the ability to write robust, maintainable code in Python.

  • You have a systematic approach to debugging complex, high-throughput systems and performance bottlenecks.

  • You are energized by building '0 to 1' infrastructure systems that stand the test of time and provide a reliable foundation for the team.

  • You possess strong communication skills and are effective at creating clear documentation for system architectures and infrastructure strategies.

  • You have a growth mindset, a keen eye for detail in code reviews, and a passion for empowering your teammates by improving developer velocity.

We hire talented and passionate people from different backgrounds because workplace diversity and inclusion is critical to our ability to serve creators worldwide. If you’re excited about a role but your past experience doesn’t match with every bullet point outlined above, we strongly encourage you to apply anyway. If you’re a creator at heart, are energized by our mission, and share our company values, we’d love to hear from you.

About Patreon

Patreon powers creators to do what they love and get paid by the people who love what they do. Our team is passionate about making this mission and our core values come to life every day in our work. Through this work, our Patronauts:

  • Put Creators First | They’re the reason we’re here. When creators win, we win.

  • Build with Craft | We sign our name to every deliverable, just like the creators we serve.

  • Make it Happen | We don’t quit. We learn and deliver.

  • Win Together | We grow as individuals. We win as a team.

 

Patreon is proud to be an equal opportunity employer. We provide employment opportunities without regard to age, race, color, ancestry, national origin, religion, disability, sex, gender identity or expression, sexual orientation, veteran status, or any other protected class. If you need a reasonable accommodation during the interview process, please let us know via email at accommodations@patreon.

 

Patreon offers a competitive benefits package including and not limited to salary, equity plans, healthcare, flexible time off, company holidays and recharge days, commuter benefits, lifestyle stipends, learning and development stipends, patronage, parental leave, and 401k plan with matching.

 

Patreon operates under a hybrid work model, where employees based in office locations are expected to come into the office two days per week, excluding sick time and paid leave. The goal of this policy is to be intentional about the in-person time we spend together to strengthen the feeling of community at Patreon. Candidates hired into remote-eligible roles are not expected to meet the same requirements.

At Patreon, we believe in fair and transparent pay. In compliance with New York and California pay transparency laws, we are sharing the expected salary range for this role.

 

The posted salary range is dependent on the location and the level. This range may encompass multiple levels within the role’s job family. The final offer will be based on candidate’s experience, skills, competencies, and geographic location, aligning with the appropriate job level within Patreon’s leveling framework. For remote employees located outside CA and NY, salary may vary based on location and local market conditions.

 

Patreon reserves the right to modify or update compensation and benefits at any time

Redirects to Patreon's application page.

Other roles

More at Patreon.