Senior Backend Engineer (ML Platform)

Backend Engineer · Senior · Full Time

New York City OfficeUSD 200k – 230k2d ago
Apply for this role

Opens Tennr's application page

Role

What you'll do.

Tennr is seeking a founding Senior ML Infrastructure Engineer to architect and scale cloud infrastructure powering AI-driven healthcare workflows. In this pivotal role, you'll design resilient systems for ML training, inference, and data pipelines while owning observability across the stack. This is an ideal opportunity for a distributed systems expert with 4+ years of production infrastructure experience who wants to grow into ML infrastructure in a fast-paced healthcare AI startup.

Responsibilities

  • Infrastructure Architecture & Scaling: Architect, build, and scale the cloud infrastructure behind ML training, inference, and data pipelines. Design infrastructure decisions that support model deployment at scale while optimizing for cost efficiency and performance in AWS environments.
  • Resilient System Design: Design and implement resilient systems for model deployment, evaluation, and monitoring that maintain reliability as traffic grows. Implement circuit breakers, graceful degradation, and fault tolerance patterns across the ML infrastructure stack.
  • Observability Implementation: Own observability across the entire stack including logging, metrics, tracing, and alerting. Build comprehensive monitoring dashboards and alerting rules to detect and respond to production issues proactively.
  • Production Troubleshooting & Optimization: Troubleshoot production incidents and continuously improve system performance and resource efficiency. Lead post-incident reviews and implement improvements to prevent recurring issues in distributed ML systems.
  • Cross-functional Collaboration: Collaborate with ML engineers, backend engineers, and product teams to integrate machine learning models cleanly with data pipelines and customer-facing products. Define and document APIs and integration patterns for ML model deployment.

Qualifications

What we look for.

Technical

  • Distributed Systems Expertise

    Demonstrated mastery in designing and implementing distributed systems with deep understanding of consistency, availability, and partition tolerance tradeoffs. Experience managing system state across multiple services.

  • Python Proficiency

    Strong production-grade Python development skills for building infrastructure tools, data pipelines, and backend services. Familiarity with async frameworks and performance optimization techniques.

  • TypeScript Proficiency

    Solid TypeScript skills for backend development, particularly in Node.js environments. Experience with type safety and scalable application architecture in JavaScript ecosystems.

  • AWS Platform Mastery

    Hands-on, production-level experience with AWS services including EC2, RDS, S3, SQS, Lambda, VPC, and infrastructure-as-code tools. Deep understanding of AWS cost optimization and architectural patterns.

  • PostgreSQL Database Administration

    Strong experience designing and maintaining PostgreSQL databases at scale including query optimization, indexing strategies, replication, and backup management. Understanding of ACID properties and transaction isolation.

  • Observability & Monitoring

    Expert-level knowledge of logging, metrics, distributed tracing, and alerting. Experience with tools like Datadog, New Relic, or similar platforms. Strong understanding of SLOs, SLIs, and error budgeting.

  • Production Incident Management

    Proven track record managing and resolving production incidents. Strong capability in root cause analysis, incident documentation, and implementing preventative measures in complex systems.

Education

  • Computer Science or Related Field

    Bachelor's degree in Computer Science, Software Engineering, or equivalent field preferred. Equivalent professional experience demonstrating mastery of computer science fundamentals will be considered.

Experience

  • Infrastructure Engineering

    Minimum 4+ years building and scaling production infrastructure in distributed systems, cloud platforms, or data engineering roles. Experience architecting systems that serve millions of requests or process large data volumes.

  • Backend Software Engineering

    Solid foundation in backend software engineering fundamentals with experience designing scalable, maintainable APIs and services. Understanding of microservices architecture, containerization, and deployment patterns.

  • Startup Mentality

    Comfortable navigating ambiguity in fast-paced startup environments with demonstrated high ownership. Track record of driving projects from concept through production deployment with minimal supervision.

  • ML Infrastructure Growth Path

    Demonstrated interest in growing into ML infrastructure and MLOps domains. While prior ML ops or infrastructure experience is not required, clear interest in expanding technical skills into machine learning systems is essential.

Skills

Required

  • Python

    Production-grade Python development for infrastructure and backend services

  • TypeScript/Node.js

    Backend development with TypeScript in Node.js environments

  • AWS

    Hands-on AWS architecture and cloud platform expertise

  • PostgreSQL

    Database design, optimization, and administration at scale

  • Distributed Systems Design

    System architecture for scalability, reliability, and performance

  • Observability Tools

    Logging, metrics, tracing, and alerting implementation and optimization

  • Infrastructure as Code

    Terraform, CloudFormation, or similar tools for declarative infrastructure management

Preferred

  • Kubernetes (k8s)

    Nice to have

    Container orchestration and Kubernetes cluster management experience

  • Inference Engines

    Nice to have

    Experience with vLLM, SGLang, TensorRT, or similar ML inference engines

  • ML Frameworks

    Nice to have

    Familiarity with PyTorch, TensorFlow, or other machine learning frameworks

  • Data Pipeline Tools

    Nice to have

    Experience with Apache Airflow, dbt, or similar data orchestration platforms

  • Healthcare Domain Knowledge

    Nice to have

    Understanding of healthcare workflows, insurance processes, or medical technology

  • Vector Databases

    Nice to have

    Experience with Pinecone, Weaviate, Qdrant, or similar vector database systems

Tech stack

Languages

PythonTypeScript

Frameworks

FastAPIExpress.jsCelery

Databases

PostgreSQLRedis

Tools

AWS EC2AWS RDSAWS S3TerraformDockerKubernetes

Other

vLLMSGLangTensorRTDatadog or New RelicGit

Compensation

Pay and benefits.

Base·USD 200,000 – 230,000

Equity·Stock options

Benefits

  • Unlimited PTO

    Flexible time off policy enabling you to manage work-life balance while maintaining productivity. Trust-based approach to vacation and professional development time.

  • 100% Paid Employee Health Benefits

    Comprehensive health insurance coverage including medical, dental, and vision plans fully funded by Tennr with zero employee premium contributions.

  • Employer-Funded 401(k) Match

    Competitive retirement savings program with employer matching contributions to help you build long-term financial security.

  • Competitive Parental Leave

    Generous parental leave policy supporting work-life transitions and family planning for all employees.

  • Modern Office Environment

    Beautiful new office space at 345 Hudson Street in Hudson Square, New York featuring collaborative spaces and modern amenities for 4 days/week in-office presence.

  • Free Lunch & Snacks

    Complimentary daily lunch and fully stocked snack pantry supporting nutrition and team bonding during working hours.

  • Equity Participation

    Stock options allowing founding team members to participate in company growth and share in long-term value creation.

Process

Interview steps.

  1. 01

    Technical Screening Call

    Initial 30-minute conversation with a recruiter to discuss your background, experience with distributed systems, and understanding of ML infrastructure challenges. Confirms alignment on role expectations and compensation.

  2. 02

    System Design Interview

    1-hour technical interview focusing on infrastructure architecture decisions. You'll design systems for ML training pipelines, inference serving, or data processing at scale. Emphasis on tradeoffs, scalability, and production reliability.

  3. 03

    Backend Engineering Deep Dive

    1-hour coding interview assessing Python and TypeScript proficiency. Problems focus on distributed systems concepts, database optimization, or infrastructure automation. Discussion of production considerations and edge cases.

  4. 04

    Infrastructure & Observability Discussion

    45-minute conversation with an ML infrastructure engineer covering monitoring strategies, incident response patterns, and your approach to building reliable systems. Discussion of past production incidents and resolution approaches.

  5. 05

    Leadership & Culture Fit

    30-minute meeting with a senior engineer or founder exploring your collaboration style, comfort with ambiguity in startups, and interest in growing into ML infrastructure domains. Assessment of alignment with Tennr's high-ownership culture.

  6. 06

    Offer & Negotiation

    Comprehensive compensation discussion including salary, equity details, and benefits. Opportunity to ask questions about technical roadmap, team structure, and long-term growth opportunities.

Full posting

Original listing.

Company Description

Today, when you go to your doctor and get referred to a specialist, your doctor sends out a referral and tells you, “They’ll be in touch soon.” So you wait. And wait. Sometimes days, weeks, or even months. Why? Because too often providers are overwhelmed with the painstakingly tedious work required to get paid by insurance companies. Powered by proprietary models, Tennr handles the complex paperwork that gets patients through the door and providers paid, helping operators get patients the right care, at the right time, in the right setting.

 

Role Description

We're looking for a founding Sr. ML Infrastructure Engineer with a strong background in distributed systems and building pipelines that scale. In this role, you'll own the infrastructure that powers Tennr's AI-driven healthcare platform - the training, inference, and data pipelines that let our models handle growing traffic and an expanding product surface.

Our ML team builds in-house, proprietary VLMs, LLMs, and other models purpose-built for hard problems in healthcare. You don't need deep ML experience to thrive here - what matters is a strong system design foundation and the interest to grow into the ML-side. If you think in systems, care about reliability, and want to expand into ML infrastructure, this is a rare chance to build foundational systems from the ground up.

 

Responsibilities

  • Architect, build, and scale the cloud infrastructure behind our ML training, inference, and data pipelines.

  • Design resilient systems for model deployment, evaluation, and monitoring that stay reliable as traffic grows.

  • Own observability across the stack - logging, metrics, tracing, and alerting.

  • Troubleshoot production issues and continuously improve performance and efficiency.

  • Collaborate with ML engineers, backend engineers, and cross-functional teams to integrate models cleanly with data pipelines and products.

 

Candidate Qualifications

  • 4+ years building and scaling infrastructure in production-distributed systems, cloud platforms, or data engineering.

  • Strong backend software engineering fundamentals, with proficiency in Python and TypeScript.

  • Hands-on experience with AWS, and PostgreSQL

  • Solid grasp of observability, reliability, and production incident response.

  • Comfortable with ambiguity and high ownership; you move fast and drive projects from idea to production in a startup environment.

  • Interested in growing into ML infrastructure - prior ML ops/infra experience is not required.

  • Nice to have: exposure to inference engines (vLLM, SGLang, TensorRT), or k8s.

 

Why Tennr?

  • Drive Impact: one of our company values is Cowboy, meaning you set the pace. You won’t just talk about things, you’ll get them done. And feel the impact.

  • Develop Operational Expertise: learn the inner workings of scaling systems, tools, and infrastructure

  • Innovate with Purpose: we’re not just doing this for fun (although we do have a lot of fun). At Tennr, you’ll join a high-caliber team maniacally focused on reducing patient delays across the U.S. healthcare system.

  • Build Relationships: collaborate and connect with like-minded, driven individuals in our Hudson Square office 4 days/week

  • Free lunch! Plus a pantry full of snacks.

 

Benefits

  • Beautiful new office at 345 Hudson Street

  • Unlimited PTO

  • 100% paid employee health benefit options

  • Employer-funded 401(k) match

  • Competitive parental leave

 
 





Redirects to Tennr's application page.

Other roles

More at Tennr.