Notion

Software Engineer, AI Platform

Notion1 weeks ago
Location

San Francisco, California

Type

Full Time

Salary

USD 180,000 – 201,000

Level

Mid

Role

Backend Engineer

Posted

Jul 14, 2026

Full TimeMid

The role

Summary

Join Notion's AI Platform team as a Software Engineer to build the foundational systems and shared primitives powering AI capabilities across the workspace. You'll own the prototyping, development, and scaling of production-grade AI infrastructure that enables product teams to ship features faster while maintaining reliability, quality, and cost efficiency. This role requires expertise in large-scale systems engineering, platform architecture, and the ability to navigate the rapidly evolving AI/ML landscape with pragmatism and extreme ownership.

What you'll do

Own AI Platform Primitives Development: Lead the prototyping, development, and scaling of core AI platform systems and shared primitives including model integrations, context management, long-running actions, and cost/performance optimization frameworks. Drive architectural decisions that balance quality, reliability, and execution speed.
Enable Product Team Velocity: Partner closely with product teams across Notion to provide well-paved paths and production-ready guardrails that reduce duplicated work and accelerate AI feature shipping. Design extensible APIs and libraries that make AI capabilities easy for other engineers to adopt and integrate into products.
Operate Critical AI Systems in Production: Maintain and evolve critical AI infrastructure systems at scale, implementing comprehensive observability and diagnostics to understand provider and model behavior. Debug production failures, optimize latency and costs, and iterate on systems with minimal user disruption while handling increasing workloads.
Manage Model and Provider Evolution: Architect and implement versioning strategies, controlled rollout mechanisms, and compatibility layers to safely adopt new models and AI providers. Establish quality gates, reliability testing frameworks, and evaluation systems that validate production readiness and enable rapid safe migration across infrastructure.
Build Infrastructure for Scalable AI: Design and implement the infrastructure layer supporting reliability and availability across provider changes. Create shared systems for handling provider failure modes, managing API rate limits, implementing retry logic, and ensuring consistent behavior across multiple LLM providers and models at production scale.
Drive Cross-Layer System Improvements: Work across multiple layers including backend services, shared libraries, infrastructure components, and product integration points to identify optimization opportunities. Decompose complex system behavior, debug failures across layers, and implement pragmatic solutions that improve overall platform efficiency and maintainability.

What we look for

Technical

Large-Scale Systems ArchitectureDemonstrated experience designing and scaling systems that handle millions of users and high-throughput operations. Understanding of distributed system challenges including consistency, availability, reliability, and performance optimization at production scale.
Platform and Infrastructure EngineeringExperience building shared infrastructure, platform layers, or service-oriented architectures that enable multiple product teams to scale faster. Proficiency in creating abstractions that reduce complexity and operational overhead for downstream consumers.
AI/ML Systems KnowledgeWorking knowledge of large language models, model serving infrastructure, and practical AI deployment challenges. Understanding of LLM behavior, provider APIs, token management, cost optimization, and quality measurement in production environments.
Backend DevelopmentStrong backend software engineering fundamentals including API design, database optimization, and service architecture. Ability to implement scalable, maintainable systems while making pragmatic tradeoffs between quality, performance, and timeline.
Observability and DebuggingProficiency with observability tooling, logging frameworks, and distributed tracing to understand system behavior. Ability to debug production issues across multiple services and layers, identify root causes, and implement lasting solutions.

Education

Computer Science FoundationBachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience demonstrating strong computer science fundamentals and problem-solving capabilities.

Experience

Software Engineering Experience2-4 years of professional software engineering experience, ideally with at least one full product cycle from development through production scaling and operational management.
Platform/Infrastructure Team ExperiencePrior experience on LLM, ML platform, data infrastructure, or backend infrastructure teams that own critical shared systems supporting multiple consumer teams.
Production System ReliabilityHands-on experience operating and scaling production systems, including on-call responsibilities, incident response, and driving system reliability improvements under load.

Skills

Required skills

Passion for AI Systems at ScaleDeep commitment to building reliable, efficient platforms that abstract away complexity for other engineers. You've worked on teams scaling critical shared systems and understand how reliability, latency, cost, and quality challenges evolve as usage grows.
Adaptability and CuriosityComfort working in rapidly evolving technical landscapes where models, providers, and requirements change frequently. You enjoy understanding how systems behave in practice, can debug across multiple layers (backend, infrastructure, libraries, product), and use AI tools effectively to enhance your productivity.
Extreme OwnershipAbility to operate in ambiguous problem spaces, align stakeholders around solutions, and drive execution with accountability. You take ownership of platform outcomes including reliability, adoption, quality, and operational excellence, working effectively across team boundaries.
Thoughtful Problem-SolvingStrong ability to decompose complex system behavior, understand context before acting, and develop clean, pragmatic solutions. Comfortable debugging across multiple system layers independently and knowing when to seek help from teammates.
Pragmatic and Business-Oriented MindsetUnderstanding that platform engineering involves constant tradeoffs between quality, latency, cost, reliability, and execution speed. Ability to prioritize based on product and business impact while balancing technical craft with operational urgency.

Nice to have

Applied AI Product DevelopmentExperience building production AI features including prompt engineering, evaluation frameworks, model integrations, and quality measurement systems. Familiarity with practical challenges of deploying LLMs and iterating on AI capabilities with real user feedback.
Data Processing Pipeline ScalingExperience designing and scaling distributed data processing pipelines at massive scale using technologies like Apache Spark or Ray. Understanding of distributed computing patterns, resource optimization, and handling failures across compute clusters.
TypeScript and Node.js EcosystemFull-stack development experience using TypeScript and Node.js, including backend frameworks, build tooling, package management, and deployment patterns. Comfort building and deploying production systems in the Node.js ecosystem.
MLOps and ML Serving InfrastructurePrevious experience building MLOps platforms, model serving infrastructure, or ML deployment systems. Understanding of model versioning, experiment tracking, feature stores, inference optimization, and managing model lifecycle in production environments.

Compensation & benefits

Salary

USD 180,000 – 201,000 (annual)

Stock options

Available

Benefits

Competitive Equity Package

Participate in Notion's equity program, aligning your success with the company's long-term growth. As an early-stage high-growth company, equity grants provide meaningful wealth-building potential.

Comprehensive Health Coverage

Medical, dental, and vision insurance plans with employer contribution toward premiums. Access to preventive care and mental health services supporting your overall wellness.

Flexible Work Arrangements

Work from Notion offices in San Francisco or New York City on Mondays, Tuesdays, and Thursdays (Anchor Days), with flexibility on other days. This hybrid arrangement balances in-person collaboration with focused work time.

Professional Development Budget

Annual learning and development budget to support continuous growth. Invest in conferences, courses, certifications, and tools that enhance your engineering skills and platform expertise.

Paid Time Off

Generous vacation and paid time off policy supporting work-life balance. Notion encourages taking time for rest and personal pursuits as part of a sustainable work culture.

Retirement Planning Support

401(k) plan with employer matching to support long-term financial security. Access to financial planning resources and retirement investment options.

Parental Leave

Comprehensive parental leave policies supporting employees during major life transitions. Flexibility in returning to work and maintaining career trajectory.

Remote Work Technology

Equipment allowance and technology budget to support effective remote work. Resources to optimize your home office setup and work environment on non-anchor days.


Interview process

  1. 1
    Initial Screening Call 30-minute conversation with a recruiter to discuss your background, interest in the AI Platform role, and alignment with Notion's culture. This call focuses on understanding your experience with large-scale systems and AI infrastructure.
  2. 2
    Technical Phone Interview 60-minute technical discussion with an engineer on the AI Platform team. Expect questions about system design, distributed systems concepts, and your approach to debugging complex problems. You may discuss a recent project or architectural challenge you've solved.
  3. 3
    System Design Interview 90-minute deep-dive into designing an AI platform system (e.g., building abstractions for multi-provider LLM integration, designing an evaluation framework, or scaling a shared AI service). You'll work through tradeoffs, scalability, and reliability considerations while the interviewer asks probing questions.
  4. 4
    Behavioral and Platform Philosophy Discussion 60-minute conversation with AI Platform team leadership exploring your approach to platform engineering, ownership mentality, and cross-team collaboration. Discussion covers how you've handled ambiguity, prioritized platform adoption, and influenced technical decisions across teams.
  5. 5
    Final Panel or Offer Discussion Meeting with hiring manager and potentially other key stakeholders to discuss role expectations, team dynamics, and company vision. This is your opportunity to ask final questions and demonstrate enthusiasm for building AI infrastructure at Notion's scale.

Apply for this position

You'll be redirected to the company's application page


Notion

Notion

View all jobs

Notion is an American productivity software providing an all-in-one workspace for notes, tasks, and databases. Popular for its flexibility, it serves individuals and teams for project management and collaboration, evolving from a no-code tool to a versatile platform.

San Francisco, California, United StatesFounded 2013notion.so

Tech Stack

Languages
TypeScriptJavaScriptPython
Frameworks
Node.jsExpress or Similar REST FrameworksApache Spark or Ray
Databases
PostgreSQLRedisVector Databases
Tools
Docker and KubernetesCI/CD PipelinesObservability and Monitoring ToolsLLM APIs and Providers
Other
ML Model ServingDistributed Systems ConceptsGit and Version Control

Interview Guides

11 guides available for Notion

Apply Now