OpenAI

Platform Engineer, Forward Deployed Engineering (FDE) -SF

OpenAI5 months ago
Location

San Francisco

Type

Full Time

Salary

USD 230,000 – 385,000

Level

Senior

Role

Backend Engineer

Posted

Feb 13, 2026

Full TimeSenior

The role

Summary

Join OpenAI's Forward Deployed Engineering team as a Platform Engineer to build next-generation platform capabilities that power enterprise AI deployments. You'll embed with customer-focused pods, architect scalable systems from scratch, and translate real-world customer patterns into durable platform abstractions. This role requires 5+ years of software or ML engineering experience with proven track record of shipping 0-to-1 products, exceptional cross-functional communication skills, and deep expertise in reliability, security, and systems design—perfect for senior engineers ready to shape the future of AI at scale.

What you'll do

Embed with Customer-Focused FDE Pods: Partner directly with customer-tagged Forward Deployed Engineering teams to provide hands-on technical leverage. Contribute substantively to architecture decisions, product shaping, system refactoring, hardening efforts, and implementation of reusable platform abstractions while preserving pod ownership of customer relationships and day-to-day execution.
Translate Cross-Customer Patterns into Platform Bets: Analyze repeated signals and feedback patterns across multiple customer deployments to identify generalizable opportunities. Synthesize observations into crisp platform hypotheses with well-defined success criteria, realistic scope estimates, and validation plans that respect customer constraints and go-to-market timelines.
Raise Engineering Standards Through Mentorship and Tooling: Establish organization-wide quality norms through high-signal code review, technical pair programming sessions, and mentorship. Design and build lightweight developer tooling that makes sound architecture, code readability, and correctness the default behavior across the entire FDE organization.
Lead Complex Platform Capabilities End-to-End: Act as Directly Responsible Individual (DRI) for high-leverage platform primitives such as the Context Platform. Own requirements gathering, technical design, implementation decisions, and production launch. Make critical tradeoffs explicit, document decisions, and maintain early customer engagement to ensure solutions address real deployment scenarios.
Collaborate Across Cross-Functional Platform Teams: Work seamlessly with B2B Product managers, customer-tagged FDE engineers, operations teams, and business partners to identify market opportunities, prioritize features, and bring platform capabilities to production. Translate technical constraints into business impact and ensure alignment between platform roadmap and customer needs.
Drive Structured Iteration and Learning: Establish instrumentation strategies, design rigorous evaluation frameworks, and implement error analysis processes to continuously improve platform capabilities. Default to systems thinking by converting ambiguous feedback, production failures, and escalations into durable product requirements rather than one-off fixes.
Ensure Reliability, Security, and Governance at Design-Time: Build platform capabilities with security, compliance, and reliability as first-class design concerns. Implement proper access controls (RBAC), maintain comprehensive auditability, establish clear data access boundaries, design safe rollout mechanisms, instrument observability, and enable rapid incident-driven hardening.
Communicate Technical Tradeoffs Across Audiences: Articulate complex technical decisions and platform tradeoffs to diverse stakeholders including engineering teams, product leadership, go-to-market leaders, and executive leadership. Simplify sophisticated concepts and connect technical implementation choices to measurable business outcomes and adoption impact.
Shape Product Requirements Based on Customer Deployments: Leverage direct exposure to production customer environments to inform platform design. Convert real-world operational insights into product requirements that address genuine customer pain points and deployment constraints rather than theoretical use cases.
Influence OpenAI Platform and Products: Contribute to OpenAI's broader platform evolution by embedding customer insights and technical innovations back into core platform offerings. Help determine what generalizes across customers, what remains customer-specific, and what "ready for handoff" criteria mean for platform maturity.
Participate in Production Incident Response and Hardening: Engage in on-call rotation for supported platforms and contribute to incident post-mortems. Drive systematic hardening efforts to prevent recurrence of production issues and continuously improve system resilience and observability.
Own Customer Technical Outcomes End-to-End: Take ownership of customer-adjacent technical projects from initial scoping through production adoption. Apply structured iteration methodologies including instrumentation, evaluation frameworks, error analysis, and continuous refinement of success metrics to maximize customer impact.
Build Reusable Abstractions and Patterns: Identify common patterns across customer deployments and design generalizable abstractions that reduce future engineering effort. Document architectural patterns, create reference implementations, and establish templates that enable faster, higher-quality customer deployments.

What we look for

Technical

Core Backend Architecture and Systems DesignExpert-level proficiency in designing scalable, reliable distributed systems. Deep understanding of system design principles including load balancing, caching strategies, database optimization, API design patterns, microservices architecture, and asynchronous processing. Proven ability to architect zero-to-one platform capabilities that become foundational infrastructure for other engineers.
Security and Compliance ArchitectureDemonstrated expertise in building security-first systems including role-based access control (RBAC) implementation, data governance, encryption strategies, secure API design, audit logging, and compliance frameworks relevant to enterprise deployments. Experience with security threat modeling and hardening production systems against common attack vectors.
Reliability and Production SystemsAdvanced proficiency in designing highly available, resilient systems. Strong background in monitoring and observability (metrics, logging, distributed tracing), incident response, error handling, graceful degradation, and systematic approaches to improving system reliability. Experience with zero-downtime deployments and canary rollout strategies.
Machine Learning Systems and DeploymentSolid understanding of ML engineering principles including model serving, prompt engineering, evaluation methodologies, error analysis frameworks, and the operational complexities of deploying and scaling ML models in production. Familiarity with tools and infrastructure for ML workflows and model lifecycle management.
API and Integration DesignExpertise in designing clean, well-documented APIs that support integrations across diverse customer environments. Understanding of versioning strategies, backwards compatibility, rate limiting, error handling, and creating intuitive abstractions that enable customer success with minimal support.
Data Pipeline and ProcessingProficiency in designing efficient data ingestion, transformation, and processing pipelines. Experience with both batch and real-time data processing, ensuring data quality, consistency, and traceability through production systems. Understanding of tools and patterns for building reliable ETL infrastructure.
Observability and Monitoring InstrumentationAdvanced ability to design comprehensive observability into systems from inception. Expertise in structured logging, metric collection, distributed tracing, and creating actionable dashboards that enable rapid issue diagnosis. Proven skill in building telemetry that answers critical business and technical questions.
Infrastructure and DevOps FundamentalsSolid understanding of cloud infrastructure (especially cloud platforms), containerization (Docker), orchestration (Kubernetes), infrastructure-as-code principles, CI/CD pipelines, and deployment automation. Experience setting up production-ready deployment and monitoring infrastructure.
Software Engineering Best PracticesRigorous commitment to code quality, maintainability, and testability. Expertise in test-driven development, comprehensive automated testing strategies (unit, integration, end-to-end), code review practices, refactoring techniques, and writing clear, well-documented code that enables team velocity.
Customer-Centric Technical ThinkingAbility to translate customer requirements into technical specifications while maintaining system integrity. Skill in scoping customer work to identify generalizable patterns, architecting solutions that balance customer-specific needs with platform reusability, and communicating technical constraints clearly to non-technical stakeholders.

Education

Bachelor's Degree in Computer Science or Related FieldStandard educational background covering fundamental computer science concepts, data structures, algorithms, systems design, and software engineering principles. Equivalent professional experience with demonstrated mastery of core computer science fundamentals is acceptable.
Advanced Systems Design KnowledgeDeep knowledge equivalent to advanced coursework or professional experience in distributed systems, databases, and large-scale system architecture. Self-taught expertise through building production systems at scale is equally valuable.

Experience

5+ Years in Software or ML Engineering at ScaleMinimum five years of professional software engineering or ML engineering experience, with significant time (at least 3+ years) building production systems at meaningful scale. Demonstrated progression in technical responsibility and impact, ideally including lead or senior individual contributor roles.
0-to-1 Product and Platform ShippingProven track record of shipping multiple 0-to-1 capabilities or products that other engineers, teams, or customers directly depend on. Evidence of taking technical ownership from concept through production adoption, including setting vision, making architectural tradeoffs, and driving user adoption.
Customer-Adjacent Technical OwnershipSubstantial experience owning technical projects with direct customer impact, from initial requirements and scoping through production adoption. Demonstrated ability to work closely with customers, understand their constraints, and design solutions that address real operational challenges rather than theoretical requirements.
Production Systems and ReliabilityHands-on experience building, operating, and hardening production systems where reliability, performance, and data security were critical design considerations. Background in on-call support, incident response, and systematic approaches to improving production stability.
High-Ambiguity and Fast-Iteration EnvironmentsThriving experience in high-ambiguity environments with rapidly changing requirements, typical of startups or product-centric companies. Demonstrated ability to make decisions with incomplete information, iterate quickly based on feedback, and maintain shipping velocity in uncertain conditions.
Structured Iteration and Data-Driven Decision MakingProven expertise in applying structured iteration methodologies including hypothesis development, instrumentation, evaluation, error analysis, and continuous metric refinement. Demonstrated ability to measure outcomes rigorously and make data-driven product and technical decisions.
Cross-Functional Collaboration and InfluenceExtensive experience working effectively across product, design, data science, go-to-market, and executive teams. Demonstrated ability to communicate complex technical ideas clearly, translate technical decisions into business impact, and drive alignment across diverse stakeholders toward shared outcomes.

Skills

Required skills

Backend Systems ArchitectureExpert-level system design including database design, API architecture, microservices patterns, scaling strategies, and decisions around when to build versus integrate with third-party solutions. Ability to architect systems that scale to enterprise customer bases.
Python or Go ProgrammingProduction-grade proficiency in at least one modern backend language, with Python or Go being highly preferred due to their prevalence in AI and infrastructure tooling. Strong fundamentals in software engineering practices including testing, documentation, and performance optimization.
Production Deployment and InfrastructureHands-on experience deploying and maintaining production systems, including working with cloud platforms (AWS, GCP, or Azure), containerization technologies, orchestration systems, and deployment automation. Understanding of infrastructure observability and monitoring.
Security and Access Control DesignAbility to design systems with security requirements at the forefront, including implementing role-based access control (RBAC), audit logging, data encryption, secure API authentication, and compliance with regulatory frameworks relevant to enterprise deployments.
API Design and IntegrationExpertise in designing clean, well-documented APIs that support diverse integration patterns. Experience with REST/GraphQL design principles, versioning strategies, handling rate limiting, designing for backward compatibility, and creating intuitive abstractions for customer integrations.
Observability, Monitoring, and DebuggingProficiency in designing comprehensive observability into systems including structured logging, metrics collection, distributed tracing, and alerting. Ability to rapidly diagnose and debug production issues using systematic approaches and tooling.
Code Review and Technical MentorshipAbility to conduct thoughtful, constructive code reviews that improve code quality and team learning. Comfort with pair programming and mentoring junior engineers to establish high technical standards across teams.
Data Analysis and Structured IterationAbility to design instrumentation, interpret data, and make data-driven product decisions. Experience with evaluation methodologies, error analysis, A/B testing, and incrementally improving systems based on evidence rather than assumptions.
Customer Communication and Technical TranslationStrong written and verbal communication skills, with ability to explain technical concepts clearly to both technical and non-technical audiences. Demonstrated ability to present platform capabilities and tradeoffs to customer stakeholders and translate feedback into technical requirements.
Systems Thinking and Problem DecompositionAbility to think holistically about complex problems, identify root causes versus symptoms, and design solutions that generalize across multiple use cases. Skill in decomposing ambiguous problems into well-scoped technical work with clear success criteria.

Nice to have

Experience with AI/ML Platforms and LLM DeploymentsBackground deploying and scaling machine learning models or large language models in production, including experience with prompt engineering, model serving infrastructure, evaluation frameworks for model quality, and operational considerations specific to AI systems.
Enterprise SaaS Product ExperienceExperience shipping enterprise software-as-a-service products, understanding enterprise buyer needs, deployment models, compliance requirements, and the operational demands of supporting large customer bases with diverse configurations and compliance needs.
Golang and Rust ExpertiseAdvanced proficiency in systems languages like Golang or Rust, particularly for building high-performance, concurrent systems. Strong background in memory safety and performance optimization would be valuable.
Database and Data InfrastructureDeep expertise in database design, selection between relational/NoSQL databases for specific use cases, query optimization, designing scalable data architectures, and experience with modern data warehousing or streaming platforms.
Kubernetes and Container OrchestrationProduction experience operating Kubernetes clusters, designing deployment strategies, managing stateful services, configuring networking, and solving operational challenges in containerized environments at scale.
Incident Response and Reliability EngineeringFormal training or substantial experience in incident response, root cause analysis, reliability engineering principles, designing for failure recovery, and building blameless post-mortem cultures that drive systematic improvements.
Product Strategy and Go-to-Market AcumenExperience working closely with product teams to shape market strategy, understand customer personas, plan product roadmaps, and make informed prioritization decisions about what to build and when to build it.
Open Source Community ContributionActive participation in open source projects, particularly infrastructure or developer tools, demonstrating ability to write code that other engineers depend on and capacity to gather and incorporate feedback from community users.
Public Speaking and Technical EvangelismComfort presenting to large audiences, writing technical blog posts, or evangelizing platform capabilities at industry conferences. Ability to clearly articulate why technical decisions matter and build enthusiasm around platform capabilities.
Leadership and Mentoring Track RecordInformal or formal leadership experience including mentoring junior engineers, leading technical initiatives, building consensus across teams, and demonstrated growth in your ability to influence at scale without requiring a formal title.

Compensation & benefits

Salary

USD 230,000 – 385,000 (annual)

Stock options

Available


Apply for this position

You'll be redirected to the company's application page