Software Engineer, API Agents
Backend Engineer · Senior · Full Time
Opens OpenAI's application page
Role
What you'll do.
Join OpenAI's API Agents team to architect and build the shared backend infrastructure powering the next generation of autonomous AI agents. This role combines deep backend and infrastructure engineering with strong product judgment, focusing on building reliable systems for agent coordination, memory, tool execution, and safe production deployment across diverse enterprise and research workflows. You'll partner with research and product teams to translate frontier AI capabilities into dependable, scalable production primitives serving software engineering, healthcare, finance, and enterprise operations.
Responsibilities
- Design and Build Core Agent Infrastructure: Architect and implement the shared agent harness and backend services that power long-running, high-value workflows. This includes designing reliable abstractions for agent coordination, context management, and tool execution that abstract complexity while enabling diverse applications across product lines.
- Develop Multi-Agent Orchestration Systems: Build reusable capabilities for agent delegation, subagent coordination, and multi-agent workflows. Design systems that handle context propagation, state management, and safe transitions between agents while maintaining observability and control throughout execution.
- Implement Production Safety and Compliance: Establish foundations for secure agent execution including identity and permissions management, sandboxed execution environments, and guardrails. Design systems that enable safe deployment of autonomous capabilities while maintaining compliance and cost efficiency.
- Build Memory and Context Retrieval Systems: Develop distributed systems for agent memory management, context retrieval, and knowledge persistence. Implement search and information retrieval infrastructure that helps agents find relevant context for informed decision-making and action.
- Establish Observability and Evaluation Frameworks: Create comprehensive monitoring, logging, and evaluation infrastructure for agent systems. Build tools and pipelines that measure task completion rates, cost efficiency, latency characteristics, and reliability metrics across agent deployments.
- Cross-Functional Partnership and Translation: Collaborate with research, Codex, infrastructure, and applied product teams to translate model capabilities and research advances into production-grade shared systems. Lead technical discussions that bridge research prototypes to reliable, usable primitives for internal and external developers.
Qualifications
What we look for.
Technical
Backend Programming Languages
Proficiency in one or more backend languages such as Python, Go, Rust, or TypeScript with ability to navigate codebases across different languages. Strong fundamentals in systems programming, concurrency patterns, and performance optimization.
Distributed Systems Design
Deep understanding of distributed systems concepts including consistency models, consensus algorithms, transaction handling, and failure modes. Experience designing systems that balance performance, reliability, and operational complexity.
API and Service Architecture
Expertise in designing robust APIs, designing service boundaries, and managing cross-service concerns. Experience with API versioning, backward compatibility, and managing complex integrations across multiple services.
Database and Storage Systems
Strong knowledge of relational databases, NoSQL systems, caching layers, and distributed storage. Experience selecting appropriate storage technologies based on access patterns, consistency requirements, and scalability needs.
Systems Observability and Reliability
Expertise in designing logging, metrics, tracing, and alerting systems. Experience establishing SLOs, error budgets, and reliability practices. Familiarity with observability platforms and the ability to diagnose complex production issues.
Security and Access Control
Understanding of identity and permissions systems, authentication and authorization patterns, secure service-to-service communication, and secrets management. Experience designing systems that maintain security posture at scale.
Education
Computer Science or Related Field
Bachelor's degree in Computer Science, Computer Engineering, or related technical field preferred. Equivalent professional experience and demonstrated systems thinking can substitute for formal credentials.
Experience
Senior Backend Engineering Experience
7+ years of professional software engineering experience (excluding internships) in backend, infrastructure, platform, or product-driven roles. Demonstrated track record of designing and shipping complex distributed systems from ambiguous requirements to reliable production services at scale.
Distributed Systems and Backend Fundamentals
Proven expertise in designing distributed systems, APIs, microservices architectures, and workflow orchestration platforms. Experience building systems that handle scale, reliability, and operational concerns including load balancing, fault tolerance, and degradation patterns.
Production Systems Ownership
Experience leading complete system lifecycle from design through production operation. Track record of owning operational excellence, security hardening, reliability improvements, and cost optimization of critical backend services.
Infrastructure and Search Systems
Hands-on experience with infrastructure engineering, distributed storage, search engines, compute orchestration, or similar systems-level work. Comfortable working across multiple technical domains including identity systems, observability platforms, and execution environments.
Developer-Focused Platform Work
Experience building abstractions, SDKs, or platform primitives that other engineers depend on. Strong developer empathy and proven ability to simplify complexity into durable, usable interfaces.
Skills
Required
Backend Systems Design
Expert-level ability to design complex backend systems from first principles, including service boundaries, data models, failure modes, and operational characteristics.
Python or Go Programming
Production-level proficiency in Python or Go for building high-performance backend services. Strong understanding of language idioms, concurrency models, and ecosystem tooling.
Distributed Systems Architecture
Deep expertise in designing and reasoning about distributed systems including trade-offs between consistency, availability, and partition tolerance. Experience with message queues, event systems, and eventual consistency patterns.
Production Operations
Experience running production systems at scale with emphasis on reliability, observability, and incident response. Comfort making operational trade-offs and designing for maintainability.
API Design and Integration
Expertise in designing clean, extensible APIs that serve diverse client needs. Experience managing API contracts and ensuring backward compatibility across versions.
Technical Leadership and Communication
Ability to drive architectural decisions through clear technical communication with engineers, researchers, and product stakeholders. Experience translating complex systems into understandable abstractions.
Preferred
Workflow Orchestration
Nice to haveExperience with workflow orchestration platforms, long-running task management, or state machine implementations. Familiarity with systems like Temporal, Airflow, or similar tools.
Search and Information Retrieval Systems
Nice to haveExperience designing or working with search infrastructure, vector databases, semantic search, or information retrieval systems. Understanding of ranking algorithms and retrieval optimization.
Computer Vision or Multimodal AI Systems
Nice to haveBackground working with computer vision systems, multi-modal AI architectures, or tool calling frameworks. Experience interfacing AI capabilities with external systems and tools.
Rust Programming
Nice to haveProficiency in Rust for building performance-critical backend services. Experience with Rust's type system for preventing entire categories of runtime errors.
Database Query Optimization
Nice to haveExperience analyzing and optimizing complex database queries, understanding query planners, and designing optimal schema structures for performance.
Kubernetes and Container Orchestration
Nice to haveHands-on experience with Kubernetes, container orchestration, infrastructure as code, and cloud-native architecture patterns for scaling compute workloads.
Open Source Contribution
Nice to haveTrack record of contributing to open source infrastructure projects demonstrating ability to collaborate on complex technical systems and communicate with diverse technical communities.
Prior AI/ML Platform Experience
Nice to haveExperience building infrastructure for machine learning systems, ML training platforms, or AI model serving systems. Understanding of unique operational requirements for AI workloads.
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 293,000 – 385,000
Equity·Stock options
Process
Interview steps.
- 01
Initial Recruiter Screening
Conversation with OpenAI talent acquisition representative to confirm background, discuss experience with distributed systems and infrastructure, and assess alignment with team needs. Typical duration 30 minutes.
- 02
Technical Systems Design Interview
Deep-dive technical conversation focusing on backend architecture and systems design. You will work through a complex distributed systems problem, discussing trade-offs, design patterns, and operational considerations. Expect questions on API design, data model choices, and handling failure modes.
- 03
Backend Engineering and Implementation
Code-focused round evaluating proficiency in Python, Go, or your primary backend language. May involve designing a service interface, implementing a complex algorithm, or optimizing existing code. Focus is on clean systems-level thinking, not algorithmic puzzles.
- 04
Infrastructure and Operational Experience
Interview covering production operations, observability, reliability practices, and infrastructure concerns. Discussion of how you've debugged production issues, designed for operational simplicity, and managed scaling challenges.
- 05
Cross-Functional Collaboration Round
Conversation with researchers, product engineers, or infrastructure team members to assess communication skills, ability to translate requirements, and product judgment. Focus on past experiences integrating research capabilities or collaborating across teams.
- 06
Leadership and Team Interaction
Potential conversation with engineering leadership to evaluate decision-making approach, agency, and how you operate in fast-moving environments. Discussion of career goals and what you're seeking in a backend engineering role.
Full posting
Original listing.
About the Team
API Agents builds the shared agent harness, tools, and infrastructure that turn OpenAI’s frontier models into systems that can reliably complete real work. We carry the capabilities behind Codex into a much broader set of products and workflows across software engineering, research, finance, healthcare, enterprise operations, and more. Our work spans search and connected context, computer use, memory, delegation and multi-agent coordination, and safe execution. Sitting at the intersection of Research, Codex, infrastructure, and applied product teams, we build reusable agent capabilities that compound across the ecosystem.
About the Role
We are looking for an experienced backend software engineer to build the core systems behind the next generation of agents. You will design reliable services and abstractions that help agents find the right context, use tools and computers, retain knowledge, coordinate over long-running workflows, and take action safely. The role combines deep backend and infrastructure work with strong product judgment, with opportunities to work across agent runtimes, orchestration, search, execution environments, identity and permissions, observability, and evaluations. This is software and systems engineering rather than model training: success comes from strong backend fundamentals, high agency, and the ability to turn fast-moving research capabilities into dependable production primitives.
In this role, you will:
Design, build, and operate the shared agent harness and backend infrastructure that power long-running, high-value workflows across OpenAI and third-party products.
Build reusable capabilities across search and connected context, computer use, memory, tool execution, delegation, subagents, and multi-agent orchestration.
Establish the foundations agents need to operate safely in production, including secure execution environments, identity and permissions, observability, evaluations, reliability, and cost and latency efficiency.
Partner with Research, Codex, infrastructure, applied product teams, and external developers to translate model progress, traces, evaluations, and real workflow needs into shared capabilities and measurable gains in task completion, cost, and latency.
Your background might look something like:
7+ years of professional engineering experience, excluding internships, in relevant backend, infrastructure, platform, or product-driven engineering roles.
Exceptional backend engineering fundamentals and a track record of leading complex systems from ambiguous problem statements to reliable production services.
Proficiency in one or more backend languages such as Python, Go, Rust, TypeScript, or similar, with the ability to move comfortably across service, platform, and product boundaries.
Strong systems design and operational judgment across distributed systems, APIs, workflow orchestration, search or storage, compute and execution, identity and permissions, or observability and reliability.
Infrastructure depth paired with product judgment: you can abstract complex systems into simple, durable primitives that other engineers and developers can build on.
Strong developer empathy and communication skills, including experience partnering with users, research teams, and product engineers to make fast-moving capabilities useful and dependable.
High agency and comfort operating in a fast-moving environment, with strong ownership of security, reliability, and operational excellence.
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity.
We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.
For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement.
Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.
To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance.
We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link.
OpenAI Global Applicant Privacy Policy
At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.
Redirects to OpenAI's application page.
Other roles
More at OpenAI.
Software Engineer, API Multimodal
Senior
Manager, Forward Deployed Engineer (FDE), Life Sciences
Manager
Software Engineer, Ads Integrity
Senior
Forward Deployed Engineer - Zurich
Senior
Manager, Forward Deployed Engineer - Tokyo
Manager