Principal Software Engineer, Agent Harness Bridge
Principal Engineer · Principal · Full Time
Opens OpenAI's application page
Role
What you'll do.
Lead the architecture and evolution of OpenAI's Agent Harness Bridge, a critical integration layer connecting the Codex harness with training infrastructure. This Principal-level role combines backend infrastructure expertise with cross-functional leadership, requiring strong systems design and API architecture skills to support large-scale training workloads while maintaining engineering quality and research enablement at the frontier of AI systems.
Responsibilities
- Design and Evolve the Agent Harness Bridge Architecture: Architect and lead the evolution of the integration layer between the Codex harness and research training infrastructure, ensuring alignment with product runtime capabilities and training system requirements. Make principled technical decisions about abstraction layers, API contracts, and system boundaries that serve both immediate training needs and long-term platform evolution.
- Own End-to-End Integration Surfaces: Take ownership of major integration points from initial architecture design through API specification, implementation, deployment orchestration, operational monitoring, and long-term maintenance. Establish clear accountability for system correctness, performance characteristics, and user experience across the research-to-infrastructure boundary.
- Build Reliable Containerized Execution Systems at Scale: Design and implement containerized runtime systems capable of executing demanding training workloads reliably at scale. Focus on resource isolation, failure handling, observability instrumentation, and graceful degradation to support the unpredictable resource demands and failure modes of large-scale distributed training.
- Lead Cross-Functional Technical Partnerships: Collaborate closely with research teams, agent systems developers, infrastructure engineers, and platform teams to understand emerging training use cases and evolving harness capabilities. Translate technical requirements into clean system designs while maintaining strong engineering discipline and preventing technical debt accumulation.
- Design Ergonomic Interfaces for Technical Users: Create clean, stable interfaces and workflows that enable highly technical internal users to move quickly without sacrificing correctness or operational safety. Apply developer experience principles to research platform design, anticipating user workflows and reducing friction in the training infrastructure feedback loop.
- Establish Durable Abstractions and Prevent Technical Debt: Build generalizable, well-documented abstractions that handle common patterns and edge cases, preventing one-off workarounds from calcifying into long-term technical debt. Establish clear ownership models and runbooks that enable other engineers to maintain and evolve systems reliably.
- Raise Engineering Standards and Operational Rigor: Elevate standards for correctness, reliability, operational maturity, and engineering judgment across this critical research-facing system. Establish best practices for error handling, monitoring, incident response, and system design that model excellence for the broader infrastructure organization.
Qualifications
What we look for.
Technical
Backend Systems Design and Implementation
Advanced expertise in designing and building scalable backend systems, including experience with distributed systems concepts, service architectures, and integration patterns. Demonstrated ability to make principled architectural trade-offs and establish patterns that scale across team and organizational boundaries.
API Design and Contract Definition
Deep proficiency in designing clean, stable, and well-documented APIs that serve both internal and external consumers. Strong experience with versioning strategies, backwards compatibility, and API evolution in fast-moving environments where breaking changes have organizational impact.
Python Proficiency
Expert-level Python development skills with experience in backend platform engineering, systems scripting, and infrastructure automation. Comfortable writing production-grade Python at scale with attention to code maintainability, performance, and testing strategies.
Infrastructure and Platform Engineering
Strong background in platform engineering, including containerization technologies (Docker, Kubernetes), infrastructure orchestration, and operational tooling. Experience building or maintaining internal developer platforms that serve multiple engineering teams or research groups.
Systems Reliability and Operational Excellence
Demonstrated expertise in designing reliable systems with rigorous attention to failure modes, observability, monitoring, and incident response. Experience establishing operational runbooks, alerts, and diagnostic tools that enable sustained reliability across complex systems.
Containerization and Runtime Execution
Hands-on experience with containerized execution environments, container orchestration, resource management, and execution runtime design. Understanding of sandbox isolation, process management, and the operational challenges of supporting demanding workloads at scale.
Education
Bachelor's Degree in Computer Science or Related Field
Bachelor's degree in Computer Science, Software Engineering, or related discipline. Equivalent experience demonstrating deep computer science fundamentals may be considered in lieu of formal degree.
Strong Foundation in Computer Science Fundamentals
Solid understanding of core computer science concepts including algorithms, data structures, distributed systems theory, and operating systems principles. Applied knowledge of these fundamentals reflected in system design decisions and engineering practice.
Experience
Senior Backend or Infrastructure Engineering
Minimum 8+ years of backend or infrastructure engineering experience with demonstrated progression to senior levels. Track record of leading design and implementation of critical systems that required cross-functional coordination and delivered measurable business or research impact.
Leading Cross-Functional Technical Initiatives
Proven experience leading technical efforts that spanned multiple teams, engineering organizations, or different functional areas (research, infrastructure, platform). Demonstrated ability to build consensus, establish clear technical direction, and execute complex changes with organizational buy-in.
Building for Technical Internal Users
Direct experience building internal platforms, developer tools, or research infrastructure that serves demanding technical users. Understanding of developer experience principles, the importance of ergonomic design for internal systems, and how to gather and incorporate feedback from sophisticated users.
Scaling Systems in Fast-Moving Environments
Track record of growing and scaling backend systems and infrastructure within organizations experiencing rapid change. Experience managing technical debt, establishing engineering discipline while maintaining velocity, and making pragmatic decisions about when to refactor versus when to push forward.
Research-Facing Infrastructure or ML Systems
Prior experience working on or near research infrastructure, machine learning systems, or scientific computing platforms. Understanding of the unique demands researchers place on infrastructure and the trade-offs between flexibility and stability in research environments.
Skills
Required
Python
Expert-level Python development with production backend experience and familiarity with infrastructure automation and systems programming patterns.
API Design and Contract Definition
Principled API design with deep understanding of versioning, backwards compatibility, and designing for scale and evolution.
Distributed Systems Design
Strong grasp of distributed systems concepts, service architecture patterns, and the operational challenges of scaling complex systems.
Systems Reliability and Observability
Expertise in designing reliable systems with comprehensive monitoring, alerting, and incident response infrastructure.
Containerization and Orchestration
Hands-on proficiency with Docker, container runtime systems, and orchestration platforms for managing containerized workloads.
Cross-Functional Leadership and Communication
Ability to lead technical discussions across diverse stakeholder groups, establish clarity on complex technical decisions, and drive consensus on architecture and implementation strategies.
Preferred
Rust
Nice to haveExperience with Rust for systems programming, particularly in performance-critical infrastructure components or where memory safety and concurrency guarantees provide architectural advantages.
Kubernetes Administration and Design
Nice to haveDeep experience operating and designing systems on Kubernetes, including custom resource definitions, operators, and scaling strategies for complex workloads.
Machine Learning Infrastructure
Nice to haveFamiliarity with ML training infrastructure, distributed training frameworks, or the specific challenges of supporting machine learning workloads at scale.
Research Infrastructure Experience
Nice to havePrior work on or adjacent to research platforms, scientific computing systems, or infrastructure designed to support experimental workflows.
gRPC and Protocol Buffers
Nice to haveExperience designing and implementing gRPC-based service architectures with Protocol Buffers, particularly in performance-sensitive or cross-language environments.
Async/Concurrent Programming Patterns
Nice to haveDeep experience with async programming models, concurrency primitives, and patterns for managing concurrent execution in Python or other languages.
Infrastructure as Code
Nice to haveProficiency with Infrastructure as Code tools and practices for managing complex infrastructure declaratively at scale.
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 347,000 – 490,000
Equity·Stock options
Benefits
Comprehensive Health Insurance
Medical, dental, and vision coverage with competitive plan options. OpenAI typically offers employer-covered premiums for employees and their dependents.
Equity Compensation
Competitive stock options or equity grants that align individual success with company growth, standard at principal engineering levels in high-growth AI companies.
Retirement Benefits
401(k) plan with employer matching contributions to support long-term financial security and retirement planning.
Unlimited Paid Time Off
Flexible vacation policy with no explicit limit, reflecting OpenAI's trust-based culture and work-life balance commitment.
Professional Development and Learning
Support for continued learning through conferences, training programs, and technical development aligned with career growth at the principal level.
Parental Leave
Comprehensive paid parental leave policies supporting both primary and secondary caregivers.
Mental Health and Wellness Support
Counseling services, mental health resources, and wellness programs to support employee wellbeing.
Commuter and Relocation Assistance
Support for Bay Area commuting and relocation assistance for those relocating to San Francisco for the role.
Process
Interview steps.
- 01
Initial Recruiter Screening
Phone or video conversation with OpenAI recruiter to assess background, experience level, motivation for the role, and general fit with OpenAI culture. Expect discussion of your infrastructure engineering background and leadership experience.
- 02
Technical Systems Design Interview
Depth technical interview focusing on systems architecture, API design principles, and infrastructure patterns. You will likely discuss past systems you have designed, trade-offs you made, and how you would approach designing the agent harness bridge architecture.
- 03
Backend Engineering Deep Dive
Detailed technical conversation with engineers on the infrastructure team covering Python proficiency, distributed systems knowledge, containerization experience, and approach to building reliable systems at scale. May include a design exercise.
- 04
Cross-Functional Collaboration Discussion
Conversation with cross-functional stakeholders (research team members, platform engineers, other infrastructure leads) to assess your ability to work effectively across teams, understand diverse needs, and communicate technical concepts to different audiences.
- 05
Leadership and Judgment Interview
Interview with senior engineering leadership assessing technical judgment, decision-making approach, experience mentoring and leading teams, and ability to establish standards and prevent technical debt in complex environments.
- 06
On-site or Final Round
Potential on-site interview in San Francisco with multiple team members for final assessment of cultural fit, communication style, and ability to thrive in OpenAI's fast-moving research environment.
Full posting
Original listing.
About the Team
OpenAI's research training infrastructure powers how our frontier models are trained and evaluated. The Agent Harness Bridge team sits at the boundary between the Codex harness, research infrastructure, and internal agent systems, ensuring the same agentic coding runtime used in product can also support large-scale training workloads.
This team owns the integration layer that connects harness capabilities into the training stack. The work is highly cross-functional and high leverage: researchers depend on it to run reliably, platform teams depend on it to evolve cleanly, and failures in this surface can materially affect training velocity and correctness.
About the Role
We're looking for a Principal Software Engineer to lead the architecture and evolution of the Agent Harness Bridge. You'll own a critical interface between research and engineering, building the systems, APIs, and operational patterns that let researchers use agentic coding infrastructure safely and effectively in training environments.
This role is ideal for a senior backend or infrastructure engineer with strong technical judgment, product sense for highly technical users, and the ability to drive execution across multiple teams. The highest-leverage work is not novel research; it is building robust infrastructure that accelerates research without compromising engineering quality.
In this role, you will
Design, build, and evolve the bridge between the Codex harness and research training infrastructure
Own major integration surfaces end-to-end, from architecture and API design through rollout, operations, and long-term maintenance
Build reliable, containerized execution systems that can support demanding training workloads at scale
Partner closely with research, agent, infrastructure, and platform teams to support new training use cases and harness capabilities
Design clean, stable interfaces and workflows for highly technical internal users who move quickly and expect strong ergonomics
Prevent one-off workarounds from becoming long-term technical debt by establishing durable abstractions and clear ownership
Raise the bar for correctness, reliability, operational rigor, and engineering judgment across a critical research-facing system
You might thrive in this role if you
Have significant experience building and scaling backend or infrastructure systems in fast-moving environments
Bring deep strength in API design, systems design, and engineering fundamentals
Are highly detail-oriented and care deeply about correctness, reliability, and operational quality
Can work directly with demanding technical users while maintaining strong engineering discipline
Have a track record of leading cross-functional technical efforts and creating clarity across organizational boundaries
Bring strong product sense and user empathy for internal platforms and developer tooling
Are motivated by enabling researchers and accelerating their work, rather than doing research yourself
Are proficient in Python and have experience with backend platform engineering; Rust experience is a plus
Location
This role is ideally based in San Francisco due to the close collaboration required with researchers and infrastructure partners.
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity.
We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.
For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement.
Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.
To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance.
We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link.
OpenAI Global Applicant Privacy Policy
At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.
Redirects to OpenAI's application page.
Other roles
More at OpenAI.
Software Engineer, API Safety
Senior
Data Engineer, Monetization Data Platform
Senior
Software Engineer, Plugin Developer Platform
Senior
Product Engineer, Full Stack - Agents
Senior
Engineering Manager, Artifacts
Manager