Data Engineer, Monetization Data Platform
Data Engineer · Senior · Full Time
Opens OpenAI's application page
Role
What you'll do.
Join OpenAI's Monetization Data Platform team as a Data Engineer to design and operate large-scale data pipelines that power product usage, billing, payments, and financial data systems. This hands-on role involves building canonical data models, establishing data quality guarantees, and partnering with cross-functional teams across Product Engineering, Finance, and GTM. You'll own end-to-end systems from instrumentation through delivery, working on distributed systems at scale while maintaining a strong focus on data accuracy, observability, and operational excellence.
Responsibilities
- Design and Operate Production Data Pipelines: Design, build, and operate large-scale streaming and batch data pipelines that process product, financial, and operational data from diverse internal and external systems. Take ownership of pipeline performance, reliability, and scalability while ensuring high-throughput data processing meets business requirements.
- Develop Canonical Data Models and Products: Create reusable, canonical data models and products for monetization domains including product usage, pricing, billing, ads, payments, revenue, and general ledger accounting. Design schemas and transformations that provide clean, trustworthy data abstractions for downstream consumers.
- Establish Data Quality and Governance Standards: Implement strong guarantees for data accuracy, completeness, freshness, lineage, reconciliation, and auditability. Build automated testing frameworks, data validation pipelines, and monitoring systems to ensure financial data integrity and compliance with regulatory requirements.
- Build Developer-Focused Platform Capabilities: Create frameworks, abstractions, and platform capabilities that improve developer productivity and enable other teams to launch, measure, and iterate on monetization products using trusted data. Design self-service tools and shared infrastructure that reduce time-to-value.
- Cross-Functional Collaboration and Data Contracts: Partner with Product Engineering, Finance, Accounting, Analytics, and GTM teams to define data contracts, instrument new features, and translate product and business requirements into robust technical solutions. Facilitate communication between data consumers and engineering implementation.
- Lead Technical Design and Complex Projects: Own technical design and delivery of complex, cross-functional projects using clear system designs and RFCs to align stakeholders. Make informed tradeoffs among speed, scalability, reliability, and maintainability, balancing near-term delivery with long-term architectural health.
- Improve Observability and Operational Excellence: Enhance observability of critical data workflows through comprehensive monitoring, proactive incident response, thorough root-cause analysis, and systematic long-term remediation. Build on-call support systems and runbooks that enable reliable platform operations.
- Elevate Engineering Excellence Across Organization: Champion engineering best practices through design-before-implementation approaches, clear documentation, knowledge sharing, and mentorship. Contribute to raising technical standards and helping the broader organization adopt data engineering best practices.
Qualifications
What we look for.
Technical
Distributed Systems Architecture
Deep knowledge of distributed systems design principles, including data consistency models, fault tolerance, scalability patterns, and tradeoffs. Experience designing systems that handle high-volume data processing at scale.
Large-Scale Data Pipeline Architecture
Extensive hands-on experience building and operating production data platforms, distributed data systems, or high-scale data pipelines. Understanding of ETL/ELT patterns, stream processing architectures, and batch processing frameworks.
Programming Language Proficiency
High proficiency in at least one general-purpose programming language such as Python, Java, or Scala. Demonstrated ability to write production-grade, maintainable code with proper testing and documentation.
Data Modeling and Architecture
Strong fundamentals in dimensional modeling, fact tables, slowly changing dimensions, and data warehouse design patterns. Experience designing schemas that balance query performance, maintainability, and analytical flexibility.
Data Quality and Governance
Proven ability to design systems with rigorous data quality controls, comprehensive observability, data lineage tracking, governance frameworks, privacy considerations, and access-control requirements.
SQL and Transformation Frameworks
Advanced SQL proficiency for complex transformations and data analysis. Experience with modern data transformation frameworks and data manipulation tools used in production environments.
Education
Bachelor's Degree in Computer Science or Related Field
Bachelor's degree in Computer Science, Software Engineering, Mathematics, Statistics, or equivalent professional experience demonstrating strong foundational knowledge.
Equivalent Professional Experience
Demonstrated equivalent learning and expertise through professional experience building production data systems, even without formal degree. Strong portfolio of technical accomplishments.
Experience
Production Data Platform Operations
Minimum 5-7 years of experience building and operating production data platforms or high-scale distributed data systems in real-world environments with significant data volume and complexity.
End-to-End Data System Ownership
Proven track record of owning complete data systems from instrumentation and ingestion through modeling, quality assurance, and delivery to end consumers. Experience managing full lifecycle of data products.
Cross-Functional Collaboration
Demonstrated ability to partner effectively with non-engineering teams including Product, Finance, Accounting, Analytics, and Business stakeholders. Success translating business requirements into technical specifications.
Complex Problem Solving in Ambiguous Environments
Experience navigating undefined problems, requirements gathering, and driving consensus among diverse stakeholders. Strong judgment in making technical tradeoffs and prioritizing initiatives.
Skills
Required
Large-Scale Data Pipeline Design
Expert-level ability to architect and implement streaming and batch pipelines that handle high data volume, maintain data consistency, and scale with business growth.
Python or Scala Programming
Production-grade proficiency in Python, Scala, or Java for data engineering tasks including ETL development, data transformation, and infrastructure-as-code.
Data Modeling
Strong expertise in designing dimensional models, star schemas, fact tables, and other advanced data structures optimized for analytical queries and business intelligence.
SQL
Advanced SQL skills for complex queries, window functions, optimization, and working with large datasets across relational and columnar data warehouses.
System Design and Architecture
Ability to design complex systems with consideration for performance, reliability, maintainability, and operational characteristics. Experience documenting designs through RFCs and technical specifications.
Data Quality and Testing
Experience implementing data quality frameworks, validation pipelines, testing strategies, and monitoring systems for production data environments.
Cross-Functional Communication
Strong ability to communicate complex technical concepts to both technical engineers and non-technical stakeholders including Finance, Product, and Analytics teams.
Preferred
Modern Data Warehouse Technologies
Nice to haveExperience with lakehouse platforms, columnar data warehouses, or modern cloud data platforms such as Snowflake, BigQuery, Databricks, or similar solutions.
Workflow Orchestration
Nice to haveHands-on experience with orchestration tools such as Airflow, dbt, Dagster, or similar platforms for managing complex data workflows and dependencies.
Stream Processing Systems
Nice to haveExperience building real-time data pipelines using Apache Kafka, Apache Flink, Spark Streaming, or other event streaming and stream processing platforms.
Monetization and Financial Data
Nice to haveBackground working with monetization, pricing, product usage, billing, payments, revenue recognition, or financial data in production systems.
Financial Controls and Reconciliation
Nice to haveFamiliarity with financial close processes, reconciliation frameworks, audit requirements, GAAP principles, and regulatory compliance in data systems.
Self-Service Data Platforms
Nice to haveExperience designing and building self-service analytics platforms, shared data frameworks, or developer tooling that multiple engineering teams consume and depend on.
Data Observability and Monitoring
Nice to haveExperience implementing comprehensive monitoring, alerting, data lineage tracking, and observability solutions for data pipelines in production.
Incident Response and Reliability
Nice to haveStrong track record of on-call support, incident response, root-cause analysis, and implementing systematic long-term remediations for production data systems.
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 230,000 – 385,000
Equity·Stock options
Benefits
Equity and Stock Options
Competitive equity compensation as part of total rewards package, allowing participation in OpenAI's growth and success.
Comprehensive Health Coverage
Medical, dental, and vision insurance coverage for employees and family members with employer contributions.
Retirement Planning
401(k) retirement savings plan with employer matching contributions to support long-term financial security.
Flexible Time Off
Generous paid time off policy enabling work-life balance, wellness breaks, and time for personal priorities.
Professional Development
Learning and development opportunities including conference attendance, training programs, and skill development in cutting-edge data engineering technologies.
Parental Leave
Paid parental leave supporting employees during family expansion and childbirth.
Mental Health and Wellness
Employee assistance programs, mental health resources, wellness programs, and fitness benefits.
Remote Work Flexibility
Flexible work arrangements supporting both in-office collaboration and remote work options where applicable.
Full posting
Original listing.
About the team
The Monetization Data Platform team builds the trusted data and platform foundations that power how the company develops, measures, and improves monetization products. We bring together product usage, pricing, billing, ads, payments, and financial data to help Product, Engineering, Finance, and GTM teams make better decisions and deliver reliable customer experiences.
We work at the intersection of data engineering, product engineering, platform engineering, Finance, and GTM. Our goal is to turn complex monetization and financial data into accurate, explainable, and timely data products while building systems that scale with the growth and complexity of the business.
About the role
We are looking for a Data Engineer to improve and build the next generation of our monetization data platform. You will own high-impact systems end to end, from product instrumentation, source ingestion, and canonical modeling through quality controls, observability, and delivery to downstream consumers.
This is a hands-on role for an engineer who enjoys solving ambiguous product and data problems, designing durable architectures, and partnering closely with Product Engineering, Finance, Accounting, and GTM. You will help define technical direction, raise the engineering bar, and turn monetization opportunities into trusted, scalable data products and platform capabilities.
In this role, you will
Design, build, and operate large streaming and batch data pipelines that process product, financial, and operational data from a variety of internal and external systems.
Develop canonical data models and reusable data products for domains such as product usage, pricing, billing, ads, payments, revenue, and the general ledger.
Establish strong guarantees for data accuracy, completeness, freshness, lineage, reconciliation, and auditability.
Build frameworks and platform capabilities that improve developer productivity and make it easier for teams to launch, measure, and iterate on monetization products using trusted data.
Partner with Product Engineering, Finance, Accounting, Analytics, and GTM teams to define data contracts, instrument new monetization features, and translate product and business requirements into robust technical solutions.
Lead the technical design and delivery of complex, cross-functional projects, using clear system designs and RFCs to align partners before implementation and making sound tradeoffs among speed, scalability, reliability, and maintainability.
Improve the observability and operational excellence of critical data workflows, including monitoring, incident response, root-cause analysis, and long-term remediation.
Command strong sense of engineering excellence, contribute to a design-before-implementation approach with clear documentation, and knowledge sharing across teams to elevate the broader engineering organization.
You might thrive in this role if you
Have deep experience building and operating production data platforms, distributed data systems, or high-scale data pipelines.
Are highly proficient in large data pipeline architecture and at least one general-purpose programming language such as Python, Java, or Scala.
Have strong fundamentals in data modeling, data architecture, distributed systems, and software engineering.
Have designed systems with rigorous data quality, observability, lineage, governance, privacy, or access-control requirements.
Can collaborate with cross-functional partners to identify needs, navigate ambiguity, and drive progress from problem definition through delivery.
Bring a product-oriented mindset and communicate clearly with technical and non-technical partners, translating customer and business problems into precise data contracts and scalable system designs.
Care deeply about correctness and operational reliability while maintaining a practical bias toward delivering value.
Bring a strong sense of engineering excellence, using clear thinking, sound judgment, and a design-before-implementation approach to create maintainable systems.
Nice to have
Experience with monetization, pricing, product usage, billing, ads, payments, revenue, or financial data.
Familiarity with financial controls, reconciliation, close processes, or audit requirements.
Experience with modern lakehouse or data warehouse technologies, workflow orchestration, streaming systems, and data transformation frameworks.
Experience building self-service data platforms, shared frameworks, or developer tooling used by other data and engineering teams.
Monetization or finance domain experience is helpful but not required. We value strong data engineering judgment, systems thinking, and the ability to learn a complex domain quickly.
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity.
We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.
For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement.
Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.
To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance.
We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link.
OpenAI Global Applicant Privacy Policy
At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.
Redirects to OpenAI's application page.
Other roles
More at OpenAI.
Software Engineer, Plugin Developer Platform
Senior
Product Engineer, Full Stack - Agents
Senior
Engineering Manager, Artifacts
Manager
Principal Software Engineer, Enterprise Technology Vertical
Principal
Tech Lead Manager, Education
Lead