Data Engineer, Monetization Data Platform

Data Engineer · Senior · Full Time

Mountain ViewUSD 230k – 385k1d ago
Apply for this role

Opens OpenAI's application page

Role

What you'll do.

Join OpenAI's Monetization Data Platform team as a Data Engineer to design and operate large-scale data pipelines that power product usage, billing, payments, and financial data systems. This hands-on role involves building canonical data models, establishing data quality guarantees, and partnering with cross-functional teams across Product Engineering, Finance, and GTM. You'll own end-to-end systems from instrumentation through delivery, working on distributed systems at scale while maintaining a strong focus on data accuracy, observability, and operational excellence.

Responsibilities

  • Design and Operate Production Data Pipelines: Design, build, and operate large-scale streaming and batch data pipelines that process product, financial, and operational data from diverse internal and external systems. Take ownership of pipeline performance, reliability, and scalability while ensuring high-throughput data processing meets business requirements.
  • Develop Canonical Data Models and Products: Create reusable, canonical data models and products for monetization domains including product usage, pricing, billing, ads, payments, revenue, and general ledger accounting. Design schemas and transformations that provide clean, trustworthy data abstractions for downstream consumers.
  • Establish Data Quality and Governance Standards: Implement strong guarantees for data accuracy, completeness, freshness, lineage, reconciliation, and auditability. Build automated testing frameworks, data validation pipelines, and monitoring systems to ensure financial data integrity and compliance with regulatory requirements.
  • Build Developer-Focused Platform Capabilities: Create frameworks, abstractions, and platform capabilities that improve developer productivity and enable other teams to launch, measure, and iterate on monetization products using trusted data. Design self-service tools and shared infrastructure that reduce time-to-value.
  • Cross-Functional Collaboration and Data Contracts: Partner with Product Engineering, Finance, Accounting, Analytics, and GTM teams to define data contracts, instrument new features, and translate product and business requirements into robust technical solutions. Facilitate communication between data consumers and engineering implementation.
  • Lead Technical Design and Complex Projects: Own technical design and delivery of complex, cross-functional projects using clear system designs and RFCs to align stakeholders. Make informed tradeoffs among speed, scalability, reliability, and maintainability, balancing near-term delivery with long-term architectural health.
  • Improve Observability and Operational Excellence: Enhance observability of critical data workflows through comprehensive monitoring, proactive incident response, thorough root-cause analysis, and systematic long-term remediation. Build on-call support systems and runbooks that enable reliable platform operations.
  • Elevate Engineering Excellence Across Organization: Champion engineering best practices through design-before-implementation approaches, clear documentation, knowledge sharing, and mentorship. Contribute to raising technical standards and helping the broader organization adopt data engineering best practices.

Qualifications

What we look for.

Technical

  • Distributed Systems Architecture

    Deep knowledge of distributed systems design principles, including data consistency models, fault tolerance, scalability patterns, and tradeoffs. Experience designing systems that handle high-volume data processing at scale.

  • Large-Scale Data Pipeline Architecture

    Extensive hands-on experience building and operating production data platforms, distributed data systems, or high-scale data pipelines. Understanding of ETL/ELT patterns, stream processing architectures, and batch processing frameworks.

  • Programming Language Proficiency

    High proficiency in at least one general-purpose programming language such as Python, Java, or Scala. Demonstrated ability to write production-grade, maintainable code with proper testing and documentation.

  • Data Modeling and Architecture

    Strong fundamentals in dimensional modeling, fact tables, slowly changing dimensions, and data warehouse design patterns. Experience designing schemas that balance query performance, maintainability, and analytical flexibility.

  • Data Quality and Governance

    Proven ability to design systems with rigorous data quality controls, comprehensive observability, data lineage tracking, governance frameworks, privacy considerations, and access-control requirements.

  • SQL and Transformation Frameworks

    Advanced SQL proficiency for complex transformations and data analysis. Experience with modern data transformation frameworks and data manipulation tools used in production environments.

Education

  • Bachelor's Degree in Computer Science or Related Field

    Bachelor's degree in Computer Science, Software Engineering, Mathematics, Statistics, or equivalent professional experience demonstrating strong foundational knowledge.

  • Equivalent Professional Experience

    Demonstrated equivalent learning and expertise through professional experience building production data systems, even without formal degree. Strong portfolio of technical accomplishments.

Experience

  • Production Data Platform Operations

    Minimum 5-7 years of experience building and operating production data platforms or high-scale distributed data systems in real-world environments with significant data volume and complexity.

  • End-to-End Data System Ownership

    Proven track record of owning complete data systems from instrumentation and ingestion through modeling, quality assurance, and delivery to end consumers. Experience managing full lifecycle of data products.

  • Cross-Functional Collaboration

    Demonstrated ability to partner effectively with non-engineering teams including Product, Finance, Accounting, Analytics, and Business stakeholders. Success translating business requirements into technical specifications.

  • Complex Problem Solving in Ambiguous Environments

    Experience navigating undefined problems, requirements gathering, and driving consensus among diverse stakeholders. Strong judgment in making technical tradeoffs and prioritizing initiatives.

Skills

Required

  • Large-Scale Data Pipeline Design

    Expert-level ability to architect and implement streaming and batch pipelines that handle high data volume, maintain data consistency, and scale with business growth.

  • Python or Scala Programming

    Production-grade proficiency in Python, Scala, or Java for data engineering tasks including ETL development, data transformation, and infrastructure-as-code.

  • Data Modeling

    Strong expertise in designing dimensional models, star schemas, fact tables, and other advanced data structures optimized for analytical queries and business intelligence.

  • SQL

    Advanced SQL skills for complex queries, window functions, optimization, and working with large datasets across relational and columnar data warehouses.

  • System Design and Architecture

    Ability to design complex systems with consideration for performance, reliability, maintainability, and operational characteristics. Experience documenting designs through RFCs and technical specifications.

  • Data Quality and Testing

    Experience implementing data quality frameworks, validation pipelines, testing strategies, and monitoring systems for production data environments.

  • Cross-Functional Communication

    Strong ability to communicate complex technical concepts to both technical engineers and non-technical stakeholders including Finance, Product, and Analytics teams.

Preferred

  • Modern Data Warehouse Technologies

    Nice to have

    Experience with lakehouse platforms, columnar data warehouses, or modern cloud data platforms such as Snowflake, BigQuery, Databricks, or similar solutions.

  • Workflow Orchestration

    Nice to have

    Hands-on experience with orchestration tools such as Airflow, dbt, Dagster, or similar platforms for managing complex data workflows and dependencies.

  • Stream Processing Systems

    Nice to have

    Experience building real-time data pipelines using Apache Kafka, Apache Flink, Spark Streaming, or other event streaming and stream processing platforms.

  • Monetization and Financial Data

    Nice to have

    Background working with monetization, pricing, product usage, billing, payments, revenue recognition, or financial data in production systems.

  • Financial Controls and Reconciliation

    Nice to have

    Familiarity with financial close processes, reconciliation frameworks, audit requirements, GAAP principles, and regulatory compliance in data systems.

  • Self-Service Data Platforms

    Nice to have

    Experience designing and building self-service analytics platforms, shared data frameworks, or developer tooling that multiple engineering teams consume and depend on.

  • Data Observability and Monitoring

    Nice to have

    Experience implementing comprehensive monitoring, alerting, data lineage tracking, and observability solutions for data pipelines in production.

  • Incident Response and Reliability

    Nice to have

    Strong track record of on-call support, incident response, root-cause analysis, and implementing systematic long-term remediations for production data systems.

Tech stack

Languages

PythonScalaJavaSQL

Frameworks

Apache SparkApache AirflowApache Kafkadbt (Data Build Tool)Apache Flink

Databases

SnowflakeBigQueryDatabricksPostgreSQLApache Iceberg

Tools

Git and Version ControlCI/CD PipelinesDatadog or Similar MonitoringDocker and ContainerizationTerraform or Infrastructure-as-Code

Other

Data Lineage and Metadata ManagementFinancial Data ConceptsChange Data Capture (CDC)Data Privacy and Compliance

Compensation

Pay and benefits.

Base·USD 230,000 – 385,000

Equity·Stock options

Benefits

  • Equity and Stock Options

    Competitive equity compensation as part of total rewards package, allowing participation in OpenAI's growth and success.

  • Comprehensive Health Coverage

    Medical, dental, and vision insurance coverage for employees and family members with employer contributions.

  • Retirement Planning

    401(k) retirement savings plan with employer matching contributions to support long-term financial security.

  • Flexible Time Off

    Generous paid time off policy enabling work-life balance, wellness breaks, and time for personal priorities.

  • Professional Development

    Learning and development opportunities including conference attendance, training programs, and skill development in cutting-edge data engineering technologies.

  • Parental Leave

    Paid parental leave supporting employees during family expansion and childbirth.

  • Mental Health and Wellness

    Employee assistance programs, mental health resources, wellness programs, and fitness benefits.

  • Remote Work Flexibility

    Flexible work arrangements supporting both in-office collaboration and remote work options where applicable.

Full posting

Original listing.

About the team

The Monetization Data Platform team builds the trusted data and platform foundations that power how the company develops, measures, and improves monetization products. We bring together product usage, pricing, billing, ads, payments, and financial data to help Product, Engineering, Finance, and GTM teams make better decisions and deliver reliable customer experiences.

We work at the intersection of data engineering, product engineering, platform engineering, Finance, and GTM. Our goal is to turn complex monetization and financial data into accurate, explainable, and timely data products while building systems that scale with the growth and complexity of the business.

About the role

We are looking for a Data Engineer to improve and build the next generation of our monetization data platform. You will own high-impact systems end to end, from product instrumentation, source ingestion, and canonical modeling through quality controls, observability, and delivery to downstream consumers.

This is a hands-on role for an engineer who enjoys solving ambiguous product and data problems, designing durable architectures, and partnering closely with Product Engineering, Finance, Accounting, and GTM. You will help define technical direction, raise the engineering bar, and turn monetization opportunities into trusted, scalable data products and platform capabilities.

In this role, you will

  • Design, build, and operate large streaming and batch data pipelines that process product, financial, and operational data from a variety of internal and external systems.

  • Develop canonical data models and reusable data products for domains such as product usage, pricing, billing, ads, payments, revenue, and the general ledger.

  • Establish strong guarantees for data accuracy, completeness, freshness, lineage, reconciliation, and auditability.

  • Build frameworks and platform capabilities that improve developer productivity and make it easier for teams to launch, measure, and iterate on monetization products using trusted data.

  • Partner with Product Engineering, Finance, Accounting, Analytics, and GTM teams to define data contracts, instrument new monetization features, and translate product and business requirements into robust technical solutions.

  • Lead the technical design and delivery of complex, cross-functional projects, using clear system designs and RFCs to align partners before implementation and making sound tradeoffs among speed, scalability, reliability, and maintainability.

  • Improve the observability and operational excellence of critical data workflows, including monitoring, incident response, root-cause analysis, and long-term remediation.

  • Command strong sense of engineering excellence, contribute to a design-before-implementation approach with clear documentation, and knowledge sharing across teams to elevate the broader engineering organization.

You might thrive in this role if you

  • Have deep experience building and operating production data platforms, distributed data systems, or high-scale data pipelines.

  • Are highly proficient in large data pipeline architecture and at least one general-purpose programming language such as Python, Java, or Scala.

  • Have strong fundamentals in data modeling, data architecture, distributed systems, and software engineering.

  • Have designed systems with rigorous data quality, observability, lineage, governance, privacy, or access-control requirements.

  • Can collaborate with cross-functional partners to identify needs, navigate ambiguity, and drive progress from problem definition through delivery.

  • Bring a product-oriented mindset and communicate clearly with technical and non-technical partners, translating customer and business problems into precise data contracts and scalable system designs.

  • Care deeply about correctness and operational reliability while maintaining a practical bias toward delivering value.

  • Bring a strong sense of engineering excellence, using clear thinking, sound judgment, and a design-before-implementation approach to create maintainable systems.

Nice to have

  • Experience with monetization, pricing, product usage, billing, ads, payments, revenue, or financial data.

  • Familiarity with financial controls, reconciliation, close processes, or audit requirements.

  • Experience with modern lakehouse or data warehouse technologies, workflow orchestration, streaming systems, and data transformation frameworks.

  • Experience building self-service data platforms, shared frameworks, or developer tooling used by other data and engineering teams.

  • Monetization or finance domain experience is helpful but not required. We value strong data engineering judgment, systems thinking, and the ability to learn a complex domain quickly.

About OpenAI

OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. 

We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.

For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement.

Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations.

To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance.

We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link.

OpenAI Global Applicant Privacy Policy

At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.

Redirects to OpenAI's application page.

Other roles

More at OpenAI.

View all 102 roles