Member of Technical Staff (Software Engineer, Data Platform)

Data Engineer · Senior · Full Time

San FranciscoUSD 220k – 405k1mo ago
Apply for this role

Opens Perplexity AI's application page

Role

What you'll do.

Perplexity AI is seeking a senior-level Member of Technical Staff to join their Data Platform team, focusing on designing and operating large-scale data infrastructure that powers AI, product features, and analytics. The ideal candidate will drive architectural decisions, build self-serve data platforms, and create robust data processing systems using cutting-edge technologies.

Responsibilities

  • Data Pipeline Architecture: Design and operate large-scale batch and streaming data pipelines that power Perplexity's product features, AI workflows, analytics, and experimentation
  • Event-Driven Systems: Build and manage event-driven and streaming systems for real-time data ingestion, transformation, and delivery, alongside batch frameworks for complex computations
  • Data Orchestration: Lead data orchestration architecture using tools like Airflow or Dagster, managing scheduling, dependencies, retries, SLAs, and end-to-end observability
  • Data Platform Development: Create self-serve data platforms enabling engineers, data scientists, and analysts to discover, define, and operate data pipelines with minimal friction
  • Technical Leadership: Drive architectural decisions across storage, compute, orchestration, and data APIs, partnering closely with product engineering and data science teams
  • Team Mentorship: Mentor engineers, review designs, and elevate the technical standards of data infrastructure through collaborative feedback and documentation

Qualifications

What we look for.

Technical

  • Data Processing Technologies

    Extensive experience with batch and streaming data processing systems at scale

  • Programming Languages

    Proficiency in Python and at least one additional backend language like Go or TypeScript

  • Data Orchestration Tools

    Deep familiarity with orchestration systems such as Airflow or Dagster

Education

  • Degree Preference

    Bachelor's or Master's degree in Computer Science, Software Engineering, or related technical field preferred

Experience

  • Industry Experience

    5+ years of software engineering experience, with strong background in production data infrastructure systems

  • ML/AI Workflow Support

    Experience supporting machine learning and AI training pipelines or evaluation systems

  • Platform Ownership

    Previous ownership of internal platforms used by multiple teams

Skills

Required

  • Streaming Technologies

    Expertise in Kafka, Kinesis, or similar real-time data streaming platforms

  • Data Quality Tools

    Familiarity with data quality, lineage, observability, and governance tooling

  • Systems Architecture

    Strong systems thinking around reliability, latency, cost, and complexity tradeoffs

Preferred

  • Cloud Platforms

    Nice to have

    Experience with cloud data platforms like Databricks, Snowflake, and open-source technologies

  • Big Data Technologies

    Nice to have

    Knowledge of Spark, Flink, dbt, Iceberg, Delta Lake, and ClickHouse

Tech stack

Languages

PythonGoTypeScript

Frameworks

AirflowDagster

Databases

SnowflakeClickHouse

Tools

KafkaSpark

Other

Delta LakeIceberg

Compensation

Pay and benefits.

Base·USD 220,000 – 405,000

Benefits

  • Competitive Compensation

    Salary range of $220K to $405K with equity options

  • Cutting-Edge Technology

    Work with advanced AI and data infrastructure at an innovative startup

  • Professional Growth

    Opportunities to mentor, lead architectural decisions, and work on complex data systems

Process

Interview steps.

  1. 01

    Initial Screening

    Phone or video call with recruiting team to discuss background and role alignment

  2. 02

    Technical Interview

    In-depth technical discussion focusing on data engineering experience, system design, and architectural approaches

  3. 03

    System Design Challenge

    Architectural design exercise demonstrating candidate's ability to solve complex data infrastructure problems

  4. 04

    Team Interviews

    Meetings with potential teammates and technical leadership to assess cultural and technical fit

  5. 05

    Final Interview

    Comprehensive review with senior leadership and final decision-making

Full posting

Original listing.

About the Role

The Data Platform team owns the end-to-end data lifecycle at Perplexity, from ingestion through processing, storage, and serving, powering product features, analytics, experimentation, AI workloads, and the company’s data lake.

The team defines the architecture for batch and streaming systems, the orchestration and observability stack, and a self-serve data platform, while thoughtfully combining platforms such as Databricks and Snowflake with open-source technologies including Spark, Kafka, Flink, Airflow, Dagster, dbt, Iceberg, Delta Lake, and ClickHouse.

In this senior/staff role, you will shape architecture, set standards, and drive the long-term technical direction of Perplexity’s data ecosystem.

Key Responsibilities

  • Design and operate large-scale batch and streaming data pipelines that directly power Perplexity product features, AI training and evaluation workflows, analytics, and experimentation.

  • Build event-driven and streaming systems (Kafka, Kinesis, PubSub, or similar) for real-time ingestion, transformation, and delivery, alongside batch frameworks for backfills, aggregations, and offline computation.

  • Lead the architecture of data orchestration using tools like Airflow or Dagster, owning scheduling, dependency management, retries, SLAs, and end-to-end observability for critical data flows.

  • Set and enforce guarantees for data correctness, freshness, lineage, and recoverability, designing systems that handle rapid scale growth, partial failures, and evolving schemas without disrupting AI workloads or product experiences.

  • Build self-serve data platforms that let engineers, data scientists, and analysts safely discover data, define contracts, and create and operate their own pipelines with minimal friction.

  • Improve developer experience through better abstractions, opinionated paved paths, and standards for data modeling, testing, validation, and deployment, treating the data platform as a product used by many teams.

  • Drive architectural decisions across storage, compute, orchestration, and data APIs, partnering closely with product engineering and data science to align the data ecosystem with Perplexity’s roadmap.

  • Mentor engineers, review designs, and raise the technical bar for data infrastructure through thoughtful feedback, documentation, and hands-on collaboration.

Qualifications

  • 5+ years (Senior) or 8+ years (Staff) of software engineering experience.

  • Strong experience building production data infrastructure systems.

  • Hands-on experience with batch and/or streaming data processing at scale.

  • Deep familiarity with data orchestration systems (Airflow, Dagster, or similar).

  • Proficiency in Python and at least one additional backend language (Go, TypeScript, etc.).

  • Strong systems thinking around reliability, latency, cost, and complexity tradeoffs.

  • Experience supporting ML/AI workflows, training pipelines, or evaluation systems.

  • Familiarity with data quality, lineage, observability, and governance tooling.

  • Prior ownership of internal platforms used by many teams.

If you’re excited about this role, we encourage you to apply even if your experience doesn’t match every qualification listed above.

Redirects to Perplexity AI's application page.

Other roles

More at Perplexity AI.

View all 15 roles