# Senior Data Engineer
**Company:** [Plaid](https://scaleengineer.com/companies/plaid)
Senior Data Engineer at Plaid, a leading fintech infrastructure company, will design and maintain scalable data systems that power data-driven decision-making across the organization. This role involves building SQL and Python-based data workflows, orchestrating complex pipelines using modern tools like DBT and Airflow, and partnering with engineering, product, and business teams to enable analytics at scale. Requires 4+ years of hands-on data engineering experience with large-scale datasets, deep expertise in cloud data warehouses, and a passion for data quality and infrastructure.
**Role:** Data Engineer
**Seniority:** Senior
**Locations:** New York City Office
**Salary:** 190800–238800 USD
[Apply](https://jobs.ashbyhq.com/plaid/eeb094c3-da73-411e-8dda-b3f4ee397f97)
Canonical: https://scaleengineer.com/jobs/plaid/senior-data-engineer
---
## Responsibilities

- Design Golden Datasets and Data Architecture: Analyze Plaid's product strategy and requirements to inform the design of golden datasets. Establish data usage principles, schema design patterns, and best practices that enable reliable analytics across the organization. Collaborate with stakeholders to understand data needs and translate them into scalable, well-documented data structures.
- Build and Maintain SQL and Python Data Pipelines: Own core data pipelines that power Plaid's data lake and data warehouse infrastructure. Write efficient, maintainable SQL queries and Python code to extract, transform, and load data at petabyte scale. Ensure pipelines handle both batch and real-time processing scenarios using technologies like Spark and Kafka.
- Establish Data Quality and SLAs: Define and enforce data quality standards, uptime guarantees, and usefulness metrics for datasets across Plaid. Implement monitoring and alerting systems to proactively detect and resolve data pipeline failures. Take ownership of previously unowned internal datasets and build comprehensive SLAs around them.
- Lead Cross-Functional Data Engineering Projects: Drive key data engineering initiatives that require collaboration across engineering, product, business intelligence, marketing, and finance teams. Partner empathetically with stakeholders to understand their data needs while balancing infrastructure requirements and business priorities.
- Optimize Data Infrastructure and Performance: Get deep into low-level data infrastructure to optimize query performance, reduce costs, and improve reliability. Continuously evaluate and implement new technologies and tools. Create proof-of-concepts that demonstrate technical advancement while considering user experience and organizational adoption.
- Advocate for Data Privacy and Integrity: Champion data privacy best practices and consumer data protection across all data systems and workflows. Act as a steward for data governance, ensuring compliance with financial regulations and internal policies. Ensure all data handling practices prioritize consumer interests and data security.

## Requirements

### education

- {"name":"Computer Science or Related Field","description":"Bachelor's degree in Computer Science, Data Science, Engineering, Mathematics, or related quantitative discipline preferred. Equivalent professional experience in data engineering or software engineering can substitute."}

### technical

- {"name":"Advanced SQL","description":"Expert-level SQL proficiency with deep understanding of query optimization, complex joins, window functions, and modern SQL patterns. Comfortable leveraging SQL as a flexible and extensible tool for sophisticated data transformations."}
- {"name":"Python Programming","description":"Strong Python programming skills for building data processing scripts, orchestrating workflows, and integrating with data infrastructure. Experience with data manipulation libraries and production-grade Python development practices."}
- {"name":"Data Orchestration Tools","description":"Proficiency with modern SQL data orchestration platforms like DBT for modeling, Airflow for workflow orchestration, and Mode or similar tools for analytics. Understanding of orchestration patterns and best practices."}
- {"name":"Cloud Data Warehouse Architecture","description":"Deep architectural knowledge of cloud-based data warehouses including Amazon Redshift, Snowflake, or Databricks. Experience with query optimization, cost management, and infrastructure scaling."}
- {"name":"Distributed Data Processing","description":"Experience with Apache Spark for large-scale data processing and transformation. Understanding of distributed computing concepts, partitioning strategies, and performance optimization."}
- {"name":"Streaming Data Platforms","description":"Experience building real-time data pipelines using Apache Kafka or similar event streaming platforms. Understanding of stream processing patterns, exactly-once semantics, and data consistency."}
- {"name":"Schema Design and Data Governance","description":"Strong appreciation for relational schema design, dimensional modeling, and star schema patterns. Experience evolving analytics schemas to handle changes and supporting complex business requirements."}
- {"name":"Data Infrastructure Management","description":"Ability to manage, deploy, and troubleshoot low-level data infrastructure. Comfortable with infrastructure-as-code, deployment automation, and infrastructure monitoring and logging."}

### experience

- {"name":"4+ Years of Data Engineering","description":"Minimum 4 years of dedicated, hands-on data engineering experience with proven expertise solving complex, large-scale data pipeline challenges in production environments."}
- {"name":"Large-Scale Data Modeling","description":"Demonstrated experience building data models and pipelines on top of massive datasets at petabyte scale (500TB+), with understanding of schema evolution on unstructured data sources."}
- {"name":"Data Warehouse and Lake Operations","description":"Hands-on experience operating and optimizing performant data warehouses and data lakes such as Redshift, Snowflake, or Databricks in production environments."}
- {"name":"Batch and Real-Time Pipeline Development","description":"Experience designing and maintaining both batch-oriented and real-time data pipelines using distributed processing frameworks like Spark and streaming platforms like Kafka."}
- {"name":"Cross-Functional Collaboration","description":"Proven ability to work effectively with diverse teams including engineers, product managers, business intelligence professionals, and business stakeholders to deliver data solutions that balance technical and business requirements."}

## Skills

### required

- {"name":"SQL","description":"Expert-level SQL for complex queries, optimization, and data transformation"}
- {"name":"Python","description":"Production-grade Python for data processing and pipeline development"}
- {"name":"DBT","description":"SQL data transformation and modeling with DBT"}
- {"name":"Apache Airflow","description":"Workflow orchestration and scheduling with Airflow"}
- {"name":"Redshift","description":"Amazon Redshift data warehouse design and optimization"}
- {"name":"Data Modeling","description":"Schema design, dimensional modeling, and analytics data architecture"}
- {"name":"Apache Spark","description":"Large-scale distributed data processing with Spark"}
- {"name":"Apache Kafka","description":"Real-time data pipeline development with Kafka"}

### preferred

- {"name":"Snowflake","description":"Experience with Snowflake data warehouse platform"}
- {"name":"Databricks","description":"Experience with Databricks platform and Delta Lake"}
- {"name":"Atlan","description":"Data governance and metadata management with Atlan"}
- {"name":"Retool","description":"Building data tools and internal applications with Retool"}
- {"name":"Mode Analytics","description":"Data exploration and visualization with Mode Analytics"}
- {"name":"Golang","description":"Understanding of Golang for integration with Plaid's backend applications"}
- {"name":"Data Privacy Frameworks","description":"Familiarity with financial data privacy regulations and compliance frameworks"}
- {"name":"Proof-of-Concept Development","description":"Experience creating technical proof-of-concepts that balance innovation with practical adoption"}

## Tech stack

### tools

- {"name":"Apache Airflow","description":"Workflow orchestration and scheduling platform"}
- {"name":"Atlan","description":"Data governance and metadata management platform"}
- {"name":"Retool","description":"Internal tool builder for data applications"}
- {"name":"Mode Analytics","description":"Data exploration, visualization, and SQL notebook platform"}

### others

- {"name":"Data Lake Architecture","description":"Design and optimization of petabyte-scale data lakes"}
- {"name":"Schema Design","description":"Dimensional modeling and analytics schema patterns"}
- {"name":"Data Quality Management","description":"SLA definition, monitoring, and data validation frameworks"}
- {"name":"Financial Data Integration","description":"Working with financial institution data from 12,000+ sources"}

### databases

- {"name":"Amazon Redshift","description":"Cloud data warehouse for analytical queries at scale"}
- {"name":"Snowflake","description":"Cloud-native data warehouse platform"}
- {"name":"Databricks","description":"Unified analytics platform with Delta Lake"}

### languages

- {"name":"SQL","description":"Primary language for data transformation, querying, and analytics"}
- {"name":"Python","description":"Data processing, pipeline orchestration, and scripting"}
- {"name":"Golang","description":"Integration context with Plaid's backend applications"}

### frameworks

- {"name":"DBT","description":"SQL-based data transformation and modeling framework"}
- {"name":"Apache Spark","description":"Large-scale distributed data processing engine"}
- {"name":"Apache Kafka","description":"Event streaming and real-time data pipeline platform"}

## Benefits

### benefits

- {"name":"Medical, Dental, and Vision Insurance","description":"Comprehensive health coverage plans for employees and eligible dependents"}
- {"name":"401(k) Retirement Plan","description":"Employer-sponsored retirement savings plan with potential company matching"}
- {"name":"Equity Compensation","description":"Stock options or equity grants as part of total compensation package"}
- {"name":"Professional Development","description":"Opportunity to learn best practices and up-level technical skills from Plaid's strong data engineering and platform teams"}
- {"name":"Collaborative Work Environment","description":"Cross-functional partnerships with teams across engineering, product, business intelligence, and business functions"}
- {"name":"Ownership and Impact","description":"High-impact role with the opportunity to carve out ownership of internal datasets, visualizations, and data strategy across Plaid"}
- {"name":"Flexible Work Arrangement","description":"Work with a company culture that values diversity, inclusion, and flexible work options"}

## Compensation

- **max:** 250000
- **min:** 180000
- **currency:** USD
- **stockOptions:** true

## Interview process

### steps

- {"name":"Initial Screening","description":"Phone or video call with a recruiter to discuss your background, experience with data engineering, and interest in the role at Plaid"}
- {"name":"Technical Screening","description":"One-on-one conversation with a senior data engineer to assess SQL proficiency, Python skills, and approach to solving data infrastructure challenges"}
- {"name":"System Design Discussion","description":"Deep-dive conversation on designing data pipelines and architectures. Discuss your experience with large-scale data systems, warehouse design, and trade-offs between different technical approaches"}
- {"name":"Cross-Functional Collaboration Panel","description":"Meetings with stakeholders from engineering, product, and business intelligence to assess collaboration style, communication skills, and ability to balance technical and business requirements"}
- {"name":"Executive or Leadership Discussion","description":"Final conversation with a senior leader on team vision, Plaid's data strategy, and how you see your role contributing to the organization's growth"}

## Full description
We believe that the way people interact with their finances will drastically improve in the next few years. We’re dedicated to empowering this transformation by building the tools and experiences that thousands of developers use to create their own products. Plaid powers the tools millions of people rely on to live a healthier financial life. We work with thousands of companies like Venmo, SoFi, several of the Fortune 500, and many of the largest banks to make it easy for people to connect their financial accounts to the apps and services they want to use. Plaid’s network covers 12,000 financial institutions across the US, Canada, UK and Europe. Founded in 2013, the company is headquartered in San Francisco with offices in New York, Seattle, Washington D.C., Raleigh, London, and Amsterdam.

Making data-driven decisions is key to Plaid's culture. To support that, we need to scale our data systems while maintaining correct and complete data. We provide golden datasets and tooling to teams across engineering, product, and business and help them explore our data quickly and safely to get the data insights they need, which ultimately helps Plaid serve our customers more effectively. In addition, Plaid will not be successful if we can't move quickly. We build the data systems and tools that enable everyone at Plaid to be data-driven, making analytics easy, obvious, and proactive across the company.

Data Engineers heavily leverage SQL and Python to build data workflows that integrate with our Golang applications. We use tools like DBT, Airflow, Redshift, Atlan, and Retool to orchestrate data pipelines and define workflows. We work with engineers, product managers, business intelligence, data analysts, and many other teams to build Plaid's data strategy and a data-first mindset.

You will be in a high impact role that will directly enable business leaders to make faster and more informed business judgements based on the datasets you build. You will have the opportunity to carve out the ownership and scope of internal datasets and visualizations across Plaid which is a currently unowned area that we intend to take over and build SLAs on. You will have the opportunity to learn best practices and up-level your technical skills from our strong DE team and from the broader Data Platform team. You will collaborate with and have strong and cross functional partnerships with literally all teams at Plaid from Engineering to Product to Marketing/Finance etc.

**Responsibilities**

* Understanding different aspects of the Plaid product and strategy to inform golden dataset choices, design and data usage principles.
* Have data quality and performance top of mind while designing datasetsLeading key data engineering projects that drive collaboration across the company.
* Advocating for adopting industry tools and practices at the right time.
* Owning core SQL and python data pipelines that power our data lake and data warehouse.
* Well-documented data with defined dataset quality, uptime, and usefulness.

**Qualifications**

* 4+ years of dedicated data engineering experience, solving complex data pipelines issues at scale.
* You’ve have experience building data models and data pipelines on top of large datasets (in the order of 500TB to petabytes)
* You value SQL as a flexible and extensible tool, and are comfortable with modern SQL data orchestration tools like DBT, Mode, and Airflow.
* You have experience working with different performant warehouses and data lakes; Redshift, Snowflake, Databricks.
* You have experience building and maintaining batch and realtime pipelines using technologies like Spark, Kafka.
* You appreciate the importance of schema design, and can evolve an analytics schema on top of unstructured data.
* You are excited to try out new technologies. You like to produce proof-of-concepts that balance technical advancement and user experience and adoption.
* You like to get deep in the weeds to manage, deploy, and improve low level data infrastructure.
* You are empathetic working with stakeholders. You listen to them, ask the right questions, and collaboratively come up with the best solutions for their needs while balancing infra and business needs.
* You are a champion for data privacy and integrity, and always act in the best interest of consumers.

Our mission at Plaid is to unlock financial freedom for everyone. To support that mission, we seek to build a diverse team of driven individuals who care deeply about making the financial ecosystem more equitable. We recognize that strong qualifications can come from both prior work experiences and lived experiences. We encourage you to apply to a role even if your experience doesn't fully match the job description. We are always looking for team members that will bring something unique to Plaid!

Plaid is proud to be an equal opportunity employer and values diversity at our company. We do not discriminate based on race, color, national origin, ethnicity, religion or religious belief, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, transgender status, sexual stereotypes, age, military or veteran status, disability, or other applicable legally protected characteristics. We also consider qualified applicants with criminal histories, consistent with applicable federal, state, and local laws. Plaid is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance with your application or interviews due to a disability, please let us know at accommodations@plaid.com.

Please review our Candidate Privacy Notice [here](https://plaid.com/legal/#candidate-privacy-notice).  

Additional compensation in the form(s) of equity and/or commission are dependent on the position offered. Plaid provides a comprehensive benefit plan, including medical, dental, vision, and 401(k). Pay is based on factors such as (but not limited to) scope and responsibilities of the position, candidate's work experience and skillset, and location. Pay and benefits are subject to change at any time, consistent with the terms of any applicable compensation or benefit plans.
