Senior Data Engineer
Data Engineer · Senior · Full Time · Remote
Opens Codat's application page
Role
What you'll do.
Senior Data Engineer at Codat, a JPMorgan-backed advisory intelligence platform for commercial banking, leading the Data and Insights team. You'll write production Python code daily, architect and maintain data pipelines, set technical direction for the Insights platform, and integrate AI throughout your workflows. This role combines hands-on engineering excellence with visible technical leadership across a fintech organization processing 350,000+ business financial connections.
Responsibilities
- Production Data Pipeline Development: Write production-quality Python code on a daily basis to design, build, and maintain robust data pipelines that power Codat's Insights products. Implement complex data transformations, ensure fault tolerance, and optimize pipeline performance for processing large-scale financial data connections across 350,000+ business customers.
- Full Project Lifecycle Ownership: Own data engineering projects end-to-end from problem discovery and domain understanding through pragmatic system design, feature shipping, and operational maintenance. Drive projects through all phases while maintaining production reliability and performance standards.
- Technical Direction and Platform Leadership: Set and lead the technical strategy for the Insights platform, making architectural decisions that shape the platform's future. Communicate technical vision and rationale clearly to engineering teams, product managers, and commercial stakeholders, ensuring alignment across the organization on technical priorities and trade-offs.
- Engineering Standards and Code Quality: Champion engineering excellence across the Data and Insights team by establishing and enforcing best practices including comprehensive testing strategies, robust observability, data quality validation, and maintainable code patterns. Conduct technical reviews and mentoring to raise overall team engineering standards.
- AI Integration and Innovation: Make artificial intelligence a core part of your daily workflow, identifying and implementing AI applications across data products and pipelines where they deliver measurable value. Explore use cases from research and prototyping to operational AI agents that diagnose and resolve pipeline issues automatically.
- Semantic Layer and MCP Foundation Building: Architect and implement the foundations for emerging Model Context Protocol (MCP) and semantic layer capabilities that enable both humans and AI systems to query and reason over Codat's data assets. Design data structures and interfaces that support next-generation data accessibility.
Qualifications
What we look for.
Technical
Python Programming
Advanced Python expertise writing production-ready, well-tested code with strong emphasis on maintainability, observability, and operational excellence. Demonstrated ability to build complex systems and not just script simple workflows.
Data Pipeline Architecture
Proven track record building data pipelines and production systems from first principles, designing data flows for reliability and performance. Experience with real engineering challenges such as handling edge cases, ensuring data quality, and managing pipeline orchestration.
SQL and Data Querying
Deep SQL expertise for complex data transformations, analysis, and optimization. Ability to write efficient queries, design schemas, and optimize for performance across large datasets.
Apache Spark and Distributed Computing
Solid experience with Apache Spark for distributed data processing, including DataFrame operations, performance tuning, and working with large-scale data workloads. Understanding of distributed computing principles and challenges.
Databricks and Delta Lake
Production experience with Databricks platform and Delta Lake format, including implementing data governance, managing data quality, and leveraging lakehouse architecture for analytics and ML workloads.
Data Orchestration Tools
Experience with modern orchestration platforms such as Dagster, Apache Airflow, or Temporal for scheduling, monitoring, and managing complex data workflows. Ability to design reliable orchestration systems.
dbt (Data Build Tool)
Practical experience with dbt for transforming data in the warehouse, including building data models, testing data quality, and managing dependencies in analytics codebases.
CI/CD and Deployment Practices
Expertise in modern continuous integration and continuous deployment practices, automated testing pipelines, and infrastructure-as-code patterns. Experience shaping CI/CD practices for engineering teams.
Docker and Containerization
Strong Docker and containerization expertise for building reproducible data environments, packaging applications, and enabling consistent deployment across development and production systems.
Cloud Infrastructure
Practical experience with cloud-based infrastructure platforms (AWS, GCP, or Azure) including services for data processing, storage, networking, and monitoring. Ability to design scalable cloud architectures.
AI and Machine Learning Integration
Evidence of effectively integrating AI into development workflows beyond code generation, including research applications, domain knowledge building, rapid prototyping, and operational AI systems. Demonstrated efficiency gains from AI tool usage.
Education
Bachelor's Degree in Computer Science or Related Field
Bachelor's degree in Computer Science, Software Engineering, Mathematics, Physics, or equivalent professional experience demonstrating strong foundational CS knowledge.
Alternative: Demonstrated Technical Expertise
Equivalent professional experience demonstrating mastery of core computer science fundamentals, distributed systems, and data engineering principles through substantial production work.
Experience
Senior-Level Data Engineering
5+ years of professional data engineering experience with 2+ years in senior or lead capacity. Track record of building greenfield data systems, making architectural decisions, and owning technical outcomes end-to-end.
Production System Ownership
Demonstrated experience owning production data systems from design through operation, including debugging production issues, optimizing performance, and ensuring reliability and data quality.
Technical Leadership
Evidence of technical leadership including influencing technical direction, mentoring junior engineers, establishing engineering practices, and communicating technical strategy to non-technical stakeholders.
Financial or Banking Domain
Preferred experience in fintech, banking, or financial services domains. Understanding of financial data structures, regulatory considerations, and banking system integration challenges is valuable.
Skills
Required
Python
Production-grade Python development with emphasis on code quality, testing, and maintainability
SQL
Advanced SQL skills for complex data transformations and queries
Data Pipeline Design
Architecture and implementation of robust, production-ready data pipelines
Apache Spark
Distributed data processing using Spark for large-scale analytics
Databricks/Delta Lake
Experience with lakehouse platforms and Delta Lake architecture
Orchestration (Dagster/Airflow/Temporal)
Workflow orchestration and scheduling for complex data processes
Docker/Containerization
Container-based deployment and environment management
Cloud Platforms
AWS, GCP, or Azure for data infrastructure and services
CI/CD Pipeline Management
Implementation of automated testing, deployment, and continuous integration practices
Technical Communication
Clear explanation of complex technical concepts to engineers and non-technical stakeholders
Preferred
dbt
Nice to haveData transformation tool for building modular, testable analytics code
AI/LLM Integration
Nice to havePractical experience applying AI beyond code generation, including research, prototyping, and operational AI agents
Semantic Layers and Text-to-SQL
Nice to haveFamiliarity with semantic layer concepts, ontologies, or natural language interfaces to data
Model Context Protocol (MCP)
Nice to haveKnowledge of emerging MCP standards for AI-ready data architecture
Financial Services Domain
Nice to haveExperience in fintech, banking, or financial services with understanding of financial data models and banking systems
Data Governance and Quality
Nice to haveExperience implementing data governance frameworks, data quality monitoring, and metadata management
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·GBP 90,000 – 110,000
Equity·Stock options
Benefits
Competitive Salary and Equity
Attractive compensation package commensurate with experience, including performance-based incentives and stock options as a funded fintech startup backed by JPMorgan, PayPal, Amex, Plaid, and Shopify
Flexible Work Arrangement
Hybrid or remote work options allowing flexibility in how and where you work, with consideration for team collaboration and company culture
Professional Development
Learning budget for conferences, courses, and training to develop new skills in data engineering, AI/ML, cloud platforms, and other emerging technologies
Health and Wellness
Comprehensive health insurance including medical, dental, and vision coverage, plus wellness programs and mental health support
Pension and Retirement
Employer-matched pension contributions ensuring long-term financial security
Annual Leave
Generous paid time off policy with flexible vacation days to maintain work-life balance
Technical Excellence Culture
Work within an engineering-focused organization that values code quality, technical depth, and continuous improvement
High-Impact Products
Influence the technical direction of products used by 350,000+ business customers and trusted by industry-leading financial institutions
Process
Interview steps.
- 01
Application Review
Your application and experience will be reviewed against the role requirements, with particular attention to your production data engineering background, technical depth in required tools, and evidence of technical leadership.
- 02
Initial Screening Call
A 30-minute conversation with the hiring manager or recruiter to understand your career background, motivation for joining Codat, and assess alignment with the role and team dynamics.
- 03
Technical Phone Screen
A 60-minute technical discussion covering data engineering concepts, your approach to designing production systems, specific experience with required technologies (Python, SQL, Spark, orchestration), and discussion of a recent data engineering project you've worked on.
- 04
Technical Deep Dive Interview
A 90-minute session with senior data engineers focusing on system design. You may be asked to design a data pipeline for a real-world fintech scenario, discuss architectural trade-offs, and explain your engineering approach to handling data quality, reliability, and scalability challenges.
- 05
Product and Leadership Round
A 60-minute conversation with product leadership, potentially including the VP of Engineering or Head of Data. Discussion will focus on your product mindset, ability to balance technical excellence with business outcomes, experience setting technical direction, and how you approach communicating technical strategy to non-technical stakeholders.
- 06
Behavioral and Cultural Fit
A 45-minute discussion exploring your collaboration style, approach to mentoring and raising team standards, experience working in high-growth environments, and cultural values alignment with Codat's mission-driven, engineering-focused organization.
- 07
Final Interview with Engineering Leadership
A 60-minute conversation with senior engineering leadership to discuss your vision for data platform evolution, thoughts on AI integration in data systems, and expectations for technical leadership, career growth, and working style.
- 08
Offer and Negotiation
Upon successful completion of all interview rounds, you'll receive a formal offer including salary, equity, and benefits package details. Codat is open to discussing flexible arrangements around start date, equity vesting, and other terms.
Full posting
Original listing.
About Codat
Codat is an advisory intelligence solution purpose-built for modern commercial banking. Through rich, specialized data, forward-looking insights, and integrated workflows, Codat empowers banking teams to deepen their relationships, grow their revenue, and simplify their day-to-day work.
Founded in 2017 and backed by JPMorgan, PayPal, Amex, Plaid, and Shopify, Codat has successfully powered over 350,000 connections to business customers’ financial systems — and is trusted by industry leaders to turn scattered information into actionable, strategic advantages in real time, every time.
The Role
We're looking for a Senior Data Engineer to join our Data and Insights team. You'll be hands-on every day, writing production code, building and maintaining data pipelines, and shipping features that turn raw data into intelligence our clients can act on. You'll work across the full project lifecycle, from understanding the problem through to delivery, and you'll care as much about code quality and operational reliability as you do about getting things shipped.
This is also a technical leadership role. As a senior member of the team, you'll set and lead the technical direction of our Insights platform. This is a visible position within engineering and across the wider business, so you'll explain your thinking clearly, share the reasoning behind it, and bring people with you. You'll do all of this while staying close to the code: it'll suit you if you want to keep building hands-on, rather than move into pure architecture or people management in the near term.
What You'll Do
Write production code every day, most likely in Python, building and maintaining the data pipelines that power our Insights products.
Own the full lifecycle of your projects, from understanding the data domain through to pragmatic design, shipping, and keeping things running reliably in production.
Set and lead the technical direction of the Insights platform, and communicate it openly across engineering and the wider business, so product and commercial colleagues understand the choices you are making and why.
Help raise engineering standards across the team and improve technical quality through strong engineering practice, including testing, observability, data quality checks, and clean, maintainable code.
Make AI your default way of working, and find opportunities to apply it across our products and pipelines where it delivers real value, from research and prototyping through to more operational uses such as agents that help diagnose and fix pipeline issues.
Help lay the foundations for our emerging MCP and semantic layer, so our data becomes something both people and AI systems can query and reason over.
What You'll Bring
Strong software engineering fundamentals: you write well-tested, production-ready Python and care about maintainability, observability, and operational excellence.
A track record of building data pipelines and production systems from the ground up, rather than mainly configuring managed services or wiring off-the-shelf tools together. You can describe complex logic you've written and the engineering problems you had to solve.
Solid experience with modern data engineering tools and patterns, with real depth in several of SQL, Spark, Databricks/Delta Lake, orchestration tools (Dagster, Airflow, Temporal), and dbt.
Comfort with modern deployment practices: CI/CD, containerisation (Docker), and cloud-based infrastructure. It's a bonus if you've shaped these for a team, not only worked within them.
A product mindset: you want to understand the business domain and use that understanding to shape what gets built, not only how. You're comfortable pushing back or proposing a different approach when your read of the data and the domain calls for it.
Strong communication skills: you can explain and build support for your ideas with peers, managers, and non-technical stakeholders, and you're comfortable holding a visible technical position and bringing people with you.
AI as a default part of how you work, with evidence of real efficiency gains and creative use beyond code generation, such as research, building domain knowledge, or prototyping.
Nice to have: exposure to the building blocks of AI-ready data, such as semantic layers, ontologies, or text-to-SQL, plus any experience applying AI operationally within data platforms or pipelines. This won't be your main focus, but it will help as our platform grows to support an MCP.
Redirects to Codat's application page.
Other roles