Principal Software Engineer - Platform

Principal Engineer · Principal · Full Time · Remote

India · RemoteUSD 250k – 320k1w ago
Apply for this role

Opens Atlan's application page

Role

What you'll do.

Principal Software Engineer - Platform at Atlan is a high-agency infrastructure leadership role focused on architecting and scaling the foundational data plane for AI-native applications. You'll design enterprise-scale platform services, own critical shared infrastructure handling billions of assets across 100K+ users, and drive technical standards while mentoring senior engineers in a remote-first, AI-native environment.

Responsibilities

  • Design and Architect Enterprise-Scale Platform Services: Design and build production platform services including RESTful APIs, infrastructure components, runtime systems, and data ingestion frameworks operating at enterprise scale. Define API contracts and service boundaries to support billions of data assets while maintaining 99.99% availability targets and ensuring seamless integration across heterogeneous systems.
  • Build AI-Ready Context Store Architecture: Architect the context store that transforms lakehouse infrastructure into AI-ready systems with multimodal capabilities supporting structured data, unstructured content, vector embeddings, and graph data. Design systems that enable AI agents and applications to operate with semantic understanding and business context.
  • Solve Multi-Tenant Isolation and Scaling Challenges: Address complex multi-tenant isolation, data segregation, and scaling problems for enterprise SaaS deployments. Design tenant-aware infrastructure, resource allocation strategies, and isolation mechanisms that ensure data security, performance predictability, and cost attribution across thousands of enterprise customers.
  • Design Data Contracts and Integration Framework: Define data contracts governing ingestion, validation, processing, routing, storage, and serving across heterogeneous backend systems. Establish contract-driven development practices and schema-first methodologies that ensure data quality, observability, and compliance throughout the data lifecycle.
  • Own Critical Shared Infrastructure Systems: Take ownership of foundational infrastructure components including Apache Iceberg and Polaris lakehouse systems, vector store implementations, graph database backends, and OLTP systems. Drive operational excellence, reliability improvements, and performance optimization across these mission-critical platforms.
  • Drive Technical Standards and Governance: Establish and evolve technical standards across platform teams through RFC-driven processes, architecture reviews, and comprehensive documentation. Influence technology decisions, define best practices for distributed systems development, and set the architectural direction for AI-native infrastructure.
  • Mentor and Elevate Engineering Teams: Mentor senior engineers and engineering managers, providing technical guidance and architectural mentorship. Act as a force multiplier by elevating the technical bar across teams, facilitating knowledge transfer, and developing the next generation of platform engineering leaders.
  • Contribute Production Code and Debug Distributed Systems: Write production-quality code leveraging AI-assisted development tools such as Claude Code and Cursor to accelerate development velocity. Debug complex distributed systems issues across Kubernetes orchestration, workflow engines, microservices, and multi-cloud deployments to resolve production incidents and improve reliability.
  • Drive Multi-Quarter Technical Initiatives: Identify, scope, and lead multi-quarter technical initiatives from concept through production deployment and scale. Demonstrate high agency in ambiguous problem spaces, navigate fast-changing priorities in a scale-up environment, and communicate asynchronously to influence stakeholders without direct authority.

Qualifications

What we look for.

Technical

  • Distributed Systems Architecture and Design Patterns

    Deep expertise in designing and implementing distributed systems patterns including microservices architectures, service mesh implementations, event-driven systems, and asynchronous communication patterns. Proficiency with orchestration platforms, consistency models, and techniques for managing system complexity at scale.

  • Multi-Tenant SaaS Architecture and Isolation Strategies

    Proven experience architecting multi-tenant systems with strong tenant isolation, resource segregation, and data governance. Understanding of isolation patterns (logical, physical, and hybrid), tenant-aware routing, fair resource allocation, and compliance with enterprise security requirements.

  • Kubernetes and Container Orchestration

    Advanced Kubernetes expertise including cluster architecture, resource management, networking, storage orchestration, and production operations. Proficiency with containerization technologies, infrastructure-as-code practices, and cloud-native deployment patterns.

  • Cloud Infrastructure and Multi-Cloud Deployment

    Extensive experience with AWS, GCP, or Azure cloud platforms. Deep understanding of compute services, managed databases, networking, observability tools, cost optimization, and multi-region/multi-cloud deployment strategies for enterprise applications.

  • Data Platform and Lakehouse Technologies

    Strong hands-on experience with modern data infrastructure including Apache Iceberg, Delta Lake, or equivalent lakehouse systems. Knowledge of data ingestion frameworks, columnar storage optimization, metadata management, and cost-effective data governance.

  • Workflow Orchestration and Event Processing

    Experience with distributed workflow orchestration platforms, temporal reasoning systems, event streaming infrastructure, and exactly-once/at-least-once processing semantics. Ability to design systems handling asynchronous, long-running operations at scale.

Education

  • Bachelor's Degree in Computer Science or Related Field

    Formal education in Computer Science, Software Engineering, or equivalent technical discipline. Strong foundation in algorithms, data structures, systems design, and computer architecture principles.

Experience

  • 10+ Years Platform Engineering and Infrastructure

    Minimum 10 years of progressive experience in platform engineering, infrastructure, or backend systems development at scale. Track record of building foundational systems that enable product teams to move faster and operate reliably.

  • Enterprise-Scale Distributed Systems Implementation

    Hands-on experience building, deploying, and operating enterprise-scale distributed systems handling significant throughput and data volume. Demonstrated ability to solve complex operational challenges and improve reliability in production environments serving thousands of customers.

  • Multi-Quarter Technical Leadership and Delivery

    Proven track record of defining, scoping, and leading multi-quarter technical initiatives from conception through production deployment at scale. Ability to navigate ambiguity, prioritize competing demands, and deliver complex infrastructure projects that enable organizational growth.

  • SaaS and Multi-Tenant System Experience

    Substantive experience building and scaling SaaS platforms with multi-tenant architectures. Understanding of tenant isolation patterns, resource fairness, data segregation, compliance requirements, and operational challenges specific to multi-tenant deployments.

Skills

Required

  • Kubernetes and Container Orchestration

    Production-level expertise in Kubernetes architecture, cluster management, networking, storage, and operations. Ability to design and optimize containerized infrastructure for reliability and cost-efficiency.

  • Distributed Systems Design

    Deep understanding of distributed systems principles, consensus algorithms, eventual consistency models, and techniques for building reliable systems at scale. Experience debugging and optimizing distributed architectures.

  • Multi-Tenant Architecture Design

    Expertise in designing and implementing multi-tenant systems with strong isolation guarantees, fair resource allocation, and compliance with enterprise security standards. Understanding of logical and physical isolation patterns.

  • Cloud Platform Infrastructure

    Advanced proficiency with AWS, GCP, or Azure. Deep knowledge of compute, storage, networking, managed services, and cost optimization strategies for enterprise cloud deployments.

  • Data Platform Architecture

    Strong experience with data platforms, lakehouse systems, and data engineering infrastructure. Understanding of data ingestion, storage optimization, metadata management, and data governance at scale.

  • Backend Systems and API Design

    Strong proficiency in backend development, RESTful API design, gRPC, and asynchronous communication patterns. Experience building scalable, maintainable systems using programming languages like Go, Python, Java, or Rust.

  • Production Systems Reliability and Observability

    Expertise in designing reliable, observable systems including metrics collection, distributed tracing, logging strategies, and incident response procedures. Experience achieving and maintaining high availability targets (99.99%+ uptime).

  • Async Communication and Technical Leadership

    Excellent async communication skills and ability to influence without direct authority. Proven leadership experience mentoring senior engineers and driving technical decisions across distributed teams.

Preferred

  • Temporal and Workflow Orchestration Systems

    Nice to have

    Hands-on experience with Temporal, Airflow, Prefect, or similar workflow orchestration platforms. Understanding of temporal reasoning, complex workflow patterns, and distributed state management.

  • Contract-Driven and Schema-First Data Architectures

    Nice to have

    Experience designing contract-driven APIs, schema-first data platforms, and data governance frameworks. Knowledge of schema evolution strategies and data quality validation at scale.

  • Vector Databases and Graph Systems

    Nice to have

    Familiarity with vector database technologies (Pinecone, Weaviate, Qdrant), graph databases (Neo4j, GraphQL), and their integration with data platforms for AI applications.

  • Data Quality and Observability Frameworks

    Nice to have

    Experience implementing data quality frameworks, data observability platforms, and cost attribution systems for data infrastructure. Understanding of data lineage, data contracts, and quality monitoring.

  • CI/CD and GitOps Practices

    Nice to have

    Expertise in designing CI/CD pipelines, implementing GitOps workflows, and automating infrastructure deployment. Experience with tools like Flux, ArgoCD, and modern deployment strategies.

  • Enterprise Compliance and Data Governance

    Nice to have

    Experience supporting enterprise workloads with strict compliance requirements including SOC2, HIPAA, GDPR, or similar standards. Understanding of audit logging, data residency, and regulatory requirements.

  • AI-Native Development and Tooling

    Nice to have

    Familiarity with AI-assisted development tools such as Claude Code, Cursor, or GitHub Copilot. Experience leveraging LLMs and AI tools to accelerate development velocity and improve code quality.

  • Apache Iceberg and Modern Lakehouse Technologies

    Nice to have

    Hands-on experience with Apache Iceberg, Delta Lake, or Polaris for building modern data lake architectures. Understanding of lakehouse design patterns, metadata management, and ACID transactions in data systems.

Compensation

Pay and benefits.

Base·USD 250,000 – 320,000

Equity·Stock options

Benefits

  • Competitive Base Salary and Performance-Based Variable Compensation

    Market-leading base salary benchmarked at the top of the market for principal-level infrastructure engineers. Performance-based variable pay tied to individual and company milestones, ensuring rewards grow in step with value creation.

  • Impact-Driven Equity

    Significant equity stake in Atlan enabling you to share in company value creation. Equity refreshes and long-term vesting schedules designed to support long-term retention and alignment with company success.

  • Comprehensive Health and Wellness Coverage

    Comprehensive health insurance from Day 1 including medical, dental, vision, and mental health coverage. Flexible health stipends and wellness programs tailored to each country of operation, supporting physical and mental wellbeing.

  • Flexible Time Off and Modern Leave Policies

    Unlimited or generous PTO policies with trust-based time off. Modern parental leave, sabbatical options, and support for personal development. Flexible work schedules enabling work-life balance in a remote-first environment.

  • Remote-First and Global Work Environment

    Work from anywhere globally with a diverse team across 15+ countries. No artificial office requirements; work in your preferred time zone with trust-first culture respecting individual autonomy and flexibility.

  • Professional Development and Learning

    Access to cutting-edge technologies, complex infrastructure challenges, and experienced mentorship. Structured learning programs, conference attendance budgets, and opportunities for technical certification and skill development.

  • Accelerated Career Growth in Scale-Up Environment

    Opportunity to develop at uncommon velocity through exposure to complex technical problems, rapid scaling challenges, and high-impact infrastructure decisions. Direct influence on product roadmap and company technical direction.

Full posting

Original listing.

Who We Are

Most companies are racing to deploy AI, but very few have the foundation to make it work reliably. Atlan is building that missing layer: the context layer for enterprise AI. We connect the business context behind data so humans and agents can operate with far more accuracy and confidence.

With backing from world-class investors including GIC, Insight Partners, Meritech, Peak XV, and Salesforce Ventures, we've earned the trust of most AI-forward enterprises like General Motors, Nasdaq, Workday and Elastic.

Come build the infrastructure that AI runs on.

About the Role

Atlan is building the context layer for AI agents and applications - transforming how data platforms power the next generation of AI. We're looking for a Principal Engineer to help architect and scale the foundational data plane that makes it easy to build apps, agents, and solutions for the AI era.

You'll work on systems handling billions of assets, serving 100K+ users, with 99.99% availability targets. This is a high-agency role where you'll define technical direction, drive multi-quarter initiatives, and pioneer AI-native development practices.

What you will do 🤔

  • Design and build platform services - APIs, infrastructure components, runtime systems, and ingestion frameworks at enterprise scale

  • Architect the context store that transforms lakehouse infrastructure into AI-ready systems with multimodal capabilities (structured, unstructured, vector, graph)

  • Solve complex multi-tenant isolation and scaling problems for enterprise SaaS

  • Design data contracts governing ingestion, validation, processing, routing, storage, and serving across heterogeneous systems

  • Own critical shared infrastructure including lakehouse (Iceberg/Polaris), vector stores, graph databases, and OLTP systems

  • Drive technical standards through RFCs, architecture reviews, and documentation

  • Mentor senior engineers and influence architecture decisions across teams

  • Write production code using AI-assisted development tools (Claude Code, Cursor)

  • Debug distributed systems issues across Kubernetes, workflow orchestration, and microservices

What makes you a match? 😍

Must Have

  • 10+ years in platform engineering, infrastructure, or backend systems at a SaaS company

  • Experience building enterprise-scale distributed systems at scale

  • Deep expertise in multi-tenant architectures and tenant isolation strategies

  • Strong Kubernetes, containerization, and cloud infrastructure skills (AWS/GCP/Azure)

  • Hands-on experience with distributed systems patterns - service mesh, event-driven architecture, orchestration

  • Track record of driving multi-quarter technical initiatives from concept through production at scale

Strong to Have

  • Experience designing contract-driven or schema-first data platforms

  • Familiarity with Temporal or similar workflow orchestration systems

  • Data quality frameworks, observability systems, and cost attribution at scale

  • Experience supporting enterprise workloads with strict compliance requirements

  • CI/CD pipeline design and GitOps practices

Who You Are

  • You embrace AI-native development and want to pioneer new engineering workflows

  • You have high agency and take ownership of ambiguous problems

  • You're a strong async communicator who can influence without authority

  • You're comfortable with fast-changing priorities in a scale-up environment

  • You act as a force multiplier—elevating the technical bar for those around you

Why Atlan?

Joining Atlan means being part of a global movement to help data teams do their life’s best work. Here’s what you can expect:

  • Competitive Compensation: We benchmark at the top of the market and keep compensation simple: strong base salary, performance‑based variable pay, and impact‑driven equity (for most roles), so your total rewards grow in step with the value you create over time.

  • AI Native Culture: Atlan is where AI-native builders come to build the systems the future of work will run on. AI isn’t an add-on, it’s woven into how we build, think, and work every day, empowering every Atlanian to move faster and create a bigger impact.

  • Health & Wellness: From Day‑1 health, dental, vision, and mental health to flexible health stipends, we design benefits offerings that lead in each country we're in.

  • Flexible Time Off & Leave Policies: We trust you to own your energy: flexible time off and modern leave so you can unplug properly, support yourself and your loved ones, and come back ready to drive an impact.

  • Accelerated Growth & Learning: Develop at an uncommon velocity through cutting-edge tech, complex implementations, and an experienced team that values mastery.

  • Global, Remote-First, High-Trust: Work from anywhere with a diverse team across 15+ countries, in a trust-first, async environment that gives you true flexibility and ownership over how you work.

More About Us

Atlan is building the shared context layer that enterprises need so AI can operate on trusted, governed context. The conversation has moved from data leaders asking: “Can we trust the data in our stack?” to businesses asking: “Can we trust AI inside the business?”

We are the missing infrastructure for businesses becoming AI-forward - the connective tissue between their data stack, operational systems, and AI agents.

To learn more, visit www.atlan.com and follow us on LinkedIn.

Equal Opportunity Employer

Atlan is committed to building an inclusive, diverse, and authentic workplace. We do not discriminate based on race, color, religion, national origin, age, disability, sex, gender identity or expression, sexual orientation, marital status, military or veteran status, or any other legally protected characteristic.

Recruitment Fraud Alert
Atlan only posts job openings through our official Careers page at atlan.com/careers. Any other listings or communications claiming to represent Atlan may be fraudulent. We never ask for payment during hiring. Please report suspicious activity to [email protected].

Redirects to Atlan's application page.