Senior Platform Engineer: Storage

Platform Engineer · Senior · Full Time · Remote

Global · RemoteUSD 150k – 225k12mo ago
Apply for this role

Opens Railway's application page

Role

What you'll do.

Railway is seeking a Senior Platform Engineer specialized in storage infrastructure to design and evolve high-performance, distributed storage systems. The ideal candidate will architect resilient storage solutions using cutting-edge technologies, focusing on creating scalable and efficient storage primitives that power Railway's cloud infrastructure platform.

Responsibilities

  • Storage System Design: Design and evolve production Ceph clusters, including hardware design, network requirements, configuration, tuning, and operational management
  • API Development: Create efficient, generalized APIs for live migrations of stateful workloads between hosts using systems and kernel features
  • Service Architecture: Design and build API and Orchestration services using Go, gRPC, ScyllaDB, and Temporal to connect storage primitives with higher-level platform features
  • Documentation and Planning: Write comprehensive Engineering Requirement Documents that transform ideas into defined tasks, implementation plans, and success monitoring
  • Storage Primitive Development: Build a suite of storage primitives supporting customer applications, internal services, and enabling advanced platform features like streaming image pulls

Qualifications

What we look for.

Technical

  • Distributed Systems

    Extensive experience in architecting and implementing fault-tolerant, resilient, and scalable distributed systems

  • Storage Systems

    Production experience with distributed block device systems like Ceph or deep understanding of network storage cluster design

  • Filesystem Knowledge

    Comprehensive understanding of current and next-generation filesystems including Ext4, ZFS, BTRFS, EROFS, and bcachefs

Education

  • Computer Science

    Bachelor's or Master's degree in Computer Science, Software Engineering, or related technical field preferred

Experience

  • System Design

    Proven track record of designing solutions with long-term scalability and anticipating system evolution

  • Startup Environment

    Experience working in fast-paced, high-ownership startup environments with ability to handle ambiguity

Skills

Required

  • Go Programming

    Strong proficiency in Go programming language for systems and infrastructure development

  • gRPC

    Experience designing and implementing gRPC-based microservices

  • Distributed Storage

    Deep understanding of distributed storage system design and implementation

Preferred

  • ScyllaDB

    Nice to have

    Experience with ScyllaDB for high-performance distributed database management

  • Temporal

    Nice to have

    Familiarity with Temporal for workflow orchestration

  • Advanced Filesystems

    Nice to have

    Knowledge of emerging filesystem technologies like bcachefs and EROFS

Tech stack

Languages

Go

Frameworks

gRPC

Databases

ScyllaDB

Tools

Temporal

Other

Ceph

Compensation

Pay and benefits.

Base·USD 150,000 – 225,000

Equity·Stock options

Benefits

  • Health Insurance

    Comprehensive health benefits covering employee and dependents

  • Equity Grants

    Strong equity compensation package for long-term value creation

  • Equipment Stipend

    Allowance for purchasing work-related technology and setup

  • Flexible Work

    Fully remote work arrangement with global team distribution

  • Professional Growth

    Commitment to employee development and career progression

Process

Interview steps.

  1. 01

    Initial Conversation

    Open-ended discussion about candidate's background, goals, and alignment with role

  2. 02

    Project Design Challenge

    Asynchronous design of a storage engine, with opportunity to ask clarifying questions

  3. 03

    Solution Review

    Detailed technical interview discussing project design, problem-solving approach, and technical depth

  4. 04

    Team Interview

    Meeting with four team members from different company sections to assess communication and collaboration

  5. 05

    Final Details Discussion

    Conversation with CEO to discuss offer details, onboarding, and position specifics

Full posting

Original listing.

Job description

Our core mission at Railway is to make software engineers higher leverage. We believe that people should be given powerful tools so that they can spend less time setting up to do, and more time doing.

Building the infrastructure which powers the Railway engine is the most core problem at Railway. As an infrastructure engineer working on stoarge, you will be directly responsible for designing software and hardware to back performant, high reliability block storage and object storage systems backing millions of applications. The solutions you build will be instrumental in not only scaling internal operations, but scaling the company to infinity and beyond!

“But the world would be a better place if more engineers, like me, hated technology. The stuff I design, if I'm successful, nobody will ever notice. Things will just work, and will be self-managing”

- Radia Perlman

Curious? Here are 3 blog posts that dive into exciting projects this team has worked on: 1, 2, 3

Want to learn about our work culture? Here is a three-part blog series that will help you see the unique ways our team works (Parts 1, 2, 3, and 4).

About The Role

For this role, you will:

  • Design and evolve multiple production Ceph clusters, from hardware design, to driving network requirements to configuring, tuning and operating clusters and their clients

  • Create efficient, generalizable APIs using systems/kernel features to provide safe, as-fast-as-possible live-migrations of stateful workload between hosts

  • Design and build API and Orchestration services to tie storage primitives to higher level primitives using Go, gRPC, ScyllaDB and Temporal

  • Write Engineering Requirement Documents to take something from idea, to defined tasks, to implementation, to monitoring it’s success

  • Design build a suite of storage primitives that can be used by customer applications, internal services and enable higher level platform features such as streaming image pulls or movable build caches

About You

  • Experience architecting and implementing distributed systems. You enjoy building fault tolerant, resilient, and scalable services

  • Production experience with distributed block device systems (e.g Ceph) or a solid understanding of network storage cluster design from first principles

  • Understanding and experience with current gen filesystems (Ext4, ZFS, BTRFS). Bonus points for next gen (EROFS, bcachefs)

  • A solid intuition about how long your solutions will last. All systems age. In startups, we can hope for 2-3 orders of magnitude, or 12-18mo.

  • The tact to implement your solution, creator monitors for it’s error boundaries, and document any requirements for when you’re not around

  • A great sense of direction and prioritization when it comes to dealing with the ambiguity of an early stage startup

  • A sense of grit to dive into a problem, implement a solution, scale that solution, and replace it when needed

  • A great set of communication skills for getting your point across, solution implemented, and beyond

We value and love to work with diverse persons from all backgrounds

Things to know

For better or worse, we're a startup; our team dynamics are different from companies of different sizes and stages.

  • We're distributed ALL across the globe, and that's only going to be more and more distributed. As a result, stuff is ALWAYS happening.

  • We do NOT expect you to work all the time, but you'll have to be diligent about your boundaries because the end of your day may overlap with the start of someone else's.

  • We're a small team, with high ownership, who are not only passionate about what we do, but seek to be exceptional as well. At the time of writing we're 21, serving hundreds of thousands of users. There's a lot of stuff going on, and a lot of ambiguity.

  • We want you to own it. We believe that ownership is a key to growth, and part of that growth is not only being able to make the choices, but owning the success, or failure, that comes with those choices.

Benefits and perks

At Railway, we provide best in class benefits. Great salary, full health benefits including dependents, strong equity grants, equipment stipend, and much more. For more details, check back on the main careers page.

Beyond compensation, there are a few things that we believe that make working at Railway truly unique:

  • Autonomy: We have very few meetings. Just a Monday and a Friday to go over the Company Board. We think your time is sacred, whether it's at work, or outside of work.

  • Ownership: We're a company with a high ownership, high autonomy culture. We hope that you'll come in, help us, and over the course of many years do the best work of your life. When we bring you onboard, we expect you to change the company.

  • Novel problems/solutions: We're a startup that's well funded, with cool problems, which lets us implement novel solutions! We abhor “busywork” and think, whether it's community, engineering, operations, etc there's always opportunity for creative and high leverage solutions.

  • Growth: We want you to grow with us, but we know that talent is loaned, so when you figure out what area you want to grow in next, whether it's at Railway or outside, we'll make sure you land there.

How we hire

No tricks. No surprises. Here's the entire process:

  1. Talk with us about the role

    • This is completely open ended and we're just trying to see who you are, what you want to do, and where you wanna go.

  2. Work on a small project to discuss in the interview

    • Asynchronously implement the following:

    • Pre-interview: Design a Storage Engine to power something like Railway's Volume

    • You can, and SHOULD! ask us questions ahead of time.

  3. Review your solution with the Team

    1. You'll sit down with someone on the team and go over the above. We'll poke into your solution, as well as get you acquainted with two more members of the team.

      1. Looking for: Learn about your problem solving skills. How you break down a problem and how you present a solution.

    2. Interview Structure (60 Minutes):

      1. Prework (submitted before your interview): Complete your solution

      2. 0-5m: introduction

      3. 5-50m: Building (or expanding) your solution

      4. 50-60m: Questions on Railway/Tech/etc

  4. Meet the Team

    1. You'll meet the Team, which will be comprised of 4 people from vastly different sections of the company.

      1. Looking for: How you work with the rest of the team and communicate.

  5. Offer and Details Chat with CEO

    1. Finally, we will go over the process, the role, and hammer out the details about your position, onboarding, and all the deets.

#Global

Redirects to Railway's application page.

Other roles

More at Railway.

View all 7 roles