Senior Platform Engineer
Platform Engineer · Senior · Full Time · Remote
Opens Solace's application page
Role
What you'll do.
Senior Platform Engineer at Solace, a Series C healthcare startup redefining patient advocacy through technology. You will design and build scalable cloud infrastructure, implement autonomous self-healing systems, and enable product teams to deploy with velocity while maintaining reliability across a healthcare platform serving millions. This role demands deep expertise in cloud infrastructure, observability, networking, and a strong foundation in Linux systems administration alongside the ability to troubleshoot complex distributed systems under pressure.
Responsibilities
- Cloud Infrastructure Architecture and Implementation: Design, build, and maintain scalable cloud infrastructure on GCP, AWS, or Azure (with GCP preference) using Infrastructure-as-Code practices with Terraform and GitOps technologies. Implement cloud landing zones, optimize resource allocation, and establish best practices for multi-region deployments that support high-availability healthcare workloads.
- Autonomous System Resilience: Engineer self-healing processes and autonomous recovery mechanisms into platform systems to detect and remediate failures without manual intervention. Build sophisticated monitoring and alerting strategies that enable systems to gracefully degrade and recover, reducing mean time to resolution (MTTR) and minimizing patient-facing incidents.
- Observability and Monitoring: Integrate and optimize observability tools such as Datadog and OpenTelemetry across infrastructure and applications. Design comprehensive dashboards and monitoring strategies, instrument application code for performance insights, and collaborate with product engineers to establish clear observability patterns that enable data-driven debugging and optimization.
- Risk Assessment and Change Management: Evaluate the risk and operational impact of infrastructure and application changes to live production environments serving healthcare users. Establish protocols for gradual rollouts, canary deployments, and rollback procedures. Work with product and data teams to ensure safe, rapid iteration while maintaining system stability and compliance requirements.
- Production Incident Response and Troubleshooting: Participate in on-call rotation providing incident response and escalation support for production systems. Conduct root cause analysis by methodically tracing failures through architectural layers and system abstractions under time pressure. Generate hypotheses, run experiments, and implement solutions that address underlying issues rather than symptoms.
- Platform Developer Experience: Act as a force multiplier for product and data engineering teams by building tools, abstractions, and low-friction experiences that accelerate their development velocity. Design internal platforms with a product mindset, document architectural decisions, and provide mentorship on infrastructure best practices to enable teams to deploy safely and frequently.
- Networking and Security Architecture: Design and implement secure, efficient networking architectures leveraging VPCs, subnets, peering, DNS resolution, load balancing, CDNs, TLS termination, and NAT gateways. Navigate the complexities of healthcare compliance and regulated industry requirements while optimizing network performance and security posture across cloud infrastructure.
- Cross-functional Stakeholder Communication: Communicate complex infrastructure decisions, trade-offs, and technical constraints to non-technical stakeholders, product managers, and engineering leadership. Translate technical requirements into business impact and articulate infrastructure roadmap priorities that align with company growth and healthcare compliance objectives.
Qualifications
What we look for.
Technical
Cloud Platform Mastery
Deep, hands-on experience with major cloud providers (GCP, AWS, or Azure). GCP experience strongly preferred. Demonstrated proficiency with Infrastructure-as-Code tools such as Terraform, CloudFormation, or equivalent. Understanding of cloud architecture patterns including VPCs, security groups, IAM policies, compute services, and managed databases.
Linux and Command Line Fluency
Expert-level command line proficiency across Linux systems. Comfortable debugging kernel-level issues, managing system resources, writing shell scripts, and optimizing system performance. Deep understanding of Linux networking, process management, and system administration fundamentals.
Observability and Monitoring
Production experience with observability platforms such as Datadog, Prometheus, Grafana, or OpenTelemetry. Ability to design metrics, logs, and traces strategies; instrument applications for observability; build meaningful dashboards; and correlate signals across the stack to identify performance bottlenecks and failure modes.
Networking Architecture
Strong foundational knowledge of networking concepts including VPCs, subnets, routing, BGP peering, DNS resolution and propagation, load balancing (L4/L7), NAT gateways, CDNs, TLS/SSL certificates, and how these components interact at scale. Experience troubleshooting network-level connectivity and performance issues.
Distributed Systems and Troubleshooting
Hands-on experience building or operating distributed systems. Proven ability to diagnose failures by tracing through multiple abstraction layers under pressure. Experience with multi-region deployments, eventual consistency models, coordination versus coordination-free architectures, and resiliency patterns.
Kubernetes or Container Orchestration
Strong understanding of Kubernetes architecture, self-healing capabilities, extensibility mechanisms, and operational best practices. Ability to deploy, troubleshoot, and optimize containerized workloads. Experience with networking (CNI), storage provisioning, and cluster security.
Education
Computer Science or Engineering Foundation
Bachelor's degree in Computer Science, Computer Engineering, or equivalent professional experience demonstrating deep systems knowledge. Strong fundamentals in algorithms, data structures, operating systems, and networking concepts that inform architectural decision-making.
Continuous Learning Mindset
Demonstrated commitment to continuous learning and technical growth. Comfortable rapidly acquiring expertise in new technologies, frameworks, and architectural patterns. Experience mentoring others and sharing knowledge across organizations indicates leadership potential.
Experience
Startup Environment
Proven track record working in early to mid-stage startup environments where infrastructure changes have immediate business impact. Comfortable with ambiguity, moving rapidly, wearing multiple hats, and making sound architectural decisions with incomplete information. Experience scaling systems from hundreds to millions of requests.
Scalable Web and Data Infrastructure
Built and operated infrastructure supporting web applications, data platforms, or both at scale. Experience with managing databases, data pipelines, warehouses, or similar stateful systems. Demonstrated ability to optimize for performance, reliability, and cost across heterogeneous technology stacks.
Production Incident Response
Substantial on-call experience responding to and triaging production incidents affecting customer-facing systems. Comfortable with high-pressure debugging, effective communication during outages, and conducting post-incident reviews to prevent recurrence. Track record of shipping reliability improvements.
DevOps and CI/CD Systems
Experience designing, building, and optimizing continuous integration and continuous deployment pipelines. Familiarity with tools such as GitHub Actions, CircleCI, Jenkins, or similar. Understanding of deployment strategies including blue-green deployments, canary releases, and infrastructure deployment automation.
Healthcare or Regulated Industry
Desirable: Prior experience working with healthcare systems, HIPAA compliance, or other highly regulated environments. Understanding of data governance, audit logging, encryption requirements, and compliance frameworks that govern patient-facing infrastructure.
Skills
Required
Cloud Infrastructure (GCP, AWS, or Azure)
Production-level expertise with at least one major cloud provider, with strong preference for Google Cloud Platform. Proficiency with Infrastructure-as-Code, networking primitives, managed services, cost optimization, and multi-region architectures.
Linux Systems Administration
Expert command line proficiency, system debugging, kernel concepts, process management, and performance tuning on Linux operating systems. Ability to diagnose and resolve low-level system issues.
Observability Tools
Hands-on experience with Datadog, Prometheus, Grafana, or similar observability platforms. Ability to design monitoring strategies, create dashboards, and correlate signals for root cause analysis.
Networking Fundamentals
Deep understanding of VPCs, subnets, routing, DNS, load balancing, L4/L7 concepts, NAT gateways, CDNs, and TLS. Comfortable diagnosing network connectivity and performance issues.
Problem Solving and Troubleshooting
Exceptional ability to diagnose complex failures by systematically exploring abstraction layers. Generates testable hypotheses, runs experiments, and implements solutions. Comfortable working under time pressure and with incomplete information.
Communication and Documentation
Clear written and verbal communication skills for both technical and non-technical audiences. Ability to articulate architectural decisions, trade-offs, and complex technical concepts to diverse stakeholders.
Preferred
Kubernetes and Container Orchestration
Nice to haveDeep understanding of Kubernetes architecture, networking (CNI plugins), storage, RBAC, and operational best practices. Experience troubleshooting containerized workloads at scale.
Terraform and GitOps Technologies
Nice to haveProficiency with Terraform or similar Infrastructure-as-Code tools. Familiarity with GitOps workflows and tooling such as ArgoCD or Flux for managing infrastructure as code.
Polyglot Programming Comfort
Nice to haveComfortable reading, debugging, and understanding operational characteristics of applications written in multiple programming languages (Python, Go, Java, TypeScript, etc.). Not required to be expert in all, but able to trace code execution and identify performance bottlenecks.
Site Reliability Engineering Expertise
Nice to haveDemonstrated track record of improving system reliability, performance, and cost across multiple layers of the stack. Experience with SRE practices including error budgeting, SLO/SLI definitions, and blameless post-mortems.
Data Infrastructure Experience
Nice to haveExperience with data platforms including warehouse design, ETL/ELT pipelines, data connectors, replication systems, and access control patterns. Understanding of data performance optimization and governance.
Distributed Systems Architecture
Nice to haveExperience designing or operating multi-region distributed systems, understanding tradeoffs between coordination-bound and coordination-free architectures, and familiarity with consistency models and resilience patterns.
Healthcare Systems and Compliance
Nice to havePrior experience with HIPAA compliance, healthcare data governance, or other heavily regulated environments. Understanding of audit logging, encryption strategies, and compliance frameworks.
Security Architecture
Nice to haveExperience designing secure infrastructure in cloud environments with attention to identity and access management, encryption in transit and at rest, secrets management, and compliance requirements.
Compensation
Pay and benefits.
Base·USD 170,000 – 220,000
Full posting
Original listing.
About Solace
Healthcare in the U.S. is fundamentally broken. The system is so complex that 88% of U.S. adults do not have the health literacy necessary to navigate it without help. Solace cuts through the red tape of healthcare by pairing patients with expert advocates and giving them the tools to make better decisions—and get better outcomes.
We're a Series C startup, founded in 2022 and backed by Inspired Capital, Craft Ventures, Torch Capital, Menlo Ventures, Signalfire, and IVP. Our U.S. based team is lean, mission-driven, and growing quickly.
Solace isn't a place to coast. We're here to redefine healthcare—and that demands urgency, precision, and heart. If you're looking to stretch yourself, sharpen your edge, and do the best work of your life alongside a team that cares deeply, you're in the right place. We’re intense, and we like it that way.
Read more in our Bloomberg funding announcement here.
About the Role
Good plumbing is quiet and invisible, and so is good infrastructure. As a Senior Platform Engineer at Solace, you will join our platform engineering team and help build the infrastructure powering the future of health care. You enable product and data engineers to move aggressively and deploy daily, build safeguards to autonomously recover from failures, and fix things when they can't.
You are a generalist who can work both independently and as part of a team, spanning high-level architecture and low-level foundations. You are curious and always learning. You look deeper than the surface at how systems work, how they fail, and how they can self-heal. You take pride in your craft, have impeccable communication skills, and you absorb feedback exceptionally well. You aren't afraid to make mistakes and learn more from them. Above all, you enjoy taking ownership and are stifled by large organizations.
What You'll Do
Help build out our cloud infrastructure and application platform
Assess risk and impact for changes to our live environment
Join our on-call rotation
Troubleshoot problems emerging from complex interactions of technology, governance, and people
Design and build autonomous self-healing processes in our systems
Be a resource for product and data engineers in your areas of expertise. Learn and grow into other areas of expertise.
Communicate with stakeholders and internal platform customers
What You Bring to the Table
Platform engineering requires a diverse and wide-ranging set of skills. We are looking for the following core skills and experiences:
Experience working in a start-up environment
Experience building scalable infrastructure for hosting web applications or data workloads
Deep troubleshooting. You relentlessly dig through architectural and abstraction layers to identify root causes, sometimes under time pressure. You have done this solo and in a team. You have generated and tested hypotheses. You have generated and implemented solutions.
Fluent on the command line and comfortable working with Linux
Strong communication skills
Cloud infrastructure. Experience working with GCP, AWS, or Azure. GCP experience preferred. Bonus for experience working with cloud landing zones. Familiar with GitOps technologies, such as Terraform.
Networking. Understanding VPCs, subnets, routes, peering, DNS, load balancers, L4/L7 routing, NAT gateways, CDNs, TLS, and how they all work together.
Observability. Integrated and worked with tools such as Datadog and OpenTelemetry for infrastructure monitoring, application performance monitoring, and database performance. Comfortable helping instrument application code, building dashboards, and working with application engineers.
And One or More of the Following Focus Areas:
Site reliability. You have improved performance, reliability, and cost across the entire stack, from front-end to back-end.
Kubernetes. Understand how to best use its self-healing and extensibility to run workloads scalably, resiliently, and securely, and fix it when it breaks.
Polyglot. Comfortable with different languages and learning new ones. Proficient at reading code, understanding what it does, and its operational characteristics during runtime.
Devops. Built, optimized, and maintained continuous integration and continuous delivery systems.
Developer Product Experience. You built tools and low-friction experiences for product and data engineers, designing and delivering them with a product mindset.
Data Infrastructure. You moved databases. Built warehouses, connectors, and replicators. Implemented access controls, optimized and fixed pipelines.
Distributed Systems. You built both coordination-bound and coordination-free systems and know the tradeoffs in terms of scalability, reliability, and resilience. You have worked with systems spanning multi-regions. You got your hands dirty fixing systems that collapsed.
Security. Worked in a highly-regulated industry such as healthcare with cloud infrastructure
Applicants must be based in the United States.
Up for the Challenge?
We look forward to meeting you.
Fraudulent Recruitment Advisory: Solace Health will NEVER request bank details or offer employment without an interview. All legitimate communications come from official solace.health emails only or ashbyhq.com. Report suspicious activity to [email protected] or [email protected].
Redirects to Solace's application page.
Other roles