Pylon

Software Engineer, Infrastructure

Pylon2 weeks ago
Location

San Francisco

Type

Full Time

Salary

USD 180,000 – 250,000

Level

Mid

Role

Infrastructure Engineer

Posted

Jul 9, 2026

Full TimeMid

The role

Summary

Join Pylon as a Software Engineer, Infrastructure to design and maintain core cloud infrastructure for a B2B post-sales platform serving 1500+ companies including Linear and Modal Labs. You'll own end-to-end infrastructure operations, build proactive monitoring and observability systems, scale architecture for rapid growth, and optimize developer velocity through CI/CD pipelines and internal tooling. This role requires 3+ years of backend or infrastructure engineering experience with strong cloud infrastructure expertise, AI-driven development practices, and a passion for building scalable, reliable systems in a fast-paced startup environment.

What you'll do

Cloud Infrastructure Design and Ownership: Design, implement, and maintain core infrastructure across AWS cloud environments with end-to-end ownership responsibility. Move beyond reactive 'keeping the lights on' approach to proactively architect systems that support Pylon's scaling needs for 1500+ enterprise customers. Establish infrastructure standards, implement infrastructure-as-code practices, and ensure architectural decisions align with company growth trajectories.
Proactive Monitoring and Observability: Build comprehensive monitoring, alerting, and observability systems that detect and prevent incidents before they impact customers. Implement distributed tracing, structured logging, and real-time dashboards across infrastructure stack. Design alert thresholds and runbooks that enable rapid incident response and reduce mean time to recovery (MTTR) for critical systems supporting B2B post-sales workflows.
System Scalability and Performance Optimization: Identify infrastructure bottlenecks and scaling limitations before they constrain product growth. Analyze system performance under load, optimize database queries, and engineer solutions that support 10x growth scenarios. Conduct capacity planning, load testing, and performance profiling to ensure systems remain responsive as customer base and data volumes expand exponentially.
Developer Velocity and CI/CD Pipeline Development: Own the design and maintenance of CI/CD pipelines that enable all engineers to ship features and infrastructure changes rapidly. Build internal tooling including deployment automation, environment provisioning, and infrastructure self-service capabilities. Reduce deployment friction, enable blue-green deployments, and establish best practices that accelerate development cycles while maintaining reliability standards.
AI-Driven Infrastructure Automation: Architect infrastructure systems with AI-first workflows in mind, leveraging agents and automation to maximize team productivity. Build infrastructure tools and agents that automate routine operational tasks, reduce manual toil, and enable a lean team to manage complex systems at scale. Implement infrastructure-as-code automation and self-healing infrastructure patterns.
Performance Troubleshooting and Optimization: Hunt down performance bottlenecks across application and infrastructure layers using profiling tools, APM solutions, and systematic analysis. Optimize query performance, reduce latency hotspots, and improve resource utilization. Document performance optimization findings and establish best practices to prevent regressions as systems evolve.
Cross-Functional Infrastructure Collaboration: Partner with backend engineers, frontend teams, and product management to understand infrastructure requirements and technical constraints. Communicate infrastructure capabilities and limitations clearly to non-infrastructure teams. Provide technical guidance on cloud architecture patterns, security best practices, and operational excellence to all engineering teams.
Cloud Cost Management and Optimization: Monitor and optimize AWS infrastructure spending through resource utilization analysis, instance right-sizing, and cost allocation strategies. Implement cost visibility tools and establish budget accountability across teams. Balance cost optimization with performance and reliability requirements to maintain efficient infrastructure spending as company scales.

What we look for

Technical

Cloud Infrastructure ExpertiseDemonstrated hands-on experience architecting and operating systems on AWS cloud platform. Proficiency with core AWS services including EC2, RDS, S3, CloudFront, VPC, security groups, IAM, and CloudWatch. Understanding of cloud networking, load balancing, auto-scaling, and multi-region deployment patterns required.
Container Orchestration and DeploymentStrong experience with Docker containerization and Kubernetes orchestration for production workloads. Ability to design scalable deployment strategies, manage container registries, implement resource requests/limits, and troubleshoot containerized application issues. Experience with declarative infrastructure patterns and GitOps deployment methodologies preferred.
Monitoring, Observability, and LoggingHands-on experience implementing comprehensive monitoring stacks using tools such as Prometheus, Datadog, New Relic, or similar platforms. Proficiency with centralized logging solutions like ELK Stack, Splunk, or CloudWatch Logs. Understanding of distributed tracing, metrics collection, alerting strategies, and observability best practices for production systems.
CI/CD Pipeline DevelopmentProven experience building and maintaining CI/CD infrastructure using tools like GitHub Actions, GitLab CI, Jenkins, or CircleCI. Ability to design automated deployment pipelines, implement testing automation, and establish release management processes. Experience with infrastructure testing frameworks and policy-as-code tools for ensuring compliance and reliability.
Infrastructure-as-Code and Configuration ManagementProficiency with IaC tools such as Terraform, CloudFormation, or Pulumi for reproducible infrastructure provisioning. Experience with configuration management practices using Ansible, Chef, or similar tools. Ability to version control infrastructure, implement code review processes for infrastructure changes, and maintain infrastructure documentation.
Scripting and AutomationStrong proficiency in Python, Bash, or Go for writing infrastructure automation scripts and tooling. Ability to automate routine operational tasks, build custom monitoring agents, and create self-service infrastructure tools. Experience with scripting for infrastructure testing, data migration, and system administration tasks.
Database Administration and OptimizationExperience with PostgreSQL, MySQL, or similar relational databases in production environments. Understanding of database performance tuning, query optimization, backup/recovery strategies, and high-availability configurations. Familiarity with database monitoring tools and ability to diagnose and resolve database performance issues.
Security and ComplianceKnowledge of cloud security best practices including network isolation, encryption in transit and at rest, identity and access management (IAM), and vulnerability scanning. Understanding of compliance requirements for B2B SaaS applications. Experience implementing security controls and conducting security reviews of infrastructure architecture.
Incident Response and ReliabilityExperience managing production incidents, conducting root cause analysis, and implementing preventive controls. Understanding of Site Reliability Engineering (SRE) principles, error budgets, and service level objectives (SLOs). Ability to design for high availability, implement graceful degradation, and maintain system reliability under adverse conditions.

Education

Computer Science or Related FieldBachelor's degree in Computer Science, Computer Engineering, Systems Engineering, or related field. Equivalent professional experience in infrastructure engineering and demonstrated mastery of core computer science concepts can substitute for formal degree.

Experience

Backend or Infrastructure EngineeringMinimum 3+ years of professional experience in backend engineering or infrastructure/DevOps roles. Track record of designing and maintaining production systems that serve millions of requests or process significant data volumes. Experience operating infrastructure for growing startups or scale-ups that navigated rapid growth phases.
Startup Environment NavigationPrevious experience working in startup environments and thriving in ambiguous, fast-paced settings with evolving requirements. Demonstrated ability to balance shipping velocity with system reliability. Experience prioritizing work in resource-constrained environments and making architectural tradeoffs based on business priorities.
Cloud Infrastructure at ScaleHands-on experience managing cloud infrastructure serving thousands of users or millions of API requests. Experience with infrastructure scaling challenges, multi-tenant architecture considerations, and managing infrastructure complexity in production environments.
AI-Driven Development PracticesExperience leveraging AI tools and agents for software development and infrastructure automation. Demonstrated ability to integrate AI-assisted coding and infrastructure tools into development workflows to increase productivity and reduce manual toil.
Cross-Functional Problem SolvingTrack record of collaborating effectively with product, backend, and frontend teams on infrastructure requirements. Ability to translate business requirements into technical infrastructure solutions and communicate technical limitations to non-technical stakeholders.

Skills

Required skills

AWS Cloud InfrastructureDeep hands-on expertise with Amazon Web Services platform, including EC2, RDS, S3, VPC, CloudFront, CloudWatch, and AWS security services. Ability to architect and operate scalable, reliable systems on AWS.
Kubernetes and Container OrchestrationProduction-grade experience with Kubernetes clusters, including deployment strategies, scaling, networking, and troubleshooting containerized workloads at scale.
Observability and MonitoringExpertise building comprehensive monitoring, logging, and alerting infrastructure using modern observability platforms and best practices for production systems.
CI/CD Pipeline DevelopmentProven ability to design and maintain automated deployment pipelines that enable rapid, reliable releases while maintaining system stability.
Infrastructure-as-CodeProficiency with Terraform, CloudFormation, or similar IaC tools to define and version control infrastructure in reproducible, auditable ways.
Python or Bash ScriptingStrong scripting capabilities for infrastructure automation, tooling development, and operational task automation.
Database AdministrationWorking knowledge of relational databases in production, including performance tuning, backup strategies, and high-availability configuration.
Linux System AdministrationDeep familiarity with Linux operating systems, command-line tools, and system troubleshooting in production environments.
Incident Management and TroubleshootingProven ability to diagnose and resolve production incidents rapidly, implement root cause analysis, and prevent future occurrences through systematic improvements.

Nice to have

Golang Backend DevelopmentExperience writing services and tooling in Go, aligning with Pylon's backend technology stack and enabling deeper collaboration with backend engineering teams.
GraphQL API DesignFamiliarity with GraphQL architecture patterns and optimization strategies, particularly relevant for API-driven infrastructure and developer tooling at Pylon.
React Frontend TechnologyUnderstanding of React-based frontend development enables better communication with frontend teams and informed infrastructure decisions for client-heavy workloads.
Agentic Systems Production ExperienceHands-on experience building, deploying, and operating AI agent systems in production environments, directly applicable to Pylon's AI-first product roadmap.
GraphQL API Performance OptimizationExperience optimizing GraphQL queries, implementing caching strategies, and resolving N+1 query problems in production systems.
B2B SaaS InfrastructureExperience operating infrastructure for B2B SaaS platforms, understanding multi-tenant architecture patterns, enterprise compliance requirements, and customer-facing reliability standards.
Terraform at ScaleAdvanced Terraform expertise including module design, state management best practices, and managing infrastructure across multiple environments.
Datadog or New Relic ExpertiseDeep proficiency with enterprise observability platforms for comprehensive infrastructure monitoring, performance optimization, and cost management.
Security and Compliance FrameworksFamiliarity with SOC 2, HIPAA, or other B2B enterprise compliance requirements relevant to security-conscious customer bases.

Compensation & benefits

Salary

USD 180,000 – 250,000 (annual)

Stock options

Available

Benefits

Unlimited PTO

Flexible time off policy with 14 company holidays plus unlimited paid time off to support work-life balance and personal well-being.

Parental Leave

Comprehensive parental leave benefits supporting employees during significant life transitions.

Commuter Benefits

Pre-tax commuter benefit program supporting employees commuting to San Francisco headquarters office.

Fitness Stipend

Annual fitness budget supporting wellness through gym memberships, fitness classes, or personal training.

Office Meals and Snacks

Daily lunch, dinner, and snacks provided at headquarters office to support employee wellness and team connection.

Annual Company Offsite

Yearly company-wide offsite event fostering team building, strategic alignment, and culture development.

Stock Options

Equity compensation providing meaningful ownership stake in Pylon aligned with employee success and company growth.

Comprehensive Benefits Package

Full suite of health, dental, vision, and other standard employee benefits comparable to leading tech companies.


Apply for this position

You'll be redirected to the company's application page