Prime Intellect

Member of Technical Staff - Compute Platform

Prime Intellect2 weeks ago
Location

San Francisco

Type

Full Time

Salary

USD 180,000 – 280,000

Level

Mid

Role

Full Stack Engineer

Posted

Jul 8, 2026

Full TimeMid

The role

Summary

Prime Intellect seeks a Member of Technical Staff for their Compute Platform team to build the open superintelligence stack that powers frontier AI research. This hybrid infrastructure and platform engineering role requires expertise in Python backend development, Rust systems programming, and Kubernetes orchestration to create AI workload management systems, distributed training infrastructure, and high-performance networking components at scale. Ideal candidates combine full-stack development capabilities with deep infrastructure automation experience and a passion for democratizing access to frontier AI capabilities.

What you'll do

Platform Development & AI Workload Management: Design and build intuitive web interfaces for managing and monitoring AI workloads across the Lab platform. Develop REST APIs and backend services in Python using async frameworks like FastAPI to enable seamless interaction with distributed compute resources. Create real-time monitoring and debugging tools that give users visibility into resource utilization, job status, and performance metrics. Implement user-facing features for resource management, job control, and workload orchestration that abstract away infrastructure complexity.
Infrastructure Automation & Systems Engineering: Design and implement distributed training infrastructure using Rust for high-performance networking and coordination components. Develop infrastructure automation pipelines using Ansible and Terraform to manage cloud resources and enable reproducible deployments. Build scheduling systems capable of orchestrating heterogeneous hardware environments including CPUs, GPUs, and TPUs. Implement container orchestration solutions using Kubernetes to manage containerized AI workloads at scale while maintaining reliability and resource efficiency.
Backend Services & System Integration: Collaborate with the engineering team to architect and implement new backend features for the platform. Enhance REST APIs and backend services to support emerging capabilities including SFT, RL, tool use, agent workflows, and continuously improving production models. Ensure seamless integration of new features into existing infrastructure while maintaining high reliability, security standards, and performance benchmarks. Build observability into services using Prometheus and Grafana to enable proactive monitoring and debugging.
Full-Stack Feature Development: Own end-to-end feature development spanning from frontend user interfaces built with TypeScript, React/Next.js, and Tailwind CSS through backend API implementation to infrastructure deployment. Develop developer tools and interactive dashboards that enable AI teams to manage complex workflows. Design and implement RESTful APIs that support both real-time interactions via WebSockets and batch operations for frontier-scale model training and evaluation.
AI Infrastructure & Observability: Build high-performance networking and coordination components optimized for AI/ML workloads. Implement comprehensive observability solutions using Prometheus, Grafana, and custom monitoring tools to track distributed system health. Design infrastructure abstractions that simplify GPU computing, ML training pipelines, and secure sandbox environments. Contribute to evaluations and deployment infrastructure that supports continuous model improvement at production scale.

What we look for

Technical

Python Backend DevelopmentStrong proficiency in Python with hands-on experience building REST APIs and async backend services using FastAPI or similar frameworks. Ability to write clean, maintainable code that scales across distributed systems. Experience with async/await patterns, dependency injection, and API design best practices.
Rust Systems ProgrammingSolid experience with Rust for systems-level programming including networking, concurrency, and performance-critical components. Understanding of memory safety, ownership patterns, and ability to write efficient, high-performance code. Experience building or contributing to infrastructure libraries and tools.
Kubernetes & Container OrchestrationHands-on experience designing and managing Kubernetes deployments at scale. Proficiency with container concepts, pod management, resource allocation, networking, storage, and observability within Kubernetes environments. Understanding of Helm charts, operators, and GitOps practices for infrastructure-as-code.
Cloud Platform EngineeringStrong experience with cloud infrastructure providers, particularly Google Cloud Platform (GCP). Proficiency with compute instances, networking, storage services, and cost optimization. Understanding of multi-region deployments, security boundaries, and service integration patterns.
Infrastructure AutomationHands-on expertise with infrastructure-as-code tools including Ansible and Terraform. Ability to design and implement automated deployment pipelines, configuration management, and infrastructure provisioning. Experience managing infrastructure state and implementing reproducible, version-controlled infrastructure changes.
Frontend DevelopmentModern web development skills using TypeScript, React or Next.js, and Tailwind CSS for building responsive, performant dashboards and developer tools. Understanding of component architecture, state management, and real-time communication patterns including WebSockets.
Distributed Systems DesignStrong understanding of distributed system principles including consensus, fault tolerance, load balancing, and eventual consistency. Experience designing systems for high availability, scalability, and reliability. Familiarity with distributed tracing and debugging techniques.
Observability & MonitoringExperience implementing comprehensive observability solutions using Prometheus for metrics, Grafana for visualization, and understanding of logging aggregation patterns. Ability to design meaningful metrics and alerts that provide actionable insights into system behavior.

Education

Bachelor's Degree in Computer Science or Related FieldFormal education in computer science, computer engineering, or closely related discipline providing foundational knowledge in algorithms, systems design, and software engineering principles.

Experience

Backend Infrastructure Engineering4-7 years of professional experience in backend development and infrastructure engineering, with significant time spent building scalable distributed systems, cloud infrastructure, and deployment automation.
Full-Stack DevelopmentDemonstrated experience spanning both backend services and frontend development, with ability to own features end-to-end from user interface through database. Comfortable context-switching between layers of the stack.
Kubernetes & Container OperationsHands-on production experience managing Kubernetes clusters at scale, deploying containerized applications, and troubleshooting container orchestration issues in real-world environments.
AI/ML Infrastructure ExperienceBackground working with AI/ML infrastructure, including GPU computing, distributed training frameworks, model serving, or ML platform development. Understanding of AI/ML workflows and performance optimization specific to these domains.

Skills

Required skills

PythonExpert-level proficiency in Python for building scalable backend services, APIs, and infrastructure automation. Strong understanding of async programming patterns, testing frameworks, and performance optimization.
KubernetesProduction-grade Kubernetes expertise including cluster setup, workload management, networking, storage, and troubleshooting. Ability to design scalable, reliable container orchestration strategies.
GCP (Google Cloud Platform)Hands-on experience with GCP services including Compute Engine, Cloud Storage, networking, and managed services. Ability to design cloud-native architectures and optimize for cost and performance.
REST API DesignStrong understanding of RESTful principles and ability to design clean, well-documented APIs that scale. Experience with API versioning, authentication, rate limiting, and lifecycle management.
Infrastructure AutomationProficiency with Ansible, Terraform, or similar infrastructure-as-code tools. Ability to codify infrastructure changes, implement deployment automation, and manage configuration drift.
Distributed SystemsSolid understanding of distributed system concepts, fault tolerance patterns, and ability to design systems that scale horizontally. Experience debugging and monitoring distributed systems in production.
React/Next.jsModern frontend development skills using React or Next.js framework. Understanding of component-based architecture, state management, and building responsive web interfaces.

Nice to have

RustSystems programming experience with Rust for building high-performance, memory-safe infrastructure components. Understanding of ownership patterns, trait systems, and async runtime design.
GPU Computing & ML InfrastructureHands-on experience with GPU workload management, distributed training frameworks like PyTorch or TensorFlow, and ML platform development. Understanding of GPU scheduling, memory optimization, and performance profiling.
High-Performance NetworkingExperience implementing or optimizing networking layers for performance-critical systems. Understanding of protocols, latency optimization, and distributed coordination patterns.
WebSocket & Real-Time SystemsExperience building real-time systems using WebSockets or similar technologies. Understanding of bidirectional communication patterns and stream processing architectures.
Observability & MonitoringDeep experience with monitoring systems including Prometheus, Grafana, and log aggregation platforms. Ability to design metrics and alerts that provide actionable operational insights.
Open-Source ContributionsActive contributions to open-source infrastructure projects, particularly in AI/ML, distributed systems, or cloud-native domains. Demonstrated ability to collaborate in open communities and understand production constraints.
Model Training & EvaluationUnderstanding of AI/ML model architecture, training workflows including SFT and reinforcement learning, evaluation methodologies, and production deployment patterns for AI systems.
Scheduled Workload OrchestrationExperience designing scheduling systems for heterogeneous environments with mixed workload types. Understanding of resource allocation strategies, priority queuing, and fairness in distributed scheduling.

Compensation & benefits

Salary

USD 180,000 – 280,000 (annual)

Stock options

Available


Apply for this position

You'll be redirected to the company's application page