Member of Technical Staff - Sandbox Platform
Member of Technical Staff · Senior · Full Time
Opens Prime Intellect's application page
Role
What you'll do.
Prime Intellect is seeking a highly skilled Member of Technical Staff to develop cutting-edge infrastructure and developer platforms for AI workload management. The role focuses on building distributed systems and sandbox infrastructure that enables researchers and enterprises to run advanced reinforcement learning at scale, requiring deep systems expertise and a passion for democratizing AI development.
Responsibilities
- Infrastructure Development: Design and implement distributed orchestration infrastructure using Go and Rust, focusing on high-performance networking and coordination components
- Platform Engineering: Build intuitive web interfaces and backend services for AI workload management, including REST APIs and real-time monitoring tools
- System Orchestration: Manage cloud resources, implement container orchestration, and create scheduling systems for heterogeneous hardware environments
- Automation and Deployment: Create infrastructure automation pipelines using Ansible and manage complex distributed systems with high reliability and performance
Qualifications
What we look for.
Technical
Systems Programming
Advanced experience with Rust and systems-level programming, including deep understanding of Linux kernel, networking, and performance optimization
Cloud Infrastructure
Expertise in cloud platforms (preferably GCP), Kubernetes, container orchestration, and infrastructure as code
Backend Development
Strong Python backend development skills, with experience in FastAPI, asynchronous programming, and RESTful API design
Education
Computer Science
Bachelor's or Master's degree in Computer Science, Software Engineering, or related technical field preferred
Experience
Systems Engineering
5+ years of experience in distributed systems, infrastructure development, and performance engineering
Platform Development
Proven track record of building developer tools, monitoring systems, and scalable web platforms
Skills
Required
Programming Languages
Proficiency in Rust, Python, and systems programming languages
Cloud Technologies
Advanced knowledge of Kubernetes, cloud platforms, and infrastructure automation
Web Technologies
Experience with TypeScript, React/Next.js, and backend API development
Preferred
AI/ML Infrastructure
Nice to haveUnderstanding of GPU computing, machine learning model architectures, and training infrastructure
Open Source
Nice to haveContributions to open-source infrastructure projects and community engagement
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 150,000 – 300,000
Equity·Stock options
Benefits
Equity Incentives
Significant stock options and equity compensation package
Professional Development
Budget for courses, conferences, and continuous learning opportunities
Visa Support
Full visa sponsorship and relocation assistance for qualified candidates
Team Events
Regular team off-sites and conference attendance opportunities
Process
Interview steps.
- 01
Initial Screening
Phone or video call with recruiting team to discuss background and role fit
- 02
Technical Assessment
Online coding challenge or take-home project focusing on systems programming and infrastructure skills
- 03
Technical Interviews
Multiple rounds of interviews with engineering team, covering systems design, coding, and problem-solving
- 04
Final Interview
Meeting with leadership to discuss vision, team fit, and potential contributions
Full posting
Original listing.
Building Open Superintelligence Infrastructure
Prime Intellect is building the open superintelligence stack - from frontier agentic models to the infra that enables anyone to create, train, and deploy them. We aggregate and orchestrate global compute into a single control plane and pair it with the full rl post-training stack: environments, secure sandboxes, verifiable evals, and our async RL trainer. We enable researchers, startups and enterprises to run end-to-end reinforcement learning at frontier scale, adapting models to real tools, workflows, and deployment contexts.
We recently raised $15mm in funding (total of $20mm raised) led by Founders Fund, with participation from Menlo Ventures and prominent angels including Andrej Karpathy (Eureka AI, Tesla, OpenAI), Tri Dao (Chief Scientific Officer of Together AI), Dylan Patel (SemiAnalysis), Clem Delangue (Huggingface), Emad Mostaque (Stability AI) and many others.
Role Impact
This is a hybrid role spanning both our infrastructure layers and developer platform. You'll work on two key areas:
The underlying sandbox infrastructure that powers our training systems
Our developer-facing platform for AI workload management
You will work on a distributed system with performance engineering at its core. The role will draw on the full breadth of your systems skills, from deep Linux kernel topics to high-level distributed system design. Expect your low-level systems fortitude to be pushed as you build infrastructure that remains fast, robust, and reliable at scale.
Core Technical Responsibilities
Infrastructure Development
Design and implement distributed orchestration infrastructure in Go and Rust
Build high-performance networking and coordination components
Create infrastructure automation pipelines with Ansible
Manage cloud resources and container orchestration
Implement scheduling systems for heterogeneous hardware (CPU, GPU, TPU)
Platform Development
Build intuitive web interfaces for AI workload management and monitoring
Develop REST APIs and backend services in Python
Create real-time monitoring and debugging tools
Implement user-facing features for resource management and job control
Technical Requirements
Infrastructure Skills
Systems programming experience with Rust
Strong Linux systems knowledge, including networking, namespacing, and performance tuning
Virtualization experience, including VMs, hypervisors, and low-level resource management
Infrastructure automation (Ansible, Terraform)
Container orchestration (Kubernetes)
Cloud platform expertise (GCP preferred)
Observability tools (Prometheus, Grafana)
Platform Skills
Strong Python backend development (FastAPI, async)
Modern frontend development (TypeScript, React/Next.js, Tailwind)
Experience building developer tools and dashboards
RESTful API design and implementation
Nice to Have
Experience with GPU computing and ML infrastructure
Knowledge of AI/ML model architecture and training
High-performance networking implementation
Open-source infrastructure contributions
WebSocket/real-time systems experience
What We Offer
Cash Compensation Range of $150-300k with significant equity incentives
Flexible work arrangement (San Francisco office preferred, remote possible for exceptional candidates)
Full visa sponsorship and relocation support
Professional development budget for courses and conferences
Regular team off-sites and conference attendance
Opportunity to shape the future of decentralized AI development
Growth Opportunity
You'll join a team of experienced engineers and researchers working on cutting-edge problems in AI infrastructure. We believe in open development and encourage team members to contribute to the broader AI community through research and open-source contributions.
We value potential over perfection - if you're passionate about democratizing AI development and have experience in either platform or infrastructure development (ideally both), we want to talk to you.
Ready to help shape the future of AI? Apply now and join us in our mission to make powerful AI models accessible to everyone.
Redirects to Prime Intellect's application page.
Other roles