# Staff Software Engineer - AI Runtime
**Company:** [Apollo GraphQL](https://scaleengineer.com/companies/apollo-graphql)
Staff Software Engineer - AI Runtime at Apollo GraphQL is an early-stage R&D role focused on designing and building infrastructure for AI agent discovery, understanding, and safe interaction with enterprise APIs through Constellation, Apollo's control plane. You'll own critical search, ranking, relevance, and query-planning systems while collaborating with a small, high-context team on genuinely novel problems where the roadmap evolves based on customer pilots and algorithmic innovations. This role demands hands-on production experience architecting retrieval and ranking systems at scale, comfort with ambiguity, and enthusiasm for defining problems at the intersection of AI agents and enterprise infrastructure.
**Role:** Staff Software Engineer
**Seniority:** Staff
**Locations:** United States
**Salary:** 192000–226000 USD
[Apply](https://jobs.ashbyhq.com/apollo-graphql/ee0698f5-04ab-4694-8364-efd3f3a37ec3)
Canonical: https://scaleengineer.com/jobs/apollo-graphql/staff-software-engineer-ai-runtime
---
## Responsibilities

- Design Enterprise-Scale API Discovery & Orchestration Systems: Architect and implement systems for API discovery, schema understanding, and safe agent-to-API orchestration at enterprise scale. This includes designing the core infrastructure that allows Constellation to map, catalog, and make intelligent decisions about API interactions in complex, multi-system environments.
- Own Search, Ranking, and Query-Planning Core Logic: Lead the development of search, ranking, relevance, and query-planning algorithms that determine how AI agents discover and select the right APIs and data sources. This is a central responsibility—design and iterate on the underlying algorithmic approaches rather than integrating third-party solutions.
- Rapid Prototyping & Algorithmic Evaluation: Prototype multiple algorithmic approaches to complex retrieval and ranking problems, evaluate their performance against real customer usage patterns, and iterate quickly based on pilot feedback. Priorities and direction shift week-to-week as customer insights inform architectural decisions.
- Build Agentic Systems & Observability Tooling: Develop tool-calling frameworks, agent orchestration logic, and comprehensive observability and evaluation tooling (LLM-as-judge patterns, distributed tracing, etc.) needed to ensure agents operate safely and predictably at enterprise scale.
- Collaborate on Architecture & Roadmap Direction: Work closely with a small, high-context team where every engineer influences system architecture and product direction. Communicate technical tradeoffs clearly to technical leadership, shape the roadmap based on learnings, and contribute to decision-making at every level—not just execution.
- Engage with Enterprise Customers & Iterate on Design: Collaborate with select enterprise customers piloting Constellation to understand real-world usage patterns, pain points, and requirements. Bring those insights directly back into system design, creating a tight feedback loop between customer needs and technical implementation.
- Lead Early-Stage R&D Problem Definition: Help define and refine problems at the intersection of AI agents and enterprise infrastructure where solutions aren't yet established. Thrive in ambiguity, guide architecture decisions when specifications are incomplete, and contribute to establishing best practices for the emerging field.

## Requirements

### education

- {"name":"Computer Science, Engineering, or Related Field (or Equivalent Experience)","description":"Bachelor's degree in Computer Science, Computer Engineering, or a related field. Equivalent professional experience demonstrating mastery of core computer science principles (algorithms, data structures, distributed systems) is equally valued. Apollo prioritizes demonstrated capability over credentials."}

### technical

- {"name":"Production-Grade Search, Ranking & Retrieval System Ownership","description":"Demonstrated hands-on experience designing, building, and iterating on search, ranking, relevance, or query-planning systems at real scale. This must include owning the underlying algorithmic logic and architecture—not simply integrating third-party search solutions. You should be able to discuss concrete tradeoffs between multiple algorithmic approaches (e.g., BM25 vs. dense retrieval, learning-to-rank vs. heuristic ranking, lexical vs. semantic search)."}
- {"name":"Hard Retrieval & Ranking Problem Solving","description":"Proven ability to tackle complex retrieval and ranking challenges. You've evaluated multiple algorithmic approaches, analyzed their performance characteristics, and made informed decisions about which strategies work best for specific use cases. You understand relevance metrics, ranking functions, and can debug why a system isn't performing as expected."}
- {"name":"Systems Architecture & API Design","description":"Strong foundation in designing scalable, distributed systems with clear API boundaries. Experience with query planning, federated query execution, or distributed database systems is valuable. You understand tradeoffs between consistency, availability, and latency in the context of API orchestration."}
- {"name":"Comfortable with Ambiguity & Rapid Iteration","description":"Ability to work effectively when problems aren't fully specified and solutions aren't established. You're energized by defining the right problem rather than executing against a fixed spec. You iterate quickly based on customer feedback, experiment with multiple approaches, and are comfortable with shifting priorities."}

### experience

- {"name":"5+ Years Production Software Engineering at Scale","description":"Substantial experience building and maintaining production systems that serve at scale. You've debugged complex systems in production, dealt with the realities of performance optimization, and understand the full lifecycle from design through operation and evolution."}
- {"name":"Hands-On Ownership of Search or Ranking Systems","description":"Direct experience owning search, ranking, relevance, or query-planning infrastructure. This could be at a search company, a data platform, an e-commerce company, or a systems infrastructure team—but you've been responsible for the core logic, not just the implementation of someone else's design."}
- {"name":"Experience Building or Optimizing Query Engines","description":"Background working with query optimization, distributed query planning, or federated queries. Experience with SQL or GraphQL query optimization, cost-based query planning, or federated query execution significantly strengthens your candidacy."}

## Skills

### required

- {"name":"Advanced Algorithm Design & Analysis","description":"Deep capability in designing and analyzing algorithms for search, ranking, and retrieval. You can reason about complexity, understand tradeoffs between different approaches (lexical vs. semantic, dense vs. sparse, retrieval vs. ranking), and implement solutions that balance accuracy, latency, and resource consumption."}
- {"name":"Distributed Systems & Scalable Architecture","description":"Strong foundation in designing systems that scale horizontally. You understand load balancing, partitioning, replication, consistency models, and the tradeoffs between them. You've debugged distributed systems issues and can reason about latency, throughput, and failure modes."}
- {"name":"Production System Debugging & Optimization","description":"Proven ability to identify and resolve production issues, optimize performance bottlenecks, and improve system reliability. You're comfortable with profiling, monitoring, and instrumentation—and you understand that production data often reveals truths that lab experiments don't."}
- {"name":"Technical Communication & Architectural Decision-Making","description":"Ability to communicate technical tradeoffs clearly to both engineers and leadership. You can articulate why one approach is better than another in specific contexts, make reasoned architectural decisions with incomplete information, and help others understand the reasoning behind complex decisions."}
- {"name":"Rapid Prototyping & Experimentation","description":"Comfortable quickly implementing proof-of-concept solutions to test hypotheses, iterate based on results, and move from prototype to production-ready systems. You balance speed with quality and know when 'good enough for learning' is appropriate versus when production rigor is required."}

### preferred

- {"name":"Production Rust Experience","description":"Hands-on experience building production systems in Rust. Rust's performance characteristics and safety guarantees make it particularly valuable for building performant, reliable infrastructure. However, this is a strong plus but not required—strong systems thinking and the ability to learn Rust quickly matter more."}
- {"name":"Agent Orchestration & Tool-Calling Framework Experience","description":"Direct experience building agent orchestration systems, tool-calling frameworks, or evaluating agentic systems. Familiarity with patterns like ReAct, Chain-of-Thought, function calling, and agent evaluation frameworks (LLM-as-judge) strengthens your candidacy for this role."}
- {"name":"RAG Systems & Retrieval Quality Ownership","description":"Experience building Retrieval-Augmented Generation (RAG) systems where you owned retrieval quality, reranking, and relevance metrics—not just vector store integration. Understanding the full pipeline from document ingestion through reranking and evaluation is highly valuable."}
- {"name":"Query Engine & Database Optimization Background","description":"Experience with distributed query engines, query optimization, cost-based planning, or database query execution. Understanding how systems like Presto, Trino, or traditional SQL databases plan and execute queries transfers directly to this role's challenges."}
- {"name":"GraphQL Implementation Experience","description":"Direct experience building GraphQL servers, schema design, or working with GraphQL infrastructure. Understanding Apollo's ecosystem and GraphQL's capabilities provides valuable context for how agents will interact with GraphQL APIs."}
- {"name":"Information Retrieval & Machine Learning for Ranking","description":"Background in information retrieval, learning-to-rank, or machine learning-driven ranking systems. Understanding metrics like NDCG, MRR, reciprocal rank, and precision@k; familiarity with learning-to-rank frameworks; or experience tuning ranking models strengthens your capability for this role."}

## Tech stack

### tools

- {"name":"Distributed Tracing (Jaeger, Datadog, NewRelic)","description":"Essential for understanding and debugging complex distributed systems and agentic workflows. Experience building and iterating on tracing systems helps create the observability needed to trust agent behavior."}
- {"name":"Profiling & Performance Analysis Tools","description":"Flamegraph, perf, Prometheus, and similar tools for identifying and optimizing performance bottlenecks. Deep profiling skills are essential for building high-performance ranking and retrieval systems."}
- {"name":"Kubernetes / Container Orchestration","description":"Infrastructure as Code and container orchestration for deploying and scaling distributed systems. Understanding containerization and orchestration helps with building infrastructure at scale."}
- {"name":"Version Control & Collaborative Development (Git, GitHub)","description":"Standard tooling for collaborative software development. Comfort with code review, branching strategies, and collaborative workflows is essential in a small, fast-moving team."}

### others

- {"name":"Information Retrieval Metrics & Evaluation","description":"Understanding of IR metrics (NDCG, MRR, Precision@K, Recall) and evaluation methodologies. These metrics are central to evaluating whether your ranking and retrieval systems are working effectively."}
- {"name":"Agent Evaluation & LLM-as-Judge Patterns","description":"Emerging patterns for evaluating agentic behavior using LLMs. Familiarity with how to assess agent correctness, safety, and usefulness is increasingly important in this domain."}
- {"name":"Query Planning & Cost-Based Optimization","description":"Concepts from database systems and distributed query engines. Understanding how to make intelligent decisions about which APIs to call, in what order, and with what parameters informs the query-planning challenges in this role."}

### databases

- {"name":"PostgreSQL / Vector Databases","description":"Relational and vector storage systems for API metadata, schema information, and embedding-based retrieval. Experience optimizing queries and indexing strategies across relational and vector stores is valuable."}
- {"name":"Search Infrastructure (Elasticsearch, Meilisearch, Typesense)","description":"Specialized systems for search and ranking. Hands-on experience with search infrastructure—especially owning search quality rather than just integration—is directly applicable to your responsibilities."}
- {"name":"Redis / Memcached","description":"In-memory caching and data structures for performance optimization. Experience with cache invalidation strategies and distributed caching patterns is valuable for optimizing agent interactions."}

### languages

- {"name":"Rust","description":"Primary systems programming language for building performant, memory-safe infrastructure at Constellation. Strong Rust experience is a significant advantage for contributing to the core runtime."}
- {"name":"TypeScript / JavaScript","description":"Used across Apollo's ecosystem for tooling, client libraries, and integration layers. Familiarity with modern JavaScript ecosystems is valuable for understanding how agents interact with GraphQL clients."}
- {"name":"Python","description":"Common language for AI/ML experimentation, algorithm prototyping, and data analysis—particularly valuable for evaluating algorithmic approaches before production implementation."}
- {"name":"Go","description":"Potential language for distributed systems and service components. Experience with Go's concurrency model and ecosystem tools is beneficial for certain infrastructure challenges."}

### frameworks

- {"name":"GraphQL","description":"Core to Apollo's platform. Deep understanding of GraphQL query execution, schema design, and federation is essential for understanding how Constellation will help agents interact with GraphQL APIs."}
- {"name":"LangChain / LlamaIndex","description":"Common frameworks for building agent orchestration and RAG systems. Familiarity with these ecosystems or similar frameworks helps contextualize the agentic systems you'll be building."}
- {"name":"FastAPI / REST Frameworks","description":"Used for building API services and integration layers. Experience with building scalable API services informs design decisions for agent-API interaction patterns."}

## Benefits

### benefits

- {"name":"Medical Coverage Options","description":"Choice of 3 Anthem Blue Cross medical plans for all U.S. employees. California residents have the additional flexibility to choose from 2 Kaiser medical plans, providing geographic and preference-based options for healthcare coverage."}
- {"name":"Dental & Vision Benefits","description":"Comprehensive dental and vision coverage provided by Sun Life Financial, ensuring complete preventive and corrective care without separate out-of-pocket costs for routine services."}
- {"name":"Equity Compensation","description":"Stock options as part of your compensation package, aligning your financial interests with Apollo's long-term success. This is particularly valuable in a high-growth company building the control plane for AI agents."}
- {"name":"Remote-First Work Arrangement","description":"Full flexibility to work from anywhere in the US or Canada with no office requirements. This supports work-life balance, eliminates commute time, and allows you to build your career from your preferred location."}
- {"name":"Competitive Market-Informed Compensation","description":"Apollo commits to providing competitive, market-informed salary ranges with consistency applied across the team in each country. Compensation decisions are based on your skills, experience, and interview performance—not geography within the US or Canada."}

## Compensation

- **max:** 280000
- **min:** 220000
- **currency:** USD
- **stockOptions:** true

## Interview process

### steps

## Full description
Are you excited by the challenge of building genuinely new infrastructure — where the problems aren't fully defined yet and the roadmap changes as fast as the field itself? Do you want to help figure out how AI agents discover, understand, and safely interact with real-world APIs? If so, we'd love to talk to you about joining Apollo's AI Systems team.

This team owns Constellation — Apollo's control plane for AI agents. Where GraphOS gives developers a unified way to query their data, Constellation gives AI agents a way to safely discover, understand, and coordinate across APIs at the enterprise level. This is genuinely early-stage R&D: we're piloting with select enterprise customers, iterating quickly, and figuring out the right architecture as we go rather than executing against a fixed multi-quarter plan.

**What you'll do**

* Design and build systems for API discovery, schema understanding, and safe agent-to-API orchestration at enterprise scale.
* Own search, ranking, relevance, or query-planning problems central to how agents find and use the right APIs and data — this is core to the role, not a peripheral concern.
* Prototype quickly, evaluate multiple algorithmic approaches, and iterate based on what you learn — priorities and direction can shift week to week as we learn from customer pilots.
* Build and evaluate agentic systems: tool-calling frameworks, agent orchestration, and the observability/evaluation tooling needed to trust what agents are doing.
* Collaborate closely with a small, high-context team where everyone influences architecture and direction — this is not a role with heavy process or deep management layers.
* Engage with select enterprise customers piloting Constellation to understand real usage patterns and bring that insight back into the system design.
* Communicate technical tradeoffs clearly to a small, technical leadership team and contribute to shaping the roadmap itself, not just executing against one.

**Who you are — Minimum requirements**

* You have hands-on, production experience owning a search, ranking, relevance, or query-planning system at real scale — not just integrating a third-party search tool, but designing and iterating on the underlying logic.
* You've tried multiple algorithmic approaches to a hard retrieval or ranking problem and can talk concretely about the tradeoffs.
* You're genuinely energized by ambiguity — you'd rather define the right problem than execute a well-specified one, and you're comfortable when priorities shift quickly.
* You communicate clearly and enjoy working in a small, fast-moving team with minimal process overhead.
* You have a growth mindset and want to be at the center of how AI agents and enterprise infrastructure intersect.

**Nice to have**

* Production experience with Rust — a strong plus, but not required.
* Experience building agent orchestration, tool-calling frameworks, or agent evaluation/observability systems (LLM-as-judge, tracing, etc.).
* Experience with RAG systems where you owned retrieval quality, not just integration.
* Background in distributed query engines, federated query planning, or database query optimization.

**About Apollo**

Whether you binge-watch a series on Netflix, plan faraway vacations from your phone, or read international news online, you've likely used Apollo's technology this week. Apollo supports some of the largest GraphQL platforms in the world.

We're not looking to rest on our laurels though — we're aiming to change how software is built. Having built the foundation that powers GraphQL at enterprise scale, we're now building the control plane for the next wave of software: AI agents that need to safely discover, understand, and act on real systems.

Apollo is intent on becoming the company where you can see your career grow through challenging work, collaborating with incredible teammates, and accomplishing the unattainable.

## 

**Compensation & Benefits**: At Apollo, we strive to provide competitive, market-informed compensation whilst ensuring consistency within the team in each country. We make hiring decisions based on your skills, experience, and our overall assessment of what we learned during the hiring process.  
  
In addition to the U.S. base salary range, we also provide equity and benefits. Apollo offers all U.S. employees a choice of 3 Anthem Blue Cross medical plans and California residents can also choose from an additional 2 Kaiser medical plans. Dental and Vision benefits are provided by Sun Life Financial.  
  
**Location**: This is a remote position that can be done from anywhere in the US or Canada (Canada has a different compensation range).  
  
**Equal Opportunity**: Apollo is proud to be an equal opportunity workplace dedicated to pursuing and hiring a talented and diverse workforce.  
  
**Privacy**: California residents applying for positions at Apollo can see our privacy policy [here](https://www.apollographql.com/CCPA-Privacy-Notice-for-Employees.pdf).  
  
**E-Verify**: Apollo is an E-Verify employer and will provide the federal government with your Form I-9 information to confirm that you are authorized to work in the U.S. For more information please visit [E-Verify](https://www.e-verify.gov).
