Software Engineer, Model Evaluation and Improvement

Machine Learning Engineer · Mid · Full Time

San Francisco, CAUSD 136k – 167k5d ago
Apply for this role

Opens Benchling's application page

Role

What you'll do.

Software Engineer at Benchling focused on evaluating and improving frontier AI models for scientific applications. You'll build datasets, evaluation systems, and data infrastructure that help large language models better reason about complex biological problems. This role bridges software engineering, biology, and frontier AI, requiring 2+ years of experience at the intersection of biology and AI with proven expertise in working with LLMs and building scalable systems for scientific data processing.

Responsibilities

  • Build Evaluation Datasets for Frontier Models: Develop high-quality datasets and benchmarks that rigorously evaluate frontier large language models on scientific reasoning tasks. Transform complex biological data, scientific workflows, and domain expertise into structured evaluation tasks that accurately measure model capabilities and identify performance gaps in scientific domains.
  • Analyze Model Failure Modes: Design and execute systematic experiments across frontier AI models including GPT-4, Claude, and emerging LLMs to understand failure modes and reasoning limitations in biological and scientific contexts. Use empirical analysis to identify specific weaknesses and generate actionable insights for model improvement and refinement.
  • Build Scalable Data Infrastructure: Engineer robust data pipelines and infrastructure systems that curate, transform, validate, and organize large volumes of scientific data into structured tasks for model evaluation. Implement automation frameworks that handle data quality assurance and enable continuous generation of new evaluation benchmarks at scale.
  • Collaborate with Frontier AI Labs: Partner with leading AI research organizations and model developers to understand emerging capabilities and limitations. Contribute to the development and validation of novel approaches for improving model reasoning on challenging scientific tasks, providing real-world biotech use cases and feedback.
  • Translate Expert Scientific Judgment into Evaluation Criteria: Work directly with domain-expert scientists to convert tacit domain knowledge into explicit, measurable evaluation criteria. Design evaluation frameworks that capture nuanced scientific judgment and enable reliable discrimination between strong and weak model behavior on real-world biotech problems.

Qualifications

What we look for.

Technical

  • Large Language Model Development Experience

    Demonstrated expertise in working with frontier LLMs including prompt engineering, fine-tuning, evaluation frameworks, and understanding model capabilities and limitations. Deep familiarity with current-generation models (GPT-4, Claude, Llama, or equivalent) and ability to design systems that effectively leverage their capabilities while working around their constraints.

  • Scientific Data Management

    Experience designing, building, and maintaining systems for scientific data processing, validation, and curation. Proficiency with structured data formats common in biology and biotech (genomic data, protein sequences, experimental protocols) and ability to transform complex scientific information into machine-readable formats.

  • Data Pipeline and ETL Engineering

    Strong experience building scalable data pipelines, data validation frameworks, and ETL processes. Comfortable with distributed data processing, handling large-scale datasets, and implementing robust error handling and data quality assurance mechanisms.

  • Python Programming for ML/Data Applications

    Advanced Python proficiency for machine learning and data science applications, including working with popular ML libraries (PyTorch, TensorFlow, scikit-learn, Hugging Face). Ability to write clean, maintainable, production-quality code for data processing and model evaluation workflows.

  • Evaluation Framework Design

    Experience designing comprehensive evaluation frameworks and benchmarks for machine learning systems. Understanding of evaluation metrics, statistical significance testing, and ability to create metrics that accurately capture model performance on complex reasoning tasks.

Education

  • Bachelor's Degree in Computer Science, Engineering, or Quantitative Field

    Formal technical education providing strong foundations in algorithms, data structures, and software engineering principles. Alternative: equivalent practical experience demonstrating mastery of these concepts through professional work.

  • Biology, Bioinformatics, or Life Sciences Background (Preferred)

    Academic or professional background in biological sciences, bioinformatics, computational biology, or related fields providing deep understanding of biological concepts, scientific methodology, and domain-specific data formats. Significantly strengthens ability to translate between scientists and AI systems.

Experience

  • 2+ Years at Biology-AI Intersection

    Demonstrated professional experience working at the intersection of biology and artificial intelligence, including evaluating and improving scientific models or LLMs for biological applications. This could include roles in biotech, computational biology, AI research applied to life sciences, or related domains.

  • LLM System Design and Development

    Hands-on experience building production systems with large language models, including prompt engineering, evaluation, fine-tuning, or RAG (Retrieval-Augmented Generation) implementations. Deep intuition for where current models excel, where they struggle, and how to architect systems around their capabilities.

  • Ambiguous Problem-Solving in Emerging Technology

    Proven track record of thriving in early-stage or rapidly evolving technical domains where specifications are unclear and technical approaches are still being discovered. Comfort with shifting priorities, high experimentation velocity, and iterative development in frontier technology areas.

Skills

Required

  • Python

    Advanced Python programming for data processing, machine learning, and building evaluation systems. Essential for implementing data pipelines and LLM evaluation frameworks.

  • LLM Evaluation and Benchmarking

    Expertise in designing and implementing evaluation frameworks for frontier models, including creating benchmarks that test reasoning capabilities on scientific tasks. Understanding of evaluation metrics, statistical methods, and best practices for rigorous model assessment.

  • Biological Domain Knowledge

    Strong understanding of molecular biology, genomics, biochemistry, or wet lab processes. Ability to understand complex scientific problems and translate them into tasks that test AI model reasoning on real biotech challenges.

  • Data Pipeline Engineering

    Experience designing and implementing ETL pipelines, data validation frameworks, and scalable data infrastructure. Proficiency with tools for data transformation, validation, and quality assurance.

  • Prompt Engineering and LLM APIs

    Hands-on experience with modern LLM APIs and prompt engineering techniques. Understanding of how to effectively interact with frontier models, design structured prompts, and extract reliable outputs for evaluation and analysis.

Preferred

  • SQL and Database Design

    Nice to have

    Experience with SQL query optimization, database design, and working with large-scale databases for scientific data management. Valuable for managing and querying structured evaluation datasets.

  • Scientific Computing Libraries

    Nice to have

    Proficiency with scientific Python libraries such as NumPy, Pandas, SciPy, and specialized bioinformatics packages. Useful for working with biological data and performing complex scientific computations.

  • Machine Learning and Fine-tuning

    Nice to have

    Experience with model fine-tuning, transfer learning, or training custom models on domain-specific tasks. Understanding of how to adapt pre-trained models for specialized applications.

  • Biotech Domain Experience

    Nice to have

    Prior experience in biotech, pharmaceutical R&D, academic research labs, or related life sciences environments. Familiarity with scientific workflows, wet lab processes, and how AI can address real scientist pain points.

  • Software Engineering Best Practices

    Nice to have

    Version control (Git), testing frameworks, code review processes, and CI/CD pipelines. Experience working on collaborative engineering teams building production systems.

  • Research Publication and Communication

    Nice to have

    Experience communicating technical findings to diverse audiences including scientists, engineers, and leadership. Ability to write clear technical documentation and potentially contribute to research publications.

Tech stack

Languages

Python

Frameworks

Hugging Face TransformersPyTorchLangChain

Databases

PostgreSQLVector Databases (Pinecone/Weaviate)

Tools

Jupyter NotebooksGit and GitHubDockerApache Airflow or DagsterWeights and Biases

Other

LLM APIs (OpenAI, Anthropic, open-source models)Scientific Data FormatsEvaluation Metrics and Statistical Methods

Compensation

Pay and benefits.

Base·USD 136,435 – 166,754

Equity·Stock options

Benefits

  • Equity Compensation

    Stock options as part of comprehensive compensation package, providing upside participation in Benchling's growth as a Series D-funded company at the forefront of AI-powered biotech platforms.

  • Comprehensive Health Insurance

    Medical, dental, and vision coverage with company contributions toward employee and family plans, meeting or exceeding industry standards.

  • Retirement Planning

    401(k) plan with company matching to support long-term financial planning and wealth building.

  • Professional Development

    Learning budget and support for attending conferences, courses, and pursuing certifications in AI, machine learning, and domain expertise areas.

  • Flexible Time Off

    Generous vacation and flexible time-off policies to support work-life balance and employee wellbeing.

  • Modern Office Environment

    Collaborative workspace in San Francisco designed for rapid experimentation, with access to cutting-edge tools and infrastructure for AI and data engineering work.

  • Parental Leave

    Paid parental leave benefits supporting employees through major life transitions.

  • AI and Technology Resources

    Access to frontier AI tools, computational resources, and partnerships with leading AI labs and model providers for conducting research and experiments.

Process

Interview steps.

  1. 01

    AI-Focused Assessment

    During the interview process, candidates will complete a focused exercise or discussion exploring how they think about and leverage AI to drive impact in their work. Bring examples of AI tools, platforms, or workflows you currently use. This reflects Benchling's core commitment to AI fluency and helps assess practical experience with frontier models.

  2. 02

    Technical Screening

    Conversation focused on your experience building with large language models, designing evaluation frameworks, and working with scientific data. Expect to discuss specific projects where you evaluated model performance or designed systems around LLM capabilities.

  3. 03

    Scientific Domain Depth Interview

    Discussion of your biological knowledge and experience translating scientific problems into AI tasks. This may involve walking through how you would design an evaluation for a specific scientific reasoning challenge.

  4. 04

    System Design and Infrastructure Discussion

    Technical conversation about building scalable data pipelines and evaluation infrastructure. You may be asked to discuss architectural decisions for handling large-scale scientific datasets or designing evaluation systems.

  5. 05

    Collaboration and Problem-Solving Round

    Conversation with team members focused on how you approach ambiguous problems, work with scientists and external partners, and thrive in rapidly evolving technical environments. Emphasis on intellectual curiosity and ability to navigate emerging technologies.

Full posting

Original listing.

We are rebuilding biotech for the AI era.

When a breakthrough is delayed, the world waits. Getting a molecule from discovery to patients, or a crop from lab to field, involves thousands of slow, manual, disconnected steps. AI has the potential to change this, compressing decades of R&D work into years. But that only happens when clean, structured scientific data and AI are built into how science gets done.

Benchling is the AI platform for biotech R&D. Scientists use Benchling to design experiments, capture structured data, and run AI agents and models directly in their workflows. Over 200,000 scientists around the world trust Benchling to power their most important work, from academic labs to Sanofi, Moderna, and more than half of the world's top 50 biopharma.

We’re building an AI scientist for our customers. We can’t do that if we haven’t built the muscle ourselves. AI fluency is the foundation we build on; it's core to how we work, and we're committed to helping every new hire integrate it into their day-to-day. As part of our interview process, you'll complete a brief AI-focused exercise or discussion so we can understand how you think about and use AI to drive impact in your role. Feel free to reference any tools, platforms, or workflows you use today.

Role Overview

We’re a team focused on making frontier AI models better at science. LLMs know an extraordinary amount of biology, but there’s still a large gap in reasoning for the real-world problems scientists face every day. We recently published some of our work here.

You’ll build the datasets, evaluations, and systems that help close that gap. You’ll work with scientists to turn complex scientific work into rigorous tasks that models can learn from and be evaluated against. You’ll partner with leading AI labs to understand where models fail and how to improve them.

This is an early and rapidly evolving area. You’ll work at the intersection of software engineering, biology, and frontier AI: finding tasks that are challenging for LLMs and valuable to scientists, designing evaluations that capture real scientific judgment, and building systems to create these tasks at scale.

 

RESPONSIBILITIES

  • Build datasets for evaluating and improving frontier models, turning complex scientific data into high-quality tasks and environments for LLMs.

  • Analyze model failure modes, running experiments across frontier models to understand where they struggle and identify opportunities for improvement.

  • Build scalable data infrastructure, creating pipelines that curate, transform, and validate large volumes of scientific data into tasks for model evaluation and improvement.

  • Collaborate with frontier AI labs, helping develop and evaluate new approaches for improving models on challenging scientific tasks.

  • Work closely with scientists, translating expert judgment into problems and evaluation criteria that can reliably distinguish strong model behavior.

QUALIFICATIONS

  • 2+ years at the intersection of biology and AI, with experience evaluating and improving scientific models or LLMs for biological applications.

  • Experience building with LLMs, with an intuition for where current models excel, where they struggle, and how to design systems around their capabilities.

  • Curiosity and excitement about frontier AI, with a desire to understand and push the capabilities of rapidly improving models.

  • Comfort working on ambiguous problems, where the playing field is rapidly shifting and the right technical approaches are still being discovered.

  • Collaborative mindset, able to work closely with engineers, scientists, and external research partners.

  • Desire to work in a fast-paced environment, where priorities can shift and rapid experimentation is encouraged.

HOW WE WORK

This is an in-person team in San Francisco built around collaborating in the office in a fast-paced environment. We’re in the office Monday through Friday.

#LI-KW1

Benchling welcomes everyone.

We believe diversity enriches our team so we hire people with a wide range of identities, backgrounds, and experiences.

We are an equal opportunity employer. That means we don’t discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. We also consider for employment qualified applicants with arrest and conviction records, consistent with applicable federal, state and local law, including but not limited to the San Francisco Fair Chance Ordinance.

Redirects to Benchling's application page.

Other roles

More at Benchling.

View all 10 roles