Software Engineer, Model Evaluation and Improvement
Machine Learning Engineer · Mid · Full Time
Opens Benchling's application page
Role
What you'll do.
Software Engineer at Benchling focused on evaluating and improving frontier AI models for scientific applications. You'll build datasets, evaluation systems, and data infrastructure that help large language models better reason about complex biological problems. This role bridges software engineering, biology, and frontier AI, requiring 2+ years of experience at the intersection of biology and AI with proven expertise in working with LLMs and building scalable systems for scientific data processing.
Responsibilities
- Build Evaluation Datasets for Frontier Models: Develop high-quality datasets and benchmarks that rigorously evaluate frontier large language models on scientific reasoning tasks. Transform complex biological data, scientific workflows, and domain expertise into structured evaluation tasks that accurately measure model capabilities and identify performance gaps in scientific domains.
- Analyze Model Failure Modes: Design and execute systematic experiments across frontier AI models including GPT-4, Claude, and emerging LLMs to understand failure modes and reasoning limitations in biological and scientific contexts. Use empirical analysis to identify specific weaknesses and generate actionable insights for model improvement and refinement.
- Build Scalable Data Infrastructure: Engineer robust data pipelines and infrastructure systems that curate, transform, validate, and organize large volumes of scientific data into structured tasks for model evaluation. Implement automation frameworks that handle data quality assurance and enable continuous generation of new evaluation benchmarks at scale.
- Collaborate with Frontier AI Labs: Partner with leading AI research organizations and model developers to understand emerging capabilities and limitations. Contribute to the development and validation of novel approaches for improving model reasoning on challenging scientific tasks, providing real-world biotech use cases and feedback.
- Translate Expert Scientific Judgment into Evaluation Criteria: Work directly with domain-expert scientists to convert tacit domain knowledge into explicit, measurable evaluation criteria. Design evaluation frameworks that capture nuanced scientific judgment and enable reliable discrimination between strong and weak model behavior on real-world biotech problems.
Qualifications
What we look for.
Technical
Large Language Model Development Experience
Demonstrated expertise in working with frontier LLMs including prompt engineering, fine-tuning, evaluation frameworks, and understanding model capabilities and limitations. Deep familiarity with current-generation models (GPT-4, Claude, Llama, or equivalent) and ability to design systems that effectively leverage their capabilities while working around their constraints.
Scientific Data Management
Experience designing, building, and maintaining systems for scientific data processing, validation, and curation. Proficiency with structured data formats common in biology and biotech (genomic data, protein sequences, experimental protocols) and ability to transform complex scientific information into machine-readable formats.
Data Pipeline and ETL Engineering
Strong experience building scalable data pipelines, data validation frameworks, and ETL processes. Comfortable with distributed data processing, handling large-scale datasets, and implementing robust error handling and data quality assurance mechanisms.
Python Programming for ML/Data Applications
Advanced Python proficiency for machine learning and data science applications, including working with popular ML libraries (PyTorch, TensorFlow, scikit-learn, Hugging Face). Ability to write clean, maintainable, production-quality code for data processing and model evaluation workflows.
Evaluation Framework Design
Experience designing comprehensive evaluation frameworks and benchmarks for machine learning systems. Understanding of evaluation metrics, statistical significance testing, and ability to create metrics that accurately capture model performance on complex reasoning tasks.
Education
Bachelor's Degree in Computer Science, Engineering, or Quantitative Field
Formal technical education providing strong foundations in algorithms, data structures, and software engineering principles. Alternative: equivalent practical experience demonstrating mastery of these concepts through professional work.
Biology, Bioinformatics, or Life Sciences Background (Preferred)
Academic or professional background in biological sciences, bioinformatics, computational biology, or related fields providing deep understanding of biological concepts, scientific methodology, and domain-specific data formats. Significantly strengthens ability to translate between scientists and AI systems.
Experience
2+ Years at Biology-AI Intersection
Demonstrated professional experience working at the intersection of biology and artificial intelligence, including evaluating and improving scientific models or LLMs for biological applications. This could include roles in biotech, computational biology, AI research applied to life sciences, or related domains.
LLM System Design and Development
Hands-on experience building production systems with large language models, including prompt engineering, evaluation, fine-tuning, or RAG (Retrieval-Augmented Generation) implementations. Deep intuition for where current models excel, where they struggle, and how to architect systems around their capabilities.
Ambiguous Problem-Solving in Emerging Technology
Proven track record of thriving in early-stage or rapidly evolving technical domains where specifications are unclear and technical approaches are still being discovered. Comfort with shifting priorities, high experimentation velocity, and iterative development in frontier technology areas.
Skills
Required
Python
Advanced Python programming for data processing, machine learning, and building evaluation systems. Essential for implementing data pipelines and LLM evaluation frameworks.
LLM Evaluation and Benchmarking
Expertise in designing and implementing evaluation frameworks for frontier models, including creating benchmarks that test reasoning capabilities on scientific tasks. Understanding of evaluation metrics, statistical methods, and best practices for rigorous model assessment.
Biological Domain Knowledge
Strong understanding of molecular biology, genomics, biochemistry, or wet lab processes. Ability to understand complex scientific problems and translate them into tasks that test AI model reasoning on real biotech challenges.
Data Pipeline Engineering
Experience designing and implementing ETL pipelines, data validation frameworks, and scalable data infrastructure. Proficiency with tools for data transformation, validation, and quality assurance.
Prompt Engineering and LLM APIs
Hands-on experience with modern LLM APIs and prompt engineering techniques. Understanding of how to effectively interact with frontier models, design structured prompts, and extract reliable outputs for evaluation and analysis.
Preferred
SQL and Database Design
Nice to haveExperience with SQL query optimization, database design, and working with large-scale databases for scientific data management. Valuable for managing and querying structured evaluation datasets.
Scientific Computing Libraries
Nice to haveProficiency with scientific Python libraries such as NumPy, Pandas, SciPy, and specialized bioinformatics packages. Useful for working with biological data and performing complex scientific computations.
Machine Learning and Fine-tuning
Nice to haveExperience with model fine-tuning, transfer learning, or training custom models on domain-specific tasks. Understanding of how to adapt pre-trained models for specialized applications.
Biotech Domain Experience
Nice to havePrior experience in biotech, pharmaceutical R&D, academic research labs, or related life sciences environments. Familiarity with scientific workflows, wet lab processes, and how AI can address real scientist pain points.
Software Engineering Best Practices
Nice to haveVersion control (Git), testing frameworks, code review processes, and CI/CD pipelines. Experience working on collaborative engineering teams building production systems.
Research Publication and Communication
Nice to haveExperience communicating technical findings to diverse audiences including scientists, engineers, and leadership. Ability to write clear technical documentation and potentially contribute to research publications.
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·USD 136,435 – 166,754
Equity·Stock options
Benefits
Equity Compensation
Stock options as part of comprehensive compensation package, providing upside participation in Benchling's growth as a Series D-funded company at the forefront of AI-powered biotech platforms.
Comprehensive Health Insurance
Medical, dental, and vision coverage with company contributions toward employee and family plans, meeting or exceeding industry standards.
Retirement Planning
401(k) plan with company matching to support long-term financial planning and wealth building.
Professional Development
Learning budget and support for attending conferences, courses, and pursuing certifications in AI, machine learning, and domain expertise areas.
Flexible Time Off
Generous vacation and flexible time-off policies to support work-life balance and employee wellbeing.
Modern Office Environment
Collaborative workspace in San Francisco designed for rapid experimentation, with access to cutting-edge tools and infrastructure for AI and data engineering work.
Parental Leave
Paid parental leave benefits supporting employees through major life transitions.
AI and Technology Resources
Access to frontier AI tools, computational resources, and partnerships with leading AI labs and model providers for conducting research and experiments.
Process
Interview steps.
- 01
AI-Focused Assessment
During the interview process, candidates will complete a focused exercise or discussion exploring how they think about and leverage AI to drive impact in their work. Bring examples of AI tools, platforms, or workflows you currently use. This reflects Benchling's core commitment to AI fluency and helps assess practical experience with frontier models.
- 02
Technical Screening
Conversation focused on your experience building with large language models, designing evaluation frameworks, and working with scientific data. Expect to discuss specific projects where you evaluated model performance or designed systems around LLM capabilities.
- 03
Scientific Domain Depth Interview
Discussion of your biological knowledge and experience translating scientific problems into AI tasks. This may involve walking through how you would design an evaluation for a specific scientific reasoning challenge.
- 04
System Design and Infrastructure Discussion
Technical conversation about building scalable data pipelines and evaluation infrastructure. You may be asked to discuss architectural decisions for handling large-scale scientific datasets or designing evaluation systems.
- 05
Collaboration and Problem-Solving Round
Conversation with team members focused on how you approach ambiguous problems, work with scientists and external partners, and thrive in rapidly evolving technical environments. Emphasis on intellectual curiosity and ability to navigate emerging technologies.
Full posting
Original listing.
We are rebuilding biotech for the AI era.
When a breakthrough is delayed, the world waits. Getting a molecule from discovery to patients, or a crop from lab to field, involves thousands of slow, manual, disconnected steps. AI has the potential to change this, compressing decades of R&D work into years. But that only happens when clean, structured scientific data and AI are built into how science gets done.
Benchling is the AI platform for biotech R&D. Scientists use Benchling to design experiments, capture structured data, and run AI agents and models directly in their workflows. Over 200,000 scientists around the world trust Benchling to power their most important work, from academic labs to Sanofi, Moderna, and more than half of the world's top 50 biopharma.
We’re building an AI scientist for our customers. We can’t do that if we haven’t built the muscle ourselves. AI fluency is the foundation we build on; it's core to how we work, and we're committed to helping every new hire integrate it into their day-to-day. As part of our interview process, you'll complete a brief AI-focused exercise or discussion so we can understand how you think about and use AI to drive impact in your role. Feel free to reference any tools, platforms, or workflows you use today.
Role Overview
We’re a team focused on making frontier AI models better at science. LLMs know an extraordinary amount of biology, but there’s still a large gap in reasoning for the real-world problems scientists face every day. We recently published some of our work here.
You’ll build the datasets, evaluations, and systems that help close that gap. You’ll work with scientists to turn complex scientific work into rigorous tasks that models can learn from and be evaluated against. You’ll partner with leading AI labs to understand where models fail and how to improve them.
This is an early and rapidly evolving area. You’ll work at the intersection of software engineering, biology, and frontier AI: finding tasks that are challenging for LLMs and valuable to scientists, designing evaluations that capture real scientific judgment, and building systems to create these tasks at scale.
RESPONSIBILITIES
Build datasets for evaluating and improving frontier models, turning complex scientific data into high-quality tasks and environments for LLMs.
Analyze model failure modes, running experiments across frontier models to understand where they struggle and identify opportunities for improvement.
Build scalable data infrastructure, creating pipelines that curate, transform, and validate large volumes of scientific data into tasks for model evaluation and improvement.
Collaborate with frontier AI labs, helping develop and evaluate new approaches for improving models on challenging scientific tasks.
Work closely with scientists, translating expert judgment into problems and evaluation criteria that can reliably distinguish strong model behavior.
QUALIFICATIONS
2+ years at the intersection of biology and AI, with experience evaluating and improving scientific models or LLMs for biological applications.
Experience building with LLMs, with an intuition for where current models excel, where they struggle, and how to design systems around their capabilities.
Curiosity and excitement about frontier AI, with a desire to understand and push the capabilities of rapidly improving models.
Comfort working on ambiguous problems, where the playing field is rapidly shifting and the right technical approaches are still being discovered.
Collaborative mindset, able to work closely with engineers, scientists, and external research partners.
Desire to work in a fast-paced environment, where priorities can shift and rapid experimentation is encouraged.
HOW WE WORK
This is an in-person team in San Francisco built around collaborating in the office in a fast-paced environment. We’re in the office Monday through Friday.
#LI-KW1
Benchling welcomes everyone.
We believe diversity enriches our team so we hire people with a wide range of identities, backgrounds, and experiences.
We are an equal opportunity employer. That means we don’t discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. We also consider for employment qualified applicants with arrest and conviction records, consistent with applicable federal, state and local law, including but not limited to the San Francisco Fair Chance Ordinance.
Redirects to Benchling's application page.
Other roles
More at Benchling.
Data Engineer
Mid
Data Engineer
Mid
Software Engineer, Platform (Developer Experience)
Mid
Software Engineer, Platform (Developer Experience)
Mid
Software Engineer, Full Stack (Document Canvas)
Mid