Machine Learning Engineer - Data Pipeline

Machine Learning Engineer · Senior · Full Time

Dublin, CA (HQ)USD 150k – 225k2mo ago
Apply for this role

Opens Articul8's application page

Role

What you'll do.

Articul8 is seeking a Machine Learning Engineer to design and develop sophisticated data processing pipelines for AI model training. The ideal candidate will be responsible for end-to-end data acquisition, processing, and quality improvement, working closely with research and engineering teams to power next-generation domain-specific AI models.

Responsibilities

  • Data Pipeline Development: Design and develop comprehensive data processing pipelines including extraction, filtering, and labeling of diverse data sources
  • Machine Learning Model Implementation: Develop and implement ML models to enhance data quality and diversity, including quality classifiers and verification models
  • Data Acquisition Engineering: Lead engineering projects focused on web crawling, data ingestion, and large-scale data processing
  • Distributed Systems Architecture: Develop and deploy highly scalable distributed systems capable of handling terabytes of data with robust indexing and search capabilities
  • Cross-Team Collaboration: Work closely with Applied Research, Technology, and Architecture teams to ensure seamless data flow and system operability
  • Infrastructure Management: Deploy solutions in Kubernetes Infrastructure-as-Code environment and perform routine system maintenance and checks

Qualifications

What we look for.

Technical

  • Deep Learning Frameworks

    Proficiency in at least one deep learning framework, such as PyTorch

  • Programming Languages

    Advanced proficiency in Python with ability to write clean, maintainable code

  • Distributed Systems

    Strong expertise in large stateful distributed systems and data processing technologies

  • Data Processing Tools

    Familiarity with distributed workload technologies like multiprocessing, Ray, Docker, and Kubernetes

Education

  • Advanced Degree

    BS/MS/PhD in Computer Science, Machine Learning, or related technical field

Experience

  • Machine Learning Project Experience

    Proven experience in machine learning projects, particularly in text or vision domains

  • Data Pipeline Development

    Demonstrated ability to build large-scale data processing pipelines and datasets

Skills

Required

  • Python Programming

    Strong programming skills in Python for machine learning and data processing

  • Data Pipeline Engineering

    Expertise in designing and implementing complex data acquisition and processing systems

  • Machine Learning Model Training

    Experience in training machine learning models to solve specific problems

Preferred

  • GitHub Contributions

    Nice to have

    Active open-source contributions and public code repositories

  • Data Crawling Tools

    Nice to have

    Experience with tools like Scrapy, Selenium, Hadoop, and Datasketch

  • Multilingual Skills

    Nice to have

    Proficiency in multiple languages to support diverse data collection

Tech stack

Languages

Python

Frameworks

PyTorchRay

Databases

Key-Value Databases

Tools

DockerKubernetesScrapy

Other

SeleniumHadoop

Compensation

Pay and benefits.

Base·USD 150,000 – 225,000

Equity·Stock options

Benefits

  • Health Insurance

    Comprehensive medical, dental, and vision coverage

  • Equity Compensation

    Stock options with potential for significant growth in AI startup

  • Professional Development

    Continuous learning opportunities, conference attendance, and skill development programs

  • Flexible Work Environment

    Supportive culture emphasizing diversity, creativity, and personal growth

Process

Interview steps.

  1. 01

    Initial Screening

    Resume and background review by recruiting team

  2. 02

    Technical Phone Screen

    Discussion of machine learning and data engineering experience with senior engineer

  3. 03

    Coding Challenge

    Take-home project involving data pipeline design and ML model implementation

  4. 04

    Onsite Technical Interviews

    Multiple rounds covering system design, machine learning concepts, and coding skills

  5. 05

    Final Leadership Interview

    Meeting with engineering leadership to assess cultural fit and long-term potential

Full posting

Original listing.

About us: 

At Articul8 AI, we relentlessly pursue excellence and create exceptional AI products that exceed customer expectations. We are a team of dedicated individuals who take pride in our work and strive for greatness in every aspect of our business. We believe in using our advantages to make a positive impact on the world and inspiring others to do the same. 

Job Description: 

We are seeking machine learning engineers to join our team full-time. As part of your role, you will help us build pipelines of data collection, data extraction, data filtering/synthetic data generation and data analysis.  You will own all work related to acquiring high-quality data to power the training of our domain-specific models end to end.  You will work closely with other researchers and engineers to empower our next generation of domain-specific models.  We value rapid prototyping, iterating, and shipping new systems quickly.  

Required Qualifications: 

  • BS/MS/PhD in Computer Science or a related field. 

  • Proficiency in at least one deep learning framework, such as PyTorch. 

  • Experience in machine learning projects in text or vision, e.g., has trained machine learning models to tackle a specific problem. 

  • Strong expertise in large stateful distributed systems and data processing. 

  • Strong proficiency in building large-scale data processing pipelines, familiar with distributed workload (e.g., multiprocessing, Ray, Docker, Kubernetes). 

  • Proficiency in at least one programming language commonly used in machine learning, such as Python and ability to write clean, maintainable code. 

  • Excellent problem-solving skills and attention to detail, especially when handling data anomalies and biases to further improve data quality. 

Key Competencies 

  • Active Github contributions are a big plus. 

  • Experience in building large-scale datasets. 

  • Familiar with at least one of the following tools for data crawling (e.g. Scrapy), data collection (e.g., VPNs, Selenium), data processing (e.g., Hadoop, Datasketch). 

  • Building bespoke data processing libraries from scratch. 

  • Keeping up with state-of-the-art techniques for preparing AI training data. 

  • Organizing and meticulously bookkeeping data across multiple clouds, of multiple modalities, and from many sources. 

  • Multilingual which contributes to enriching the language diversity crucial for robust model training. 

Responsibilities: 

  • Design and develop data processing pipelines, including data extraction, data filtering, data labeling, etc. 

  • Implement machine learning models to improve the quality and diversity of data (especially in the data extraction stage), e.g., quality classifier, document layout model, code verification model, etc. 

  • Own and lead engineering projects in the area of data acquisition, including web crawling, data ingestion, and processing. 

  • Collaborate with our Applied Research, Technology, and Architecture teams to ensure smooth data flow and system operability. 

  • Develop and deploy highly scalable distributed systems capable of handling terrabytes of data. 

  • Architect and implement algorithms for data indexing and search capabilities. 

  • Build and maintain backend services for data storage, including work with key-value databases and synchronization. 

  • Deploy solutions in a Kubernetes Infrastructure-as-Code environment and perform routine system checks. 

By joining our team, you become part of a community that embraces diversity, inclusiveness, and lifelong learning. We nurture curiosity and creativity, encouraging exploration beyond conventional wisdom. Through mentorship, knowledge exchange, and constructive feedback, we cultivate an environment that supports both personal and professional development. 

Your future experience at Articul8 will include continuous learning and growth opportunities as we embark on an exciting journey to disrupt the status quo. If you're excited about joining a team that's passionate about making a difference, we want to hear from you. 

If you're ready to join a team that's changing the game, apply now to become a part of the Articul8 team. Join us on this adventure and help shape the future of Generative AI in the enterprise. 

Redirects to Articul8's application page.

Other roles

More at Articul8.