Data Annotation Specialist, Software Engineering
Data Annotation Specialist · Mid · Contract · Remote
Opens Cohere's application page
Role
What you'll do.
As a Data Annotation Specialist at Cohere, you will evaluate frontier AI models' ability to execute complex coding tasks, debug code, and navigate repository architecture while providing critical feedback that directly shapes model development. This contract role requires 3-5 years of software engineering expertise and proficiency in Python, Java, JavaScript, Go, and SQL to assess code generation quality, label machine-written outputs, and track performance trends across model trajectories.
Responsibilities
- Code Evaluation and Assessment: Evaluate foundation AI models' ability to respond to complex coding requests, workflows, and code base-related questions using available development tools and frameworks. Assess accuracy of generated code responses and model-generated debugging solutions.
- Agent Trajectory Analysis: Analyze and assess agent trajectories and model capabilities for code generation and debugging requests. Review decision-making patterns and logic applied by AI models when completing complex software engineering tasks.
- Complex Task Execution: Prompt models to complete intricate coding tasks spanning multiple programming languages and architectural patterns. Validate the correctness and efficiency of generated solutions against industry best practices.
- Data Labeling and Quality Assurance: Label, proofread, and refine machine-written and human-written software engineering outputs. Identify inconsistencies, logical errors, and performance issues in AI-generated code and documentation.
- Performance Metrics and Reporting: Report quality and performance trends related to model and agent behavior. Track metrics on code generation accuracy, debugging capability, and model trajectories to inform model improvement initiatives.
Qualifications
What we look for.
Technical
Python Proficiency
Expert-level knowledge of Python including OOP principles, libraries, frameworks, and debugging practices essential for evaluating AI-generated code quality.
Multi-Language Programming
Proficient knowledge of Java, JavaScript, Go, and SQL (any dialect) to assess code generation and debugging across diverse technology stacks.
API Design and Deployment
Demonstrated experience designing, building, and deploying REST APIs or other API architectures to evaluate model-generated API solutions and architecture reviews.
Code Repository Architecture
Understanding of repository structures, version control systems (Git), branching strategies, and codebase navigation to assess model's ability to work within complex projects.
Debugging Techniques
Proficiency in debugging methodologies, code analysis tools, and error resolution strategies to evaluate model-generated debugging solutions for correctness and efficiency.
Education
Software Engineering Background
Formal study or degree in Computer Science, Software Engineering, or related field providing theoretical foundation in algorithmic design and software architecture.
Experience
Software Engineering Experience
3-5 years of professional software engineering experience in production environments, including hands-on coding, code reviews, and system design contributions.
Code Agent Evaluation Experience
Prior experience working with code agents such as OpenAI's Codex, Anthropic's Claude Code, or similar tools, or evaluating agent trajectories and model outputs in AI systems.
Skills
Required
Python
Advanced proficiency in Python for evaluating code quality, logic correctness, and best practices in AI-generated solutions.
Java
Competent knowledge of Java including OOP patterns, concurrency, and debugging for assessing code generation across enterprise environments.
JavaScript
Proficiency in JavaScript including ES6+ features, async programming, and client-server interactions for evaluating web and Node.js applications.
Go
Understanding of Go programming language including goroutines, channels, and systems programming to evaluate performance-critical code generation.
SQL
Competent SQL skills across multiple dialects for evaluating database query generation, optimization, and data manipulation code.
Code Review and Analysis
Strong ability to review code for correctness, performance, security vulnerabilities, and adherence to software engineering best practices.
Debugging and Troubleshooting
Advanced debugging capabilities including log analysis, breakpoint debugging, and systematic problem-solving methodologies.
Preferred
LLM and Code Agent Experience
Nice to haveHands-on experience evaluating or working with large language models for code generation, particularly tools like Cursor, OpenCode, Claude Code, or similar AI coding assistants.
Machine Learning Fundamentals
Nice to haveBasic understanding of machine learning concepts, training workflows, and how model outputs are evaluated for quality and bias.
API Design Expertise
Nice to haveAdvanced knowledge of API design patterns, REST conventions, GraphQL, or other API technologies for evaluating model-generated API solutions.
Technical Documentation
Nice to haveExperience writing or reviewing technical documentation, which aids in evaluating AI-generated explanations and code comments.
Quality Assurance Methodology
Nice to haveBackground in QA practices, test case design, or quality metrics that inform structured evaluation of model outputs.
Tech stack
Languages
Frameworks
Databases
Tools
Other
Compensation
Pay and benefits.
Base·CAD 41,600 – 41,600
Benefits
Performance Incentives
Additional compensation opportunities based on assessment quality, project completion rates, and consistent delivery standards.
Flexible Remote Work
Work from anywhere within Canada with complete flexibility over your work environment and schedule (minimum 16 hours weekly commitment).
BYOD Arrangement
Bring Your Own Device (laptop) arrangement, eliminating equipment constraints and allowing you to work on your preferred development setup.
Extended Contract Duration
12-month contract providing stability and sustained engagement with a cutting-edge AI research company at the forefront of enterprise AI development.
Access to Frontier AI Models
Hands-on experience evaluating and working with state-of-the-art foundation models, providing competitive industry exposure and direct contribution to AI advancement.
Portfolio Building
Opportunity to develop deep expertise in AI model evaluation and code generation assessment, valuable for advancing your career in AI and machine learning fields.
Process
Interview steps.
- 01
Resume and Writing Sample Review
Initial screening phase where the Talent Team evaluates your resume, professional background, and submitted writing samples to assess communication clarity and technical depth. Emphasize your specific software engineering experience, programming language proficiencies, and any prior work with AI or code generation tools.
- 02
Virtual Annotation Test
Comprehensive technical assessment evaluating both written communication and coding skills through multiple components including a coding take-home assignment, writing samples, and language-based technical tasks. This stage directly mirrors work responsibilities, requiring you to demonstrate code evaluation ability, problem-solving approach, and attention to detail in assessing code quality.
- 03
Video Screening Interview
Conversational video call with the Operations Team to discuss your background, motivation for the role, contractor arrangement understanding, and availability. Prepare to articulate your experience with code agents, machine learning concepts, and why you're drawn to contributing to frontier AI model development.
- 04
Independent Contractor Agreement
Upon selection, you'll receive and sign the Independent Contractor Agreement. Ensure you understand contractor implications including no employment benefits, project-based engagement, potential work fluctuation, and requirements to maintain IP confidentiality and declare competing external relationships.
Full posting
Original listing.
Who are we?
Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.
We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that.
We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft.
We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us!
Why this role?
This role will focus on evaluating coding tasks, requiring you to review and debug code, navigate repository architecture, and analyze model trajectories. Your work will contribute to our model development efforts and the logic our models apply when completing task requests.
Please note: This is a part-time independent contractor position available within Canada. We seek candidates who are able to commit to 16 hours per week minimum at a 40 CAD/hour contract rate. This role is BYOD 💻 - Bring Your Own Device (laptop). Remote work within Canada. 12 month contract. Performance incentives included!
As a Data Annotation Specialist, you will:
Evaluate the model's ability to respond to coding requests, workflows, and code base-related questions using available tools.
Assess agent trajectories and model capabilities for code generation and debugging requests.
Prompt models to complete complex coding tasks and review the accuracy of generated responses.
Label, proofread, and improve machine-written and human-written software engineering-related outputs.
Report quality and performance trends related to model/agent behaviour and project assignments.
You may be a good fit if you have:
Subject matter expertise in software engineering - you have both studied in this field and have 3-5 years of relevant industry experience.
Proficient knowledge and understanding of Python and the following programming languages: Java, JavaScript, Go, and SQL (any dialect).
Prior experience designing, building, and deploying APIs
Prior experience working with code agents (OpenCode, Claude Code, Codex, Cursor) or evaluating agent trajectories is a big plus.
The Candidate Journey:
Initial Screening - Once you have submitted your application, our Talent Team will review your resume and writing samples.
Virtual Annotation Test - This assignment will test your written and technical skills through various language-based tasks, such as a coding take-home assessment, writing sample, and more.
Video Screen - If selected to move forward, you will have a short video call with a member of our Operations Team!
Offer - Independent Contractor Agreement
As an independent contractor, you maintain control over how you complete your work and may work with multiple clients simultaneously. We request that you declare any external work relationships with Cohere’s direct competitors and always maintain the IP confidentiality of the Cohere project. Independent contractors are not eligible for health benefits or other benefits provided to employees. Compensation for services is provided to contractors by self-invoicing for services provided pursuant to the terms of our agreement with the contractor.
It is important to understand that, as an independent contractor, continuous work is not guaranteed. The client-contractor relationship is fundamentally project-based, meaning engagements may be temporary, periodic, or intermittent based on our organizational needs and project availability. As an independent contractor, you should anticipate fluctuations in workflow and, therefore, compensation for services when Cohere does not require as many hours of services in a week.
Prospective candidates, please be advised: this role involves working with human-generated and model-generated tasks that may involve exposure to not safe for work (NSFW) text content as part of data annotation tasks, including explicit, offensive, or other inappropriate material.
If the above qualifications do not perfectly align with your experience, we still encourage you to apply!
We value and celebrate diversity and strive to create an inclusive work environment for all. We welcome applicants from all backgrounds and are committed to providing equal opportunities. Should you require any accommodations during the recruitment process, please submit an Accommodations Request Form, and we will work together to meet your needs.
Redirects to Cohere's application page.
Other roles
More at Cohere.
Forward Deployed Engineer, Infrastructure Specialist (France)
Mid
Software Engineer, Integrations
Mid
Forward Deployed Engineer, Infrastructure Specialist (Singapore)
Senior
Forward Deployed Engineer, Infrastructure Specialist (South Korea)
Mid
Engineering Manager, FDE Agentic Platform
Manager