Posted 13 days ago
Content Evaluator – Bilingual (French and English)- Flexible Hours
AI Summary
Bilingual evaluator conducting human quality assessments of AI-generated customer support responses on Instagram, WhatsApp, and Messenger, scoring interactions against multi-tier rubrics and verifying facts against business knowledge bases.
About this role
Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale. We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.
Role Overview
We are seeking detail-oriented evaluators to conduct human quality evaluations for an enterprise AI customer support product on Instagram, WhatsApp, and Messenger. You will evaluate and benchmark AI model responses against complex evaluation rubrics using provided business knowledge bases.
What You’ll Own:
- Model Evaluation: Review and score AI-generated customer interactions across Foundational, Experiential, and Operational dimensions (e.g., Action Fidelity, Faithfulness, Hallucination, Compliance, Tone, and Handoff).
- Intent & Fact Verification: Benchmark both informational (R1) and transactional (R2) customer queries against authoritative business sources (FAQs, product catalogs, SOPs) within the task UI.
- Quality Assurance: Participate in dual-review processes and daily calibration audits to ensure inter-rater agreement and establish ground-truth performance targets.
- Performance Targets: Deliver precise evaluation
You’ll Thrive in This Role If You Have:
- Customer Service Background: Prior experience in customer service, call centers, retail, or handling customer communications via email, chat, or phone (highly prioritized).
- English Proficiency: Exceptional written English skills with a strong command of tone, brand voice, grammar, and nuance.
- French Proficiency: Exceptional written French skills with a strong command of tone, brand voice, grammar, and nuance.
- Analytical Precision: Ability to strictly follow multi-tier evaluation guidelines, complex logic trees, and technical rubrics without deviation.
- Tech Adaptability: Comfort using dedicated web-based tools and labeling interfaces.
Please be aware of recruitment scams involving individuals or organizations falsely claiming to represent employers. Innodata will never ask for payment, banking details, or sensitive personal information during the application process. To learn more on how to recognize job scams, please visit the Federal Trade Commission’s guide at https://consumer.ftc.gov/articles/job-scams.
If you believe you’ve been targeted by a recruitment scam, please report it to Innodata at verifyjoboffer@innodata.com and consider reporting it to the FTC at ReportFraud.ftc.gov.
Skills
Explore related jobs
More jobs at Innodata Inc.
- Data Program Manager, AI/ML DeliveryHybrid - New Jersey
- AI Voice Evaluation SpecialistRemote - Alabama
- AI Agentic Workflow ReviewerIn Office - San Jose, California
- Quality Lead, Agentic AI Workflow EvaluationIn Office - San Jose, California
- Engagement Manager, Agentic AI Workflow EvaluationsIn Office - San Jose, California
- Research Data ScientistRemote - United States
Browse these categories
Market data for this role
All reports →- SeriesRole reportsOne role family at a time: how many openings, what changed this week, who is hiring, what it pays.
- SeriesSalary reportsWhat employers publish in job postings, by level and workplace. Not self-reported pay.
- Market overviewState of tech hiring, September 2026: up 4.8%Tech hiring rose 4.8% month over month in September 2026, with 411,122 new listings. Customer support and account executive roles led the growth.