About this opportunity
Overview
Innodata is hiring detail-oriented Voice Specialists to evaluate AI model performance through real-time, voice-based conversations. In this role, you will interact with two AI models using the same assigned scenario, compare their responses, and produce structured evaluations based on defined quality criteria. The goal is to make the comparison fair and consistent so the stronger conversational experience can be identified.
This position is remote and the listed hourly pay range is $20-$27 per hour, depending on experience, skills, and qualifications.
What You’ll Do
- Review an assigned scenario and roleplay a natural conversation with two different AI models.
- Keep the scenario, conversational approach, and interaction style as consistent as possible across both models.
- Maintain a similar number of turns with each model to support a fair comparison.
- Record your voice during each interaction and closely observe response quality and behavior.
- Evaluate each model across five defined dimensions.
- Identify and categorize relevant error clusters that may appear in the audio or conversation.
- Compare Model A and Model B using the evaluation framework and your observations.
- Select the model with the stronger overall performance.
- Write a detailed rationale that explains your final preference using specific examples from both conversations.
- Apply the evaluation guidelines consistently across different scenarios and model interactions.
Requirements
- Bachelor’s degree.
- Strong attention to detail and the ability to notice subtle differences in conversational quality.
- Excellent listening and comprehension skills.
- Strong written communication skills with the ability to explain observations clearly and objectively.
- Ability to follow detailed evaluation guidelines consistently.
- Comfort speaking naturally and roleplaying different conversational scenarios.
- Ability to compare two interactions fairly without personal preference affecting the evaluation.
- Strong critical-thinking and analytical skills.
- Reliability and consistency when completing structured evaluation tasks.
- Familiarity with AI assistants, voice-based AI, or conversational systems.