About this opportunity
Role
Develop frontier AI benchmarks and datasets, working with research partners, customers, and internal data, product, and engineering teams.
Responsibilities
Design datasets for frontier model training and evaluation.
Analyze benchmark results and communicate research findings.
Apply current LLM evaluation research to workflows.
Collaborate across research, data operations, product, engineering, and strategy.
Represent research through publications, talks, reports, and customer engagements.
Requirements
Strong research background in AI/ML evaluation, NLP, or related fields.
Experience with rigorous experimental design and evaluating training/evaluation data.
Excellent technical and non-technical communication skills.
Comfortable with fast-paced, cross-functional environments.
Interest in AI data services and startup/GTM strategy.
PhD in ML, NLP, or related field preferred; equivalent industry/research experience considered.