Back to opportunities
Surge AI

Research Engineer, Coding Evaluation & Training Data

Posted 2026-09-24
Work Type
Remote
Location
Global
Compensation
Not specified

About this opportunity

Overview

Surge is hiring a Research Engineer, Coding Evaluation & Training Data to help build and run systems that teach frontier models how to code. This role sits at the intersection of software engineering and product, with a focus on training data quality, agentic evaluation, and system design for top AI labs.

The work centers on creating coding data projects that reflect real-world software engineering, including RL environments, evaluation schemes, and workflows used to assess and improve model coding capabilities.

What You'll Do

  • Own coding data projects end to end, from scoping and pilot design through execution, iteration, and scale-up.
  • Design agentic training workflows and task structures that mirror real SWE work, such as refactoring, debugging, code review, large-repo navigation, and tool use.
  • Define and refine rubrics, golden sets, and reward signals that measure true engineering value beyond simple compilation success.
  • Review data and worker output with strong engineering judgment and decide what meets the standard for frontier training.
  • Design and run qualification processes for coding workers, including hands-on assessments of coding ability.
  • Set up or partner on technical environments such as containers, repositories, test harnesses, sandboxes, and code execution infrastructure.
  • Work with technical staff at client companies to turn high-level training goals into concrete projects and environments.
  • Collaborate with Surge engineering, product, and operations to improve coding data products, internal tools, and execution processes.

What We're Looking For

  • 3–6+ years of professional software engineering experience building and maintaining real systems.
  • Strong coding ability in at least one mainstream programming language and comfort working in production codebases.
  • Strong judgment for good engineering practices, including correctness, code quality, and how real teams work.
  • Ability to reason about and debug technical environments, including containers, dependencies, and automated test setups.
  • Interest in owning projects end to end, including scoping, workflow design, execution, and continuous improvement.
  • Excellent written and verbal communication skills, including the ability to speak credibly with senior client engineers and translate fuzzy goals into actionable plans.
  • Interest in AI/ML systems and in how data, evaluation, and reward design improve agentic coding capabilities.

Skills & expertise

Software Engineering Programming/Coding AI/ML Technical Evaluation Model Evaluation Prompt Evaluation Containerization Test Automation

Related Opportunities

Trending Right Now

View all

Don't miss the next opportunity

Get a curated Friday digest of new AI and remote roles based on your location and interests — free, no spam.

Get Weekly Opportunities