AI Safety & Governance Full-time

Task Development Engineer

Model Evaluation and Threat Research

Posted

Aug 10, 2026

Location

Remote

Type

Full-time

Compensation

Up to $624000

Mission

What you will drive

  • In this role, you'll develop challenging, novel evaluation tasks for frontier AI models that remain difficult as model capabilities grow.
  • Conduct quality assurance on existing tasks to verify solvability and ensure models receive only necessary information.
  • Baseline and score task completions using your domain expertise to validate evaluation quality.
  • Improve task development infrastructure by identifying inefficient workflows and implementing enhancements.
  • Stay current with evaluation methodologies and frontier AI capabilities to inform effective task design.

Profile

What makes you a great fit

  • In this role, you'll develop challenging, novel evaluation tasks for frontier AI models that remain difficult as model capabilities grow.
  • Conduct quality assurance on existing tasks to verify solvability and ensure models receive only necessary information.
  • Baseline and score task completions using your domain expertise to validate evaluation quality.
  • Improve task development infrastructure by identifying inefficient workflows and implementing enhancements.
  • Stay current with evaluation methodologies and frontier AI capabilities to inform effective task design.

About

Inside Model Evaluation and Threat Research

Visit site →

Model Evaluation and Threat Research (formerly Alignment Research Center, Evaluations) is a project focused on evaluating the capabilities and alignment of advanced ML models.