Task Development Engineer
Model Evaluation and Threat Research
Posted
Aug 10, 2026
Location
Remote
Type
Full-time
Compensation
Up to $624000
Mission
What you will drive
- In this role, you'll develop challenging, novel evaluation tasks for frontier AI models that remain difficult as model capabilities grow.
- Conduct quality assurance on existing tasks to verify solvability and ensure models receive only necessary information.
- Baseline and score task completions using your domain expertise to validate evaluation quality.
- Improve task development infrastructure by identifying inefficient workflows and implementing enhancements.
- Stay current with evaluation methodologies and frontier AI capabilities to inform effective task design.
Profile
What makes you a great fit
- In this role, you'll develop challenging, novel evaluation tasks for frontier AI models that remain difficult as model capabilities grow.
- Conduct quality assurance on existing tasks to verify solvability and ensure models receive only necessary information.
- Baseline and score task completions using your domain expertise to validate evaluation quality.
- Improve task development infrastructure by identifying inefficient workflows and implementing enhancements.
- Stay current with evaluation methodologies and frontier AI capabilities to inform effective task design.
About
Inside Model Evaluation and Threat Research
Model Evaluation and Threat Research (formerly Alignment Research Center, Evaluations) is a project focused on evaluating the capabilities and alignment of advanced ML models.