Application Guide
How to Apply for Task Development Engineer
at Model Evaluation and Threat Research
๐ข About Model Evaluation and Threat Research
Model Evaluation and Threat Research (METR) is at the forefront of AI safety, dedicated to rigorously evaluating the capabilities and alignment of advanced machine learning models. As a nonprofit project spun out of the Alignment Research Center, it offers a unique opportunity to directly influence the safe development of frontier AI. Working at METR means collaborating with a team of experts who are shaping how the world understands and mitigates AI risks.
About This Role
As a Task Development Engineer, you will design challenging evaluation tasks that push the boundaries of AI capabilities, ensuring these tasks remain difficult as models evolve. Your work directly informs the AI community about model limitations and safety, making your contributions critical to METR's mission. This role combines technical skill with creative problem-solving to build robust evaluations that can withstand the rapid advancement of AI.
๐ก A Day in the Life
A typical day might involve brainstorming new task ideas with the team, prototyping a task in Python, and running it against a frontier model to test its difficulty. You might also spend time reviewing tasks from other engineers, providing feedback on their solvability and clarity, and iterating on your own tasks based on baseline results. Collaborating with researchers to align tasks with current AI safety concerns is also a regular part of the day.
๐ Application Tools
๐ฏ Who Model Evaluation and Threat Research Is Looking For
- Someone with a strong background in AI/ML, including hands-on experience with state-of-the-art models and an understanding of their strengths and weaknesses.
- A creative problem-solver who can design novel tasks that test complex reasoning, creativity, or other high-level capabilities in models.
- Detail-oriented with experience in quality assurance, ensuring tasks are solvable and free of unintended hints.
- Proficient in programming (e.g., Python) to build task infrastructure and automate scoring, with a knack for optimizing workflows.
๐ Tips for Applying to Model Evaluation and Threat Research
Highlight any experience you have designing or contributing to evaluation tasks, benchmarks, or adversarial testing of AI models.
Showcase your ability to think like a modelโdescribe a task you've created that specifically targeted a known limitation of current models.
In your resume and cover letter, provide concrete examples of how you've improved workflows or built tools to streamline development processes.
Research METR's published evaluations and mention specific ones you admire, explaining how you would extend or improve upon them.
Tailor your application to emphasize your commitment to AI safety and alignment, as METR is a mission-driven organization.
โ๏ธ What to Emphasize in Your Cover Letter
["Your passion for AI safety and how your skills align with METR's mission of rigorous model evaluation.", "Specific examples of tasks or evaluations you've developed, detailing the challenges and your solutions.", 'Your technical expertise in AI/ML and programming, and how it enables you to create robust, scalable evaluation tasks.', 'Your collaborative spirit and ability to work with a distributed team, as the role is remote and requires cross-functional coordination.']
Generate Cover Letter โ๐ Research Before Applying
To stand out, make sure you've researched:
- โ Read METR's published evaluations and blog posts to understand their methodologies and the types of tasks they've created.
- โ Familiarize yourself with the broader AI safety landscape, including other organizations like Anthropic, OpenAI, and ARC, to see how METR fits in.
- โ Study recent frontier model capabilities (e.g., GPT-4, Claude, Gemini) to understand current strengths and weaknesses.
- โ Look into the team members on METR's website to understand their backgrounds and potential collaborators.
๐ฌ Prepare for These Interview Topics
Based on this role, you may be asked about:
โ ๏ธ Common Mistakes to Avoid
- Don't submit a generic cover letter that doesn't mention METR or AI safety; show you've done your homework.
- Avoid focusing solely on academic achievements without demonstrating practical application of your skills to evaluation tasks.
- Don't underestimate the importance of infrastructure; METR values engineers who can build tools, so don't neglect to highlight your software development skills.
๐ Application Timeline
This position is open until filled. However, we recommend applying as soon as possible as roles at mission-driven organizations tend to fill quickly.
Typical hiring timeline:
Application Review
1-2 weeks
Initial Screening
Phone call or written assessment
Interviews
1-2 rounds, usually virtual
Offer
Congratulations!
Ready to Apply?
Good luck with your application to Model Evaluation and Threat Research!