How to apply for Head of Evals
Trajectory Labs, PBC
About Trajectory Labs, PBC
Trajectory Labs, PBC is a public benefit corporation building RL environments for frontier AI labs to train robust, secure, and reliable models. Unlike typical AI startups, their core mission is AI safety, and they work directly with leading labs to red-team and evaluate the most advanced models. This is a chance to be at the forefront of AI safety, shaping how frontier models are tested and improved.
About the role
As Head of Evals, you will lead the design and execution of safety evaluations for frontier AI models, ensuring red-teaming data quality and building scalable evaluation pipelines. You'll set the strategy for which safety tests to prioritize, automate judgment patterns into agent skills, and communicate critical safety findings to frontier labs. This role is pivotal in ensuring that advanced AI systems are robust, secure, and reliable before deployment.
A typical day
A typical day might involve reviewing red-teaming transcripts to identify quality issues, coding automation scripts to streamline evaluation pipelines, and meeting with frontier lab partners to discuss safety findings. You'll also spend time strategizing which new evaluations to build next and mentoring team members on best practices for model assessment.
Who Trajectory Labs, PBC is looking for
- 2+ years of software engineering experience, with a strong track record of using LLMs and coding agents to automate complex workflows independently.
- Experience building evaluations, benchmarks, or grading pipelines, ideally in AI safety or red-teaming contexts.
- Familiarity with prompt injection, red teaming, or AI safety research, and the ability to identify failure modes in model behavior.
- Strategic thinker who can prioritize which safety tests to build next and scale data collection and model assessment.
- Excellent communicator, able to translate technical safety findings into actionable insights for frontier lab partners.
Tips for this application
- Highlight concrete examples of how you've used LLMs or coding agents to automate workflows, especially in evaluation or red-teaming pipelines.
- Showcase any experience with designing or implementing evaluations, benchmarks, or grading systems—even if not directly in AI safety.
- Demonstrate familiarity with AI safety concepts like prompt injection or red teaming by referencing specific projects or research you've engaged with.
- Emphasize your ability to set strategy and prioritize in a fast-paced startup environment, perhaps by describing a time you scaled a data collection or assessment process.
- Tailor your resume and cover letter to Trajectory Labs' mission; mention why building RL environments for frontier AI labs excites you and how you can contribute to their safety goals.
What to cover in your cover letter
In your cover letter, focus on: (1) your hands-on experience automating workflows with LLMs and coding agents, (2) any past work building evals, benchmarks, or grading pipelines, (3) your understanding of AI safety challenges like prompt injection or red teaming, and (4) your strategic approach to prioritizing safety tests and scaling evaluation efforts. Also, express a clear passion for Trajectory Labs' mission to train robust, secure, and reliable models.
Draft a cover letterResearch before applying
- Explore Trajectory Labs' website to understand their specific approach to building RL environments and their partnerships with frontier AI labs.
- Read up on recent AI safety evaluation frameworks and red-teaming methodologies, especially those used by organizations like Anthropic, OpenAI, or DeepMind.
- Familiarize yourself with prompt injection techniques and existing benchmarks for model safety, such as those from HELM or other eval suites.
- Look into Trajectory Labs' public benefit corporation status and any published research or blog posts to align with their mission and culture.
Likely interview topics
Based on the job description, expect questions about:
- How would you design an evaluation to test a frontier model's robustness against prompt injection attacks?
- Describe a time you automated a manual judgment process into an agent skill or pipeline. What was the impact?
- What metrics or signals do you use to assess the quality of red-teaming data, and how would you ensure consistency?
- How would you prioritize which safety evaluations Trajectory Labs should build next, given limited resources?
- Can you walk us through a failure mode you've identified in an LLM and how you communicated it to stakeholders?
Common mistakes to avoid
- Avoid generic applications that don't demonstrate specific experience with evaluations or red-teaming; this role requires hands-on technical skills.
- Don't overlook the startup nature of the role—candidates who seem inflexible or unable to prioritize in a fast-changing environment may be less appealing.
- Steer clear of downplaying AI safety knowledge; even if you're not an expert, showing genuine interest and some familiarity is crucial.
Deadline
No deadline is listed. Roles without a deadline usually close once the employer has enough candidates, so apply soon if you are interested.