Application Guide

How to Apply for Head of Evals, AI Red Teaming

at Trajectory Labs, PBC

๐Ÿข About Trajectory Labs, PBC

Trajectory Labs, PBC is a startup that builds RL environments for frontier AI labs to train robust, secure, and reliable models. They are unique in their focus on red-teaming and evaluation quality, directly impacting the safety of next-generation AI systems. Working here means shaping how AI models defend against prompt injection and other vulnerabilities, with a direct line to frontier model developers.

About This Role

As Head of Evals, AI Red Teaming, you will lead the design and quality of safety evaluations deployed to frontier model developers. You'll review and approve red-teaming tasks, transcripts, and grades, and design new evaluation methodologies to address robustness gaps. This role is impactful because you'll directly influence how AI models learn to defend against prompt injection, a critical security concern.

๐Ÿ’ก A Day in the Life

You might start by reviewing a batch of red-teaming transcripts, calibrating grades and flagging inconsistencies. Later, you'd collaborate with engineers to refine an agent-based checker that automates part of the review process. You'd also spend time designing a new evaluation environment targeting a recently discovered prompt injection vulnerability, ensuring it's robust and scalable.

๐ŸŽฏ Who Trajectory Labs, PBC Is Looking For

  • Has deep experience with LLM evaluation, including designing benchmarks, building LLM judges, or creating eval environments.
  • Possesses calibrated judgment for assessing red-teaming transcripts and grades, with sustained attention to detail across hundreds of reviews.
  • Is fluent with LLMs and coding agents, and can build agent-based automation tools to scale review processes.
  • Has prior experience in prompt injection, red-teaming, or adversarial testing of AI systems, and can lead evaluation quality initiatives.

๐Ÿ“ Tips for Applying to Trajectory Labs, PBC

1

Highlight specific examples of red-teaming or evaluation work you've done, especially involving prompt injection or adversarial attacks.

2

Demonstrate your ability to build automation tools for scaling review processesโ€”mention any agent-based systems you've developed.

3

Show familiarity with frontier AI labs and their safety needs; reference recent developments in AI red-teaming.

4

Emphasize your experience with LLM judges or evaluation methodologies, providing concrete metrics or outcomes.

5

Tailor your resume to include keywords like 'RL environments', 'prompt injection', 'eval design', and 'agent-based automation'.

โœ‰๏ธ What to Emphasize in Your Cover Letter

['Your hands-on experience designing and reviewing red-teaming evaluations, particularly for prompt injection defenses.', "Concrete examples of how you've scaled evaluation quality using automation or agent-based tools.", 'Your understanding of the gaps in current model robustness and how your methodologies address them.', "Why you're excited to work at a startup focused exclusively on RL environments for frontier AI labs."]

Generate Cover Letter โ†’

๐Ÿ” Research Before Applying

To stand out, make sure you've researched:

  • โ†’ Explore Trajectory Labs' website and any published materials to understand their specific approach to RL environments and red-teaming.
  • โ†’ Research recent papers or blog posts on prompt injection and AI red-teaming to speak fluently about current challenges.
  • โ†’ Look into frontier AI labs' safety initiatives and how they use evaluations to improve model robustness.
  • โ†’ Investigate common evaluation frameworks and tools (e.g., HELM, BIG-bench) to understand the landscape.
Visit Trajectory Labs, PBC's Website โ†’

๐Ÿ’ฌ Prepare for These Interview Topics

Based on this role, you may be asked about:

1 How would you design an evaluation to test a model's robustness against a novel prompt injection technique?
2 Describe a time you had to calibrate your judgment across many red-teaming transcripts. How did you ensure consistency?
3 What agent-based automation tools have you built to scale evaluation review? Walk us through the architecture.
4 How do you stay updated on the latest red-teaming and prompt injection research, and how would you apply it here?
5 Given a set of eval results showing a model fails on certain robustness tests, how would you diagnose the root cause and improve the eval?
Practice Interview Questions โ†’

โš ๏ธ Common Mistakes to Avoid

  • Being too generic about AI safetyโ€”show you understand the nuances of red-teaming and evaluation design specifically.
  • Failing to demonstrate hands-on experience with LLMs or coding agents; this role requires technical fluency.
  • Overlooking the startup nature of the companyโ€”avoid emphasizing preference for large-team structures or rigid processes.

๐Ÿ“… Application Timeline

This position is open until filled. However, we recommend applying as soon as possible as roles at mission-driven organizations tend to fill quickly.

Typical hiring timeline:

1

Application Review

1-2 weeks

2

Initial Screening

Phone call or written assessment

3

Interviews

1-2 rounds, usually virtual

โœ“

Offer

Congratulations!

Ready to Apply?

Good luck with your application to Trajectory Labs, PBC!