Application Guide

How to Apply for Research Lead, Evaluations and Benchmarks

at Alice

๐Ÿข About Alice

Alice is a trust, safety, and security AI company focused on measuring and mitigating frontier risks in language models. Unlike many AI safety orgs that are purely research-focused, Alice ships production-grade evaluations every 2-3 weeks, meaning your work directly impacts how frontier labs assess and deploy their models. This is a rare chance to lead benchmark development at the intersection of cutting-edge research and real-world safety impact.

About This Role

As Research Lead for Evaluations and Benchmarks, you will own the entire lifecycle of AI safety benchmarksโ€”from taxonomy design and quality control to release decisions and reproducibility verification across frontier labs. You'll lead a team of in-house researchers and freelance subject-matter experts, coordinate with the CTO office on quarterly roadmaps, and maintain weekly contact with external lab researchers. This role is pivotal because the benchmarks you ship will shape how the industry measures and mitigates frontier risks.

๐Ÿ’ก A Day in the Life

A typical day might involve reviewing the latest evaluation results from your team, verifying reproducibility by running code yourself, and meeting with freelance subject-matter experts to refine a taxonomy. You might also coordinate with the CTO office on quarterly roadmap priorities, attend a virtual conference on AI safety, or have a weekly call with a frontier lab researcher to discuss benchmark adoption. The pace is fast, with a constant focus on shipping high-quality evaluations every few weeks.

๐ŸŽฏ Who Alice Is Looking For

  • **PhD or Masters in CS/ML with 3+ years building safety evaluations for language models in production**โ€”not just academic research, but hands-on experience shipping evals that are used by frontier labs or similar high-stakes environments.
  • **Published AI safety/security researcher with 5+ publications, including at least two as lead author**โ€”your work should demonstrate deep engagement with frontier risk taxonomies and evaluation methodologies.
  • **Strong engineering and taxonomy-building skills**โ€”you can read and write code to verify reproducibility, and you have a track record of creating clear, actionable evaluation frameworks that others can adopt.
  • **Proven team leader and project manager**โ€”experience directing researchers and freelancers, managing timelines, and coordinating with executive stakeholders to deliver evaluations on a rapid 2-3 week cadence.

๐Ÿ“ Tips for Applying to Alice

1

**Highlight your production evaluation experience**: In your resume and cover letter, explicitly mention evaluations you've built that are used in production settings, including the specific risks they measure and how they've been adopted by labs or industry.

2

**Showcase lead-author publications**: List your 5+ publications with clear indication of lead authorship, and briefly describe how each relates to frontier risk evaluation or taxonomy development.

3

**Demonstrate rapid shipping ability**: Provide concrete examples of evaluations you've shipped on tight timelines (e.g., every 2-3 weeks) and how you managed quality control and reproducibility under that pressure.

4

**Emphasize cross-lab collaboration**: Mention any experience coordinating with multiple frontier labs or external researchers to verify reproducibilityโ€”Alice values this highly as they aim for cross-lab consistency.

5

**Tailor to Alice's trust & safety focus**: Research Alice's public materials and recent evaluations, then explain in your application how your background aligns with their mission and how you could contribute to their benchmark taxonomy.

โœ‰๏ธ What to Emphasize in Your Cover Letter

["**Your track record of shipping production safety evaluations**: Describe specific benchmarks you've led from design to release, including the risks measured and their impact.", "**Leadership and project management**: Detail how you've managed researchers and freelancers, coordinated with executives, and delivered on quarterly roadmaps.", '**Taxonomy and quality ownership**: Explain your approach to building taxonomies and ensuring benchmark quality, including how you verify reproducibility across labs.', '**Ecosystem engagement**: Mention your weekly contacts with lab researchers, conference attendance, and how you stay current on AI safety research to inform benchmark design.']

Generate Cover Letter โ†’

๐Ÿ” Research Before Applying

To stand out, make sure you've researched:

  • โ†’ **Alice's published evaluations and blog posts**: Understand their existing benchmark suite, taxonomy, and public stance on frontier risks to see how you can contribute.
  • โ†’ **Alice's leadership and CTO office**: Learn about the backgrounds of key leaders, especially those in the CTO office, to understand their priorities and how you'd coordinate on roadmaps.
  • โ†’ **Competitor and collaborator landscape**: Identify other orgs doing AI safety benchmarks (e.g., METR, ARC Evals) and how Alice differentiates or collaborates, so you can speak to the ecosystem in interviews.
  • โ†’ **Recent AI safety incidents and regulations**: Be prepared to discuss how emerging risks and policy changes might influence benchmark design and release decisions.
Visit Alice's Website โ†’

๐Ÿ’ฌ Prepare for These Interview Topics

Based on this role, you may be asked about:

1 **Designing a benchmark for a specific frontier risk**: You might be asked to outline a taxonomy and evaluation plan for a novel risk (e.g., deception, cyberoffense) and explain how you'd ensure reproducibility across labs.
2 **Managing rapid release cycles**: How would you handle shipping evaluations every 2-3 weeks while maintaining quality and coordinating with a distributed team?
3 **Verifying reproducibility**: Describe your process for reading evaluations yourself and verifying that results are consistent when run by different frontier labs.
4 **Leading researchers and freelancers**: Share examples of how you've directed a team, managed timelines, and handled underperformance or shifting priorities.
5 **Staying current in AI safety**: What recent developments in AI safety research have caught your attention, and how would you incorporate them into Alice's benchmark roadmap?
Practice Interview Questions โ†’

โš ๏ธ Common Mistakes to Avoid

  • **Focusing only on academic research**: Alice needs someone who has shipped production evaluations, so overemphasizing theoretical work without practical deployment experience is a red flag.
  • **Ignoring the rapid release cadence**: If you don't address how you'll manage the 2-3 week shipping cycle, you may seem unprepared for the pace.
  • **Neglecting leadership and coordination**: This role requires managing researchers and freelancers; failing to highlight team leadership or executive coordination experience could hurt your candidacy.

๐Ÿ“… Application Timeline

This position is open until filled. However, we recommend applying as soon as possible as roles at mission-driven organizations tend to fill quickly.

Typical hiring timeline:

1

Application Review

1-2 weeks

2

Initial Screening

Phone call or written assessment

3

Interviews

1-2 rounds, usually virtual

โœ“

Offer

Congratulations!

Ready to Apply?

Good luck with your application to Alice!