AI Safety & Governance Full-time

Research Lead, Evaluations and Benchmarks

Alice

Posted

Sep 16, 2026

Location

Remote (US)

Type

Full-time

Mission

What you will drive

  • In this role, you'll ship benchmarks every 2-3 weeks measuring frontier AI safety risks and harms.
  • Own the quality bar by reviewing evaluations for novelty, clarity, and reproducibility across labs.
  • Manage the release process by directing freelancers and keeping researchers on schedule.
  • Set the quarterly roadmap with the CTO and research leads based on emerging needs.
  • Stay current with the AI safety ecosystem by reading research, maintaining lab contacts, and attending conferences.

Profile

What makes you a great fit

  • In this role, you'll ship benchmarks every 2-3 weeks measuring frontier AI safety risks and harms.
  • Own the quality bar by reviewing evaluations for novelty, clarity, and reproducibility across labs.
  • Manage the release process by directing freelancers and keeping researchers on schedule.
  • Set the quarterly roadmap with the CTO and research leads based on emerging needs.
  • Stay current with the AI safety ecosystem by reading research, maintaining lab contacts, and attending conferences.

About

Inside Alice

Visit site →

Alice is a trust, safety, and security AI company.