AI Safety & Governance
Full-time
Research Lead, Evaluations and Benchmarks
Alice
Posted
Sep 16, 2026
Location
Remote (US)
Type
Full-time
Mission
What you will drive
- In this role, you'll ship benchmarks every 2-3 weeks measuring frontier AI safety risks and harms.
- Own the quality bar by reviewing evaluations for novelty, clarity, and reproducibility across labs.
- Manage the release process by directing freelancers and keeping researchers on schedule.
- Set the quarterly roadmap with the CTO and research leads based on emerging needs.
- Stay current with the AI safety ecosystem by reading research, maintaining lab contacts, and attending conferences.
Profile
What makes you a great fit
- In this role, you'll ship benchmarks every 2-3 weeks measuring frontier AI safety risks and harms.
- Own the quality bar by reviewing evaluations for novelty, clarity, and reproducibility across labs.
- Manage the release process by directing freelancers and keeping researchers on schedule.
- Set the quarterly roadmap with the CTO and research leads based on emerging needs.
- Stay current with the AI safety ecosystem by reading research, maintaining lab contacts, and attending conferences.
About
Inside Alice
Alice is a trust, safety, and security AI company.