Research Lead, Evaluations and Benchmarks
Alice
Last seen in the source feed: Oct 09, 2026.
Source listings can change. Check the employer's page for current availability and application requirements.
Posted
Sep 14, 2026
Location
Remote (US)
Type
Full-time
Job description
Role details
Lead the design and delivery of AI safety benchmarks measuring frontier risks, shipping evaluations every two to three weeks.
- Own benchmark quality, taxonomy, and release decisions; read evals yourself and verify reproducibility across frontier labs.
- Direct in-house researchers and freelance subject-matter experts; manage timelines and coordinate with CTO office on quarterly roadmaps.
- Stay current on AI safety research ecosystem; maintain weekly contact with lab researchers and attend conferences.
- Requires PhD or Masters in computer science/ML plus 3+ years building safety evaluations for language models in production.
- Requires 5+ publications in AI safety/security with lead authorship on at least two; strong engineering and taxonomy-building skills.
About
Inside Alice
Alice is a trust, safety, and security AI company.