This listing is no longer available on Remote Impact

The original description remains here for reference. See the current related opportunities below or browse the newest remote impact jobs.

Safeguards Enforcement Analyst, Violence and Extremism

Anthropic

Location
Remote (US)
Type
Full-time
Compensation
$285,000 – $330,000 per year
Posted
Jul 27, 2026

Job description

  • Build workflows to assess AI model behavior and identify misuse attempts in violence and extremism.
  • Architect automated enforcement systems and review workflows that scale while maintaining high accuracy.
  • Develop evals measuring model performance and identifying regressions to inform policy improvements.
  • Review flagged content to drive enforcement decisions and surface emerging misuse patterns and gaps.

The difference you'll make

This role directly prevents AI-enabled violence and extremism by building systems to detect and mitigate misuse, ensuring AI technologies are used safely and ethically.

What makes you a great fit

  • Experience in trust and safety, content moderation, or enforcement operations.
  • Strong analytical skills to identify misuse patterns and emerging threats.
  • Ability to design scalable workflows and automated systems.
  • Knowledge of extremist movements and regulatory landscape.

Benefits

Competitive compensation and benefits package. Specific details not provided.

About Anthropic

Website

Anthropic is an AI safety company dedicated to building reliable, interpretable, and steerable AI systems. They focus on ensuring AI benefits society and avoids catastrophic risks.