AI Safety & Governance Full-time

Safeguards Enforcement Analyst, Violence and Extremism

Anthropic

Posted

Jul 27, 2026

Location

Remote (US)

Type

Full-time

Compensation

Up to $330000

Mission

What you will drive

  • Build workflows to assess AI model behavior and identify misuse attempts in violence and extremism.
  • Architect automated enforcement systems and review workflows that scale while maintaining high accuracy.
  • Develop evals measuring model performance and identifying regressions to inform policy improvements.
  • Review flagged content to drive enforcement decisions and surface emerging misuse patterns and gaps.

Impact

The difference you'll make

This role directly prevents AI-enabled violence and extremism by building systems to detect and mitigate misuse, ensuring AI technologies are used safely and ethically.

Profile

What makes you a great fit

  • Experience in trust and safety, content moderation, or enforcement operations.
  • Strong analytical skills to identify misuse patterns and emerging threats.
  • Ability to design scalable workflows and automated systems.
  • Knowledge of extremist movements and regulatory landscape.

Benefits

What's in it for you

Competitive compensation and benefits package. Specific details not provided.

About

Inside Anthropic

Visit site →

Anthropic is an AI safety company dedicated to building reliable, interpretable, and steerable AI systems. They focus on ensuring AI benefits society and avoids catastrophic risks.