How to apply for Cyber Evaluations Engineer

Anthropic

About Anthropic

Anthropic is a frontier AI research and product company focused on alignment, policy, and security. The company publishes its research and takes public positions on AI safety, which is unusual for a lab at this scale. This role sits inside that safety work rather than in a general software function.

About the role

You would build and run evaluations that measure cyber-relevant capabilities and safeguard robustness in Anthropic's models. The work includes per-release robustness testing before launches, analysis of jailbreaks and prompt bypasses, and designing detection probes for cyber misuse in production. This role directly feeds the cyber policy team's detection architecture, so your results shape what ships and what gets blocked.

A typical day

A typical day could involve writing or tuning an eval, running it against a model checkpoint, and analyzing where the results diverge from expectations. You might spend part of the day building detection probes for cyber misuse and reviewing jailbreak data with the policy team. The job post does not describe daily routines in detail, so ask about team structure, release cadence, and how eval requests are prioritized during interviews.

Who Anthropic is looking for

  • Has hands-on experience designing and running security or capability evaluations, not just consuming existing benchmarks
  • Can write tooling to run and score evals, and is comfortable maintaining that tooling over time
  • Understands jailbreaks, prompt injection, and bypass techniques well enough to probe them systematically
  • Can turn policy lines into detection logic and measure precision and coverage of that logic
  • Communicates findings clearly to both engineers and policy stakeholders without overselling results

Tips for this application

  • Name specific evals or red-team exercises you have built or run, with what you measured and what you found.
  • Describe any detection or abuse-prevention systems you have worked on, including how you measured false positives and coverage.
  • Reference Anthropic's published work on evaluations, red-teaming, or cyber capabilities if you have read it, and say what you would test differently.
  • Show that you can write the tooling yourself. Link to code, internal tools, or scoring pipelines you have built.
  • Address the 80,000 Hours concern about working at a frontier AI lab directly in your application rather than avoiding it.

What to cover in your cover letter

['A concrete example of an evaluation you designed end to end, including the scoring method and how you reported results.', 'Your experience with jailbreaks, prompt bypasses, or safeguard robustness testing, with specifics on what you tested.', 'How you have turned a policy or rule into a working detection probe and measured its precision and coverage.', "Why you want to work on cyber evaluations specifically at Anthropic, given the company's stated mission and the concerns raised about frontier lab work."]

Draft a cover letter

Research before applying

  • Read Anthropic's published research on evaluations, red-teaming, and responsible scaling, and note what is missing on the cyber side.
  • Read the 80,000 Hours career review on working at an AI lab, linked in the job post, so you can speak to the concerns it raises.
  • Look at Anthropic's usage policy and any public statements on cyber misuse to understand what the policy team is translating into detection logic.
  • Check what is publicly known about Anthropic's model release process and per-release testing, so you can ask informed questions about the eval pipeline.
Anthropic website

Likely interview topics

Based on the job description, expect questions about:

  • How you would design a capability evaluation for a specific cyber-relevant skill and validate that it measures what you claim.
  • How you would test safeguard robustness ahead of a model launch, including what failure modes you would prioritize.
  • A walkthrough of a jailbreak or prompt bypass you have analyzed, and what detection logic you would build from it.
  • How you would measure precision and coverage of a detection probe over time, and what you would do when coverage drops.
  • How you would communicate a negative or ambiguous evaluation result to a policy team that needs a clear answer.
Practise interview questions

Common mistakes to avoid

  • Treating this as a general security engineering role. The job is evaluations and detection, not incident response or infrastructure hardening.
  • Claiming evaluation experience without naming the eval, the metric, or the result. Vague references to benchmarking are easy to spot.
  • Ignoring the ethical concerns raised in the job post. Not addressing them looks like you have not read the listing carefully.
  • Proposing detection ideas without a way to measure precision and coverage. The role explicitly requires measuring both.

Deadline

No deadline is listed. Roles without a deadline usually close once the employer has enough candidates, so apply soon if you are interested.