How to apply for Data Center Global Repairs Program Support

Anthropic

About Anthropic

Anthropic is a frontier AI research and product company working on alignment, policy, and security. It posts specific opportunities it considers high impact, and it links to external reading about the risks of working at a frontier AI lab so candidates can make an informed choice.

About the role

This role owns the end-to-end hardware repair program across Anthropic's data centers. You set repair SLAs, prioritization rules, escalation paths, and reporting, and you are accountable for repair turnaround time and compute returned to service. The scope covers server, GPU/accelerator, network, and optics break-fix, RMA and reverse logistics with OEMs and ODMs, and spares and repair inventory.

A typical day

The job post does not describe a daily routine. Expect to ask about the split between strategy work, vendor and site operations management, and reporting. You can also ask how the Repairs Lead works with site operations partners, engineering, and equipment partners day to day.

Who Anthropic is looking for

  • Has hands-on experience running hardware operations at scale in data centers, including HPC environments.
  • Has managed break-fix and RMA processes with OEMs and ODMs, plus reverse logistics and spares inventory.
  • Can define repair SLAs, prioritization rules, and escalation paths, and hold site operations partners and repair vendors to those standards.
  • Can read failure trends and drive upstream fixes with engineering and equipment partners.
  • Works well in complex, fast-paced environments and is comfortable being accountable for turnaround time across multiple sites.

Tips for this application

  • Name the exact hardware categories you have supported (servers, GPUs/accelerators, network gear, optics) and the scale (number of sites, racks, or units).
  • Quantify repair outcomes: average turnaround time, SLA attainment, compute returned to service, or cost per repair.
  • Describe your work with OEMs and ODMs on RMA and reverse logistics, including which vendors and what you changed in the process.
  • Show one example where you turned a failure trend into an upstream fix with engineering or an equipment partner.
  • Read the 80,000 Hours career review linked in the job post before applying, and be ready to explain your reasoning about working at a frontier AI lab.

What to cover in your cover letter

['Your direct experience defining or running a repair program across multiple data center sites.', 'Specific metrics for repair turnaround time and compute returned to service that you owned.', 'Your experience with RMA, reverse logistics, and spares inventory management with OEMs and ODMs.', 'How you have driven failure trends upstream into engineering or equipment changes.']

Draft a cover letter

Research before applying

  • Read the Anthropic mission statement and the linked 80,000 Hours career review on working at an AI lab.
  • Review Anthropic's published work on alignment, policy, and security to understand the company's priorities.
  • Check what is publicly known about Anthropic's data center and compute infrastructure plans.
  • Look up common HPC data center repair workflows, RMA processes, and optics/GPU failure modes to prepare for technical questions.
Anthropic website

Likely interview topics

Based on the job description, expect questions about:

  • How you would define global repair SLAs, prioritization rules, and escalation paths for a growing fleet.
  • A time you improved repair turnaround time or compute returned to service, and the numbers involved.
  • How you manage OEM and ODM relationships for RMA and reverse logistics, including disputes and delays.
  • How you set and enforce standards with site operations partners and repair vendors across regions.
  • How you turn failure data into upstream fixes with engineering and equipment partners.
Practise interview questions

Common mistakes to avoid

  • Applying without reading the linked career review and being unable to discuss concerns about working at a frontier AI lab.
  • Describing hardware operations in general terms without naming specific hardware, vendors, or scale.
  • Claiming ownership of repair SLAs or turnaround metrics without being able to explain how they were measured or improved.

Deadline

No deadline is listed. Roles without a deadline usually close once the employer has enough candidates, so apply soon if you are interested.