Safety Research Grants
Thinking Machines
Posted
Sep 11, 2026
Location
Remote
Type
Full-time
Compensation
$50000 - $50000
Deadline
⏰ Sep 25, 2026
Mission
What you will drive
- Up to $50,000 in credits for safety research on Thinking Machine's open-weight models.
- Directions include: differential defensive capabilities, hazardous data filtering, and tamper-resistant training.
- Investigate alignment failure modes from fine-tuning, reward hacking, and oversight gaming.
- Estimate worst-case and marginal risk. Forecast safety-relevant scaling trends.
- Use adversarial fine-tuning to test whether safeguards persist or merely suppress.
Profile
What makes you a great fit
- Up to $50,000 in credits for safety research on Thinking Machine's open-weight models.
- Directions include: differential defensive capabilities, hazardous data filtering, and tamper-resistant training.
- Investigate alignment failure modes from fine-tuning, reward hacking, and oversight gaming.
- Estimate worst-case and marginal risk. Forecast safety-relevant scaling trends.
- Use adversarial fine-tuning to test whether safeguards persist or merely suppress.
About
Inside Thinking Machines
Thinking Machines is an AI startup founded by Miri Murati.