Research Engineer, Value Persistence Through Reinforcement Learning
Compassion Aligned Machine Learning
Last seen in the source feed: Oct 02, 2026.
Source listings can change. Check the employer's page for current availability and application requirements.
Posted
Sep 08, 2026
Location
Remote
Type
Full-time
Compensation
$110004 – $110004 per year
Job description
Role details
- In this role, you'll research whether values instilled through midtraining persist after reinforcement learning.
- Build and run training pipelines combining midtraining on synthetic corpora with RLVR or GRPO post-training.
- Implement synthetic corpus generation with well-matched controls at required scale and quality.
- Run evaluations and interpretability methods to understand value persistence and erosion mechanisms.
- Maintain reproducible configurations, logging, and experiment tracking while co-authoring papers.
About
Inside Compassion Aligned Machine Learning
Compassion Aligned Machine Learning conducts research on aligning AI with the well-being of all sentient beings.