Application Guide

How to Apply for Research Engineer, Value Persistence Through Reinforcement Learning

at Compassion Aligned Machine Learning

🏢 About Compassion Aligned Machine Learning

Compassion Aligned Machine Learning is a unique research organization dedicated to aligning AI with the well-being of all sentient beings, setting it apart from purely profit-driven labs. Working here means contributing to cutting-edge alignment research with a strong ethical focus, in a remote-first environment that values deep technical work and publication.

About This Role

This role focuses on a critical alignment question: do values instilled during midtraining survive subsequent reinforcement learning? You'll design and run experiments combining synthetic data midtraining with RLVR/GRPO post-training, and use interpretability to understand value erosion. Your findings will directly inform how we build robust, value-aligned AI systems.

💡 A Day in the Life

A typical day might start with a stand-up meeting to coordinate with remote teammates, followed by writing code for a new training run. You'd analyze results from previous experiments, perhaps using interpretability tools to visualize value drift, and spend time reading recent papers or discussing findings with colleagues. The afternoon could involve iterating on synthetic data generation or preparing a paper draft.

🎯 Who Compassion Aligned Machine Learning Is Looking For

  • -
  • S
  • t
  • r
  • o
  • n
  • g
  • b
  • a
  • c
  • k
  • g
  • r
  • o
  • u
  • n
  • d
  • i
  • n
  • d
  • e
  • e
  • p
  • l
  • e
  • a
  • r
  • n
  • i
  • n
  • g
  • a
  • n
  • d
  • e
  • x
  • p
  • e
  • r
  • i
  • e
  • n
  • c
  • e
  • w
  • i
  • t
  • h
  • t
  • r
  • a
  • i
  • n
  • i
  • n
  • g
  • l
  • a
  • r
  • g
  • e
  • l
  • a
  • n
  • g
  • u
  • a
  • g
  • e
  • m
  • o
  • d
  • e
  • l
  • s
  • ,
  • i
  • n
  • c
  • l
  • u
  • d
  • i
  • n
  • g
  • f
  • a
  • m
  • i
  • l
  • i
  • a
  • r
  • i
  • t
  • y
  • w
  • i
  • t
  • h
  • R
  • L
  • H
  • F
  • a
  • n
  • d
  • v
  • a
  • r
  • i
  • a
  • n
  • t
  • s
  • l
  • i
  • k
  • e
  • R
  • L
  • V
  • R
  • /
  • G
  • R
  • P
  • O
  • .
  • -
  • P
  • r
  • o
  • v
  • e
  • n
  • a
  • b
  • i
  • l
  • i
  • t
  • y
  • t
  • o
  • b
  • u
  • i
  • l
  • d
  • a
  • n
  • d
  • m
  • a
  • n
  • a
  • g
  • e
  • c
  • o
  • m
  • p
  • l
  • e
  • x
  • t
  • r
  • a
  • i
  • n
  • i
  • n
  • g
  • p
  • i
  • p
  • e
  • l
  • i
  • n
  • e
  • s
  • w
  • i
  • t
  • h
  • r
  • e
  • p
  • r
  • o
  • d
  • u
  • c
  • i
  • b
  • i
  • l
  • i
  • t
  • y
  • a
  • n
  • d
  • e
  • x
  • p
  • e
  • r
  • i
  • m
  • e
  • n
  • t
  • t
  • r
  • a
  • c
  • k
  • i
  • n
  • g
  • (
  • e
  • .
  • g
  • .
  • ,
  • u
  • s
  • i
  • n
  • g
  • t
  • o
  • o
  • l
  • s
  • l
  • i
  • k
  • e
  • W
  • e
  • i
  • g
  • h
  • t
  • s
  • &
  • B
  • i
  • a
  • s
  • e
  • s
  • ,
  • M
  • L
  • f
  • l
  • o
  • w
  • )
  • .
  • -
  • E
  • x
  • p
  • e
  • r
  • i
  • e
  • n
  • c
  • e
  • g
  • e
  • n
  • e
  • r
  • a
  • t
  • i
  • n
  • g
  • s
  • y
  • n
  • t
  • h
  • e
  • t
  • i
  • c
  • d
  • a
  • t
  • a
  • a
  • t
  • s
  • c
  • a
  • l
  • e
  • ,
  • w
  • i
  • t
  • h
  • a
  • t
  • t
  • e
  • n
  • t
  • i
  • o
  • n
  • t
  • o
  • c
  • o
  • n
  • t
  • r
  • o
  • l
  • c
  • o
  • n
  • d
  • i
  • t
  • i
  • o
  • n
  • s
  • a
  • n
  • d
  • q
  • u
  • a
  • l
  • i
  • t
  • y
  • .
  • -
  • R
  • e
  • s
  • e
  • a
  • r
  • c
  • h
  • m
  • i
  • n
  • d
  • s
  • e
  • t
  • :
  • c
  • o
  • m
  • f
  • o
  • r
  • t
  • a
  • b
  • l
  • e
  • w
  • i
  • t
  • h
  • o
  • p
  • e
  • n
  • -
  • e
  • n
  • d
  • e
  • d
  • q
  • u
  • e
  • s
  • t
  • i
  • o
  • n
  • s
  • ,
  • d
  • e
  • s
  • i
  • g
  • n
  • i
  • n
  • g
  • e
  • v
  • a
  • l
  • u
  • a
  • t
  • i
  • o
  • n
  • s
  • ,
  • a
  • n
  • d
  • u
  • s
  • i
  • n
  • g
  • i
  • n
  • t
  • e
  • r
  • p
  • r
  • e
  • t
  • a
  • b
  • i
  • l
  • i
  • t
  • y
  • t
  • o
  • o
  • l
  • s
  • t
  • o
  • a
  • n
  • a
  • l
  • y
  • z
  • e
  • m
  • o
  • d
  • e
  • l
  • i
  • n
  • t
  • e
  • r
  • n
  • a
  • l
  • s
  • .

📝 Tips for Applying to Compassion Aligned Machine Learning

1

Tailor your resume to highlight any prior work on value alignment, safety, or interpretability, even if it was a side project.

2

In your cover letter, explicitly connect your experience with midtraining and RL post-training to the research question of value persistence.

3

Show that you've read the company's research by mentioning a specific paper or blog post from their website.

4

If you have a GitHub or portfolio, include links to code that demonstrates your ability to build reproducible training pipelines.

5

Emphasize remote collaboration skills, as the team is distributed; mention any experience with async communication tools like Slack or Notion.

✉️ What to Emphasize in Your Cover Letter

- Your technical expertise in RL training and synthetic data generation, with concrete examples of past projects. - Your alignment with the company's mission of compassion for all sentient beings and your motivation to work on value persistence. - Your research approach: how you would tackle the question of value persistence through experiments and interpretability.

Generate Cover Letter →

🔍 Research Before Applying

To stand out, make sure you've researched:

  • → Read the company's website and any publications to understand their current research directions and philosophical stance.
  • → Look for any blog posts or papers by the founder or team on value alignment, to reference in your application.
  • → Explore recent work on value persistence and RL from other labs (e.g., Anthropic, DeepMind) to show you're up-to-date.
  • → Understand the basics of RLVR (Reinforcement Learning with Verifiable Rewards) and GRPO (Group Relative Policy Optimization) to discuss them confidently.
Visit Compassion Aligned Machine Learning's Website →

💬 Prepare for These Interview Topics

Based on this role, you may be asked about:

1 How would you design an experiment to test whether values instilled via midtraining persist after RL? What controls would you use?
2 Describe your experience with RLVR or GRPO; what challenges did you face and how did you overcome them?
3 How would you ensure the synthetic corpus is high-quality and unbiased for this research?
4 What interpretability methods would you use to detect value erosion, and why?
5 How do you approach reproducibility in your training pipelines, and what tools do you prefer?
Practice Interview Questions →

⚠️ Common Mistakes to Avoid

  • Don't submit a generic application that doesn't mention the specific research question of value persistence.
  • Avoid overemphasizing engineering over research; this role requires a scientific mindset.
  • Don't ignore the ethical mission; if you don't care about AI safety, this isn't the right fit, and it will show.

📅 Application Timeline

This position is open until filled. However, we recommend applying as soon as possible as roles at mission-driven organizations tend to fill quickly.

Typical hiring timeline:

1

Application Review

1-2 weeks

2

Initial Screening

Phone call or written assessment

3

Interviews

1-2 rounds, usually virtual

✓

Offer

Congratulations!

Ready to Apply?

Good luck with your application to Compassion Aligned Machine Learning!